Method and apparatus for monitoring network system, and computing device

By designing health measurement indicator models and algorithms, we have achieved group monitoring and personalized perspective display of large-scale network systems, solving the problem of lack of overall health measurement and visualization in traditional monitoring solutions, and improving management efficiency and monitoring effects.

WO2024221423A9PCT designated stage expired Publication Date: 2025-09-25BOE TECHNOLOGY GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/091680
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-28
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Monitoring solutions for large-scale network systems lack overall health measurement standards and intuitive visual displays, and are unable to meet the personalized monitoring perspective requirements of different roles, resulting in excessive alarm information and difficult management.

Method used

By designing health measurement indicator models and algorithms, we can achieve group monitoring of network systems, provide personalized perspectives for different roles, and display the health of various levels through the monitoring platform, including the health of systems, server clusters, services, etc.

Benefits of technology

It realizes scientific management of network systems, provides a basis for quickly locating problems, meets the monitoring needs of different roles, and improves the intuitiveness and efficiency of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023091680_25092025_PF_FP_ABST
    Figure CN2023091680_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and an apparatus for monitoring a network system, and a computing device. The network system comprises at least one resource cluster. The method comprises: in response to a health degree display triggering event, displaying health degrees of a monitored object on a first display interface, different health degrees corresponding to different marks, wherein the monitored object comprises at least one of the following: the network system, one or a plurality of resource clusters in the at least one resource cluster, one or a plurality of resources comprised in the network system, or one or a plurality of services provided by the network system. The health degrees of various levels can be displayed, and a scientific basis can also be provided for the response time of the operation and maintenance personnel based on the different health degrees, so that rapid problem positioning can be achieved, thereby solving the management problem.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus and computing device for monitoring network system Technical Field

[0001] The present application relates to the field of computers, and more specifically, to a method, apparatus, computing device, and medium for monitoring a network system. Background Art

[0002] The construction of large-scale projects, such as industrial parks, involves large-scale network systems. These systems involve clusters of multiple servers (virtual machines or physical machines) to provide the resources needed by the network system. These servers typically deploy or run containers, processes, ports, logs, plug-ins, and other components (collectively referred to as network application system assets) to provide various services.

[0003] Typically, it is necessary to access various assets in a network system in order to manage and monitor them, and to evaluate the operating status of various asset nodes in the network application system.

[0004] Due to the high complexity of large-scale network systems, including long network links, multiple operating environments, and a wide variety of monitored objects, traditional monitoring solutions meticulously monitor each asset node in the network system, often resulting in a large number of alarms. They lack holistic, scientific standards and algorithms for measuring the health of each level (system, server, service, etc.), lacking both local and global perspectives, and lack intuitive visualization methods. Furthermore, large-scale network systems involve a wide range of staff roles, such as developers, computer room operators, equipment operators, and application operators. Different personnel have different responsibilities and perspectives, and staff in different roles may only need to focus on specific types of monitored objects, resulting in different requirements for grouping monitored objects.

[0005] Therefore, a method is needed that can monitor and intuitively display the network system and can meet the personalized monitoring perspectives of different roles.

[0006] Summary of the Invention

[0007] According to one aspect of the present application, a method for monitoring a network system is provided, wherein the network system includes at least one resource cluster, and the method includes: in response to a health display trigger event, displaying the health of a monitored object on a first display interface, wherein different health levels correspond to different marks, and wherein the monitored object includes at least one of the following: the network system, one or more resource clusters in the at least one resource cluster, one or more resources included in the network system, or one or more services provided by the network system.

[0008] According to another aspect of the present application, a device for monitoring a network system is provided, wherein the network system includes at least one resource cluster, and the device includes: a data acquisition module for collecting various indicator data associated with the health of a monitored object from the network system and a public network to which the network system is connected; a monitoring server including: a display management module for displaying the health of a monitored object on a first display interface in response to a health display trigger event, wherein different health states correspond to different marks, and wherein the monitored object includes at least one of the following: the network system, one or more resource clusters in the at least one resource cluster, one or more resources included in the network system, or one or more services provided by the network system.

[0009] According to another aspect of the present application, a computing device is provided, comprising: one or more processors, and one or more memories on which a computer program is stored. When the computer program is executed by the one or more processors, the one or more processors can implement the steps of the method described above.

[0010] The embodiments of the present application provide a solution for monitoring network systems, which proposes the concept of health from the perspective of end users. By designing a health measurement indicator model and algorithm, it can ultimately display the health of various levels (for example, the comprehensive health of the system, and further refined server cluster health, server health, service health, public network portal health, etc.). Based on different health levels, it can also provide a scientific basis for the response time of operation and maintenance personnel, enable rapid positioning, and solve management problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To explain the principles of the present disclosure, embodiments of the present disclosure will be described in conjunction with the accompanying drawings. It should be understood that the elements shown in the figures may be implemented as various forms of hardware, software, or a combination thereof. Alternatively, these elements may be implemented in a combination of hardware and software on one or more appropriately programmed general-purpose computer devices.

[0012] FIG1 shows a schematic diagram of a monitoring scenario according to an embodiment of the present application.

[0013] FIG2 shows a schematic diagram of a resource configuration interface for implementing resource configuration according to an embodiment of the present application.

[0014] FIG3 shows a schematic diagram of an alarm configuration interface for implementing alarm policy configuration according to an embodiment of the present application.

[0015] FIG4 is a schematic diagram of an alarm details interface according to an embodiment of the present application.

[0016] FIG5 shows a schematic diagram of a health display interface showing health according to an embodiment of the present application.

[0017] FIG6 shows a schematic diagram of a picture viewing display interface according to an embodiment of the present application.

[0018] FIG7 shows a schematic diagram of a display interface for various indicator data values ​​of a resource according to an embodiment of the present application.

[0019] FIG8 shows a schematic diagram of a statistical data display interface according to an embodiment of the present application.

[0020] FIG9 shows a schematic flow chart of a method for monitoring a network system according to an embodiment of the present application.

[0021] FIG. 10 is a schematic diagram showing a display interface in which indicator data values ​​of multiple monitoring indicators of a focused resource cluster are overlaid on the display interface shown in FIG. 5 .

[0022] FIG11 shows a chart of all monitoring indicators required when calculating the comprehensive health of a network system according to an embodiment of the present application.

[0023] 12-14 are schematic diagrams of a process for calculating the comprehensive health of a network system according to an embodiment of the present application.

[0024] FIG15 shows a structural block diagram of a device according to an embodiment of the present application.

[0025] FIG16 shows a structural block diagram of a computing device according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] The present disclosure will be described more fully below with reference to the accompanying drawings, in which embodiments of the present disclosure are shown. However, the present disclosure can be implemented in many different forms, and the present disclosure should not be construed as being limited to the embodiments set forth herein. Throughout the text, similar reference numerals are used to represent similar elements.

[0027] The terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that when used herein, the term "comprising" specifies the presence of stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0028] Unless otherwise defined, the terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs. The terms used herein should be interpreted as having the same meaning as that in the context of this specification and the relevant art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such herein.

[0029] As mentioned above, large-scale network systems are involved in the construction of large-scale projects such as industrial parks. As an example and not a limitation, these large-scale projects can be applied to scenarios such as smart industrial parks, smart banks, or smart transportation, thus involving a variety of businesses: park search, passenger flow statistics, parking lot management, license plate recognition, face recognition, cross-border tracking, and alarm prompts. Each business is a service, corresponding to its own big data management, algorithm call, and model interface. Therefore, the network system used to provide these businesses or services has a complex layout and needs to include a large number of network resources (for example, servers) to implement. For example, for a smart industrial park, different departments or project groups can each have a corresponding set of network resources, and / or the different services provided (for example, face recognition services or parking lot management services) are also provided by their respective corresponding sets of resources.

[0030] At the same time, the performance, operating status, and health of the network system and its various resources are crucial to the final service or business. For example, if the performance of a set of resources providing facial recognition services degrades and the service cannot be provided normally, access control may not be possible. Therefore, it is necessary to monitor the network system and its resources, resource clusters, and / or services as monitoring objects to monitor the real-time status of the network system. In addition, other asset nodes of the network system (containers, processes, ports, logs, plug-ins, etc. deployed or running on the server providing resources) can also be monitored on demand to assist in determining the status of the network system.

[0031] Traditional monitoring solutions monitor each asset node in the network system in great detail, resulting in a large number of alarms, a lack of local and global perspectives, and a lack of intuitive visual display methods. They also fail to consider the different grouping requirements of staff with different roles for monitoring objects.

[0032] Therefore, the embodiment of the present application provides a solution for monitoring network systems, which groups resources by group and by label, and both labels and groups can be set in multiple ways, so as to meet the personalized monitoring perspectives of different roles and departments, and realize flexible grouping of resources; in addition, the embodiment of the present application also proposes the concept of health from the perspective of the end user, and by designing a health measurement index model and algorithm, it can ultimately display the health of each level (for example, the comprehensive health of the system, and further refined server cluster health, server health, service health, public network entrance health, etc.), based on different health levels, it can also provide a scientific basis for the response time of operation and maintenance personnel, and can quickly locate and solve management problems. In addition, in addition to displaying health, other indicator data of the system collected can also be displayed in different ways, as well as statistical data on the indicator data of the system's resources, etc.

[0033] In the context of this application, resources refer to the hardware and software configurations that can be provided by servers (physical and / or virtual machines) to implement various functions of the network system, such as various processing devices (e.g., CPU, MCU, DSP, ASIC, etc.) and various storage devices (e.g., memory or disk, etc.) in the server. Therefore, for ease of description, in some places in this application, resources may be used equivalently or interchangeably with servers.

[0034] In the context of this application, the performance, operating status, and health of each monitored object in the network system (for example, resources, resource clusters, services, public network entrances, etc.) depend on the corresponding influencing factors, and these influencing factors can be reflected or measured through corresponding indicators. Therefore, the performance, operating status, and health of the monitored object can be determined by collecting indicator data values.

[0035] The following will provide a detailed introduction to the solution for monitoring a network system according to an embodiment of the present application in conjunction with Figures 1 to 16.

[0036] FIG1 shows a schematic diagram of a monitoring scenario according to an embodiment of the present application.

[0037] FIG1 is a schematic diagram illustrating a scenario in which a monitoring platform 100 is used to monitor multiple network systems (e.g., located in the same campus), where the multiple network systems are represented by system A, system B, and system N. Although FIG1 shows only three network systems, it should be understood that a greater or lesser number of network systems is possible, and similar expressions below should be understood in this manner.

[0038] The network system monitoring solution described in this article can be implemented for each network system, meaning each network system is monitored independently. Furthermore, when determining system health based on monitoring data, the overall health of all network systems can be determined based on the system health calculated for each network system.

[0039] As shown in Figure 1, system A may include multiple server clusters (represented by cluster S1, cluster S2, and cluster S2, respectively), each of which includes the same or different numbers of servers. In this application, resources generally refer to servers, so server clusters can also be called resource clusters. In this application, servers can be cloud servers or local servers, physical servers or virtual servers. This application does not limit this. It is the basis for providing services, so servers can be regarded as resources.

[0040] Optionally, a server cluster is obtained by grouping multiple servers in a network system according to the type of service deployed. For example, each server (resource) is provided with a resource tag associated with the service type, so that a server cluster providing a certain service can be determined by tag screening. In this case, the server cluster is also called a service cluster. In addition, a server cluster can also be obtained by grouping multiple servers in a network system according to belonging to the same organization or having the same operating environment. For example, multiple servers can be grouped according to at least one dimension of department, project, or environment. In this way, each server cluster may provide multiple services.

[0041] The multiple network systems (system A, system B, system N) in FIG1 can be connected to a public network, for example, connected to a public network entrance via a secure connection network device, to perform data communication with the public network.

[0042] The monitoring platform 100 may include a monitoring server, a storage device (TSDB2), a processor, a database (such as redis and mysql), a plug-in, an interface, a program, an interface display component (for example, Web UI), etc. Optionally, the monitoring server may include at least a portion of these storage devices, processors, and databases, and may instead or additionally include additional processors and storage devices, etc., for implementing the functions of the various modules shown (for example, data acquisition module, health management module, etc.). In addition, the monitoring platform can also be connected to a display screen, which is used to display various interfaces. In this application, various interfaces may refer to content presentation pages that can provide content to users, including web pages, client interfaces, etc. The monitoring server can call various interfaces to obtain data from the outside. The monitoring platform 100 is used to collect various indicator data of the network system, and can process, analyze, calculate, display, etc. the collected indicator data.

[0043] For example, when collecting indicator data for servers in a network system (e.g., servers 1 / 2 / 3 in Figure 1), an agent can be deployed on each server in the network system to collect the server's indicator data, including relevant data such as the network, central processing unit (CPU) / memory / disk, etc. The agent then transmits the server's indicator data to the monitoring server of the monitoring platform (e.g., through a message queue or remote call push mode) and stores it in a corresponding storage device (e.g., TSDB2). The collected indicator data of various indicators are stored in the storage device in a predetermined data format.

[0044] For example, when collecting metrics data for services provided by a network system, different exporter plugins can be deployed based on the type of service. In some cases, these plugins are deployed independently, while in other cases, they are bundled with the corresponding service. The Prometheus program on the monitoring platform then uses the HTTP services exposed by various services to capture data and store it in the corresponding storage device (such as TSDB1). The monitoring platform server then calls the Prometheus interface via the HTTP protocol to transfer the monitored service-related data from TSDB1 to TSDB2.

[0045] In addition, for other indicator data of other monitoring objects of the network system (for example, containers, logs, processes on the server, etc.), corresponding programs, interfaces, etc. can be deployed on the monitoring platform in a similar manner to obtain the corresponding indicator data in collaboration with the programs, components, etc. at the monitoring objects.

[0046] In addition, the monitoring platform 100 can also obtain indicator data of the public network entrance, such as indicator data related to connectivity and bandwidth usage. For example, the method of collecting indicator data of the public network entrance can be similar to the method of collecting indicator data of the service. For example, the data can be collected through a dedicated agent or plug-in installed at the public network entrance, and then the monitoring server obtains the collected data from it.

[0047] The monitoring server or other processor of the monitoring platform 100 can process, analyze, calculate, and display the collected indicator data.

[0048] For example, the monitoring server of the monitoring platform may include a data acquisition module, a user and resource management module, an alarm management module, a health management module, and a display management module (for example, further divided into a picture viewing module, a monitoring dashboard module, and a large-screen display management module, etc.). In addition to the monitoring server, the monitoring platform may also include other plug-ins, interfaces, programs, etc. to facilitate data exchange with the outside world. Data can be transferred between modules, and at least a portion of each module can be integrated into another module, or can be further divided into more modules. Each module can be implemented by combining at least a portion of the hardware of the monitoring server included in the monitoring platform and the corresponding program or instructions stored therein. For example, each module can be implemented by a processor and a computer program stored in a memory. The processor may be, for example, a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor may be a microprocessor or any conventional processor, and may be an X84 architecture or an ARM architecture. The memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory, etc. It should be noted that the memory of the present application is intended to include, but is not limited to, these and any other suitable types of memory.

[0049] The data acquisition module can obtain the indicator data of each monitoring object and public network entrance collected by the method described above; the user and resource management module is used to realize the entry and role and grouping of users, and group the various resources in the network system (for example, various servers); the alarm management module is used to realize timely notification when an alarm event occurs; the health management module is used to determine the health of each level of the system or the health of the public network entrance, etc.; the display management module can be used for user interaction and / or the process of processing the information or data input by the user, and generate display data (graphics, images, text or charts, etc.) of various health of the network system, various indicator data, various statistical data of the entire network system, etc., and can cooperate with the UI component to display it to the user on the display interface of the display screen. In addition, the monitoring platform 100 can also provide functions for users to perform various configuration operations (for example, resource management configuration, alarm configuration, etc. to be described later). For example, the display management module can also present a corresponding configuration interface to the user on the display interface of the display screen, so that the user can enter information on the configuration interface, and then the display management module processes the information received from the user input.

[0050] For ease of explanation, Figure 2 shows a schematic diagram of the display management module cooperating with the user and resource management module to implement resource management. It should be noted that this is only a list of the modules that mainly perform the process, and does not exclude other related modules from also performing the process together.

[0051] As shown in FIG2 , a resource configuration interface is shown on the display screen (the user has successfully logged in). This interface can be in the form of a web page or a client interface. In response to user operations (e.g., through an external input device (mouse, keyboard, or trackball, etc.) or the user touching an option on the screen, etc.), resource groups are first created, and then each resource is mounted into the corresponding resource group. The left column of the resource configuration interface can show a list of each resource group that has been created. In response to the user's selection operation, each resource in the selected resource group can be displayed on the resource configuration interface. For example, grouping can be performed by at least one dimension of department, project, and operating environment.

[0052] Optionally, when displaying each resource, the resource configuration interface may also display information about each resource (including resource identifier, resource name, etc.). Optionally, each resource may also be assigned one or more resource tags (in the form of key=value) for logical grouping, thereby allowing resources that may not belong to the same department or project to be grouped together using the resource tags. For example, resource tags may be assigned based on the type of service the resource is deployed for.

[0053] Optionally, the monitoring platform can also implement graph-viewing operations related to resources or resource clusters (presenting indicator data values ​​of resources or resource clusters in the form of graphs, etc.), and / or obtain the content of ports, processes, plug-ins and logs on the server. In this way, in addition to displaying the grouping information of resources, the resource configuration interface can also display trigger prompts (for example, the trigger prompts can be corresponding text icons, etc.) for triggering display operations of graph-viewing operations related to resources or resource clusters (presenting indicator data values ​​of resources or resource groups in the form of graphs, etc.), ports, processes, plug-ins and logs, etc., so that the monitoring platform can also respond to respective display trigger events (for example, in response to the user's click operation on the corresponding icon or location area, such as the click operation on the "View Picture" text icon shown in Figure 2) and display the collected or processed relevant content accordingly.

[0054] Optionally, the content on the resource configuration interface shown in Figure 2 can switch with user operations (for example, the user slides up or down or drags a slider (not shown) for the display area of ​​the resource group, etc.), for example, when the display screen is small.

[0055] In addition, FIG3 shows a schematic diagram of the collaboration between the presentation management module and the alarm management module to implement alarm policy configuration.

[0056] As shown in Figure 3, the alarm configuration interface is shown on the display screen. The alarm configuration interface may include multiple input boxes or selection boxes for users to input or select information such as the alarm policy name, alarm level, trigger mode, statistical period, trigger conditions (supporting logical operators, functions, comparison operators, etc.), or resource filtering, notification method and recipient, etc., to form an alarm strategy for predetermined resources.

[0057] This demonstrates how label filtering can be used to select resources or resource clusters with predetermined labels from resources in the network system. Alarm policies can then be set for these resources or resource clusters. This means that different alarm policies can be applied to different resources or resource clusters. For example, different label settings can be used to set alarm policies for a specific resource, resource cluster, or even the entire network system.

[0058] For example, FIG3 schematically illustrates that the alarm policy is for the resource cluster of Project 1, and the alarm level is Level 2. The targeted resource cluster can be filtered by the label SYSTEM=IOT-CQST (e.g., resource labels by service type, and may also include additional labels indicating the organization to which the resource belongs). In addition, since the alarm policy here is still for the resources of Project 1, an additional label alertsource=Project1 is set, indicating that when the overall CPU usage of the resources in the resource cluster based on Project 1 and with the label SYSTEM=IOT-CQST (e.g., the average of the CPU usage of each of these resources) exceeds 85% within a statistical period (e.g., 60 seconds), the alarm trigger condition is determined to be met, and thus a trigger can be performed according to Mode 1 (e.g., the touch mode may include a nightingale trigger mode or a prometheus trigger mode). Alternatively, it can be set to determine that the alarm trigger condition is met when the overall CPU usage of these resources exceeds 85% within a statistical period (e.g., 60 seconds) (which can be implemented based on the happen function). Optionally, when configuring an alarm policy, you can also configure additional information about the recipients of the alarm notification.

[0059] Optionally, the content displayed on the alarm configuration interface shown in FIG3 can be switched with user operations (for example, the user slides up or down the display area or drags a slider (not shown), etc.), for example, when the display screen is small. Optionally, the alarm management module can also cooperate with the health management module, so that when an alarm event occurs and an alarm notification needs to be sent to the receiving personnel, the health of the resource or resource cluster targeted by the alarm event, or the comprehensive health of the network system, etc. (the health calculation method will be described later) is further attached to the alarm notification, so that the operation and maintenance personnel can understand the overall situation of the resource or resource cluster or network system targeted by the alarm event.

[0060] In addition, the display management module can also cooperate with the alarm management module to display alarm details when the triggering conditions of the alarm are met.

[0061] As shown in Figure 4, an alarm details interface is shown on the display screen. Optionally, the alarm details interface may correspond to the configured alarm strategy shown in Figure 3.

[0062] Assuming that the indicator data value of the monitoring indicator of the resource cluster meets the alarm trigger condition (for example, the overall CPU usage of the resource cluster is greater than 85% within the statistical period (for example, 60s)), the alarm details are displayed on the display interface. The alarm details may include alarm trigger condition information, the indicator data value that meets the alarm trigger condition, and related information of the resource cluster it is targeting.

[0063] For example, the upper part of Figure 4 shows the basic information of the current alarm event (for example, alarm level, event occurrence time, whether the trigger is completed, and notification personnel information), the expression of the alarm trigger condition, the label information of the resource cluster that meets the alarm trigger condition, and the group information (the resource cluster may belong to multiple groups), etc.

[0064] In addition, the lower part of Figure 4 shows information on the indicator data values ​​of the resource cluster that meets the alarm triggering conditions. Specifically, the information on these indicator data values ​​is shown in the form of a list and a chart, respectively. For example, when displayed in a list, the indicator data values ​​at the time points when multiple alarm events occur are shown in chronological order according to the occurrence of the alarm event (for example, the overall CPU usage of the resource cluster is greater than 85% within the statistical period (for example, 60s)); when displayed in a chart, the indicator data values ​​within a predetermined time period can be displayed according to the timeline, and a reminder mark (for example, a vertical dotted line) can be displayed at the time point when the alarm event first occurs.

[0065] Optionally, the content on the alarm details interface shown in FIG4 can be switched in response to user operations (e.g., the user swipes up or down on the display area or drags a slider (not shown), etc.), for example, when the display screen is small. Furthermore, the display management module can collaborate with the health management module to display the health of the monitored object (resource, resource cluster, provided service, or network system).

[0066] The calculation method of health will be described in detail later. The health management module can calculate the corresponding health based on the indicator data values ​​of each monitoring indicator (corresponding to each influencing factor related to the performance or operating status of the targeted resource or resource cluster) collected by the data acquisition module, and then the display management module displays the health, so that the operating status of the monitored object can be intuitively displayed on the display interface. For example, the display management module can determine the corresponding RGB value based on the current health calculated by the health management module, and display the current health using the color corresponding to the determined RGB value.

[0067] FIG5 is a schematic diagram showing a health status display interface showing health status.

[0068] As shown in Figure 5, the health display interface shows the health of four resource clusters. These four resource clusters can be grouped by belonging to the same organization and can correspond to the resources included in Project Group 1 of Department 1, Project Group 1 of Department 2, Project Group 2 of Department 1, and Project Group 2 of Department 2, respectively. These four resource clusters have different health levels (for example, 23, 93, 55, and 100 points, respectively, on a 100-point scale) and can be displayed using different colors.

[0069] In addition, the real-time status of the resource cluster (focus on the resource cluster) corresponding to project group 1 of department 1 is also shown on the right, that is, the health of each resource in the resource cluster is shown (for example, a numerical value plus a mark of a different color or a mark of a different shape, etc.), and additionally, more information about the server corresponding to each resource can also be shown, such as server ID, IP address, CPU count, memory, total disk size, etc.

[0070] In addition, the network system may include a large number of clusters, and the health display interface may only be able to display the health of a part of the clusters at a time. Therefore, the health display interface can also provide an interface for users to input switching information. For example, the user can slide in the area displaying the health of the resource cluster or click on the interface element indicating the switch below the area.

[0071] Additionally, in order to enable users to have a more comprehensive understanding of the various resources and their groupings of the network system, the total number of resource clusters of the network system (for example, 28 clusters) and the total number of resources (servers) (for example, 128 servers) can also be displayed on the health interface.

[0072] Optionally, the content on the health display interface shown in FIG5 can be switched with user operations (for example, the user slides up or down on the display area or drags a slider (not shown) and so on). In addition, when presented in the form of a web page, the real-time status of the resource cluster of interest on the right side of FIG5 can be independently displayed on another page (another display interface) that is different from the display page (first display interface) of the health of the resource cluster. In addition, the image viewing module in the display management module can also cooperate with the user and resource management module and the data acquisition module to display the indicator data value of a certain monitoring indicator of certain resources within a predetermined time period, which can be used for troubleshooting and analysis.

[0073] For example, a graph display interface can be displayed in response to user input (e.g., clicking the graph icon shown in FIG2 ). The user can enter information indicating the resources to be viewed on the graph display interface, and can filter by, for example, metric name, resource tag, and time period, thereby viewing the indicator data value of a monitoring metric of a resource cluster within a certain time period through the graph display interface. Of course, filtering can also be performed by metric name, resource group identifier, and time period.

[0074] Optionally, the indicator data values ​​of the monitoring indicators of each resource in the filtered resource cluster within a predetermined time period may be displayed in the form of a graph (eg, a line segment).

[0075] As shown in Figure 6, the picture viewing module of the display management module can display a picture viewing display interface for users to input monitoring indicators and time periods after processing the received input information. For example, the picture viewing display interface can be displayed in response to the selection operation of the "view picture" text icon shown in Figure 2 or other ways to trigger the display. An input box or selection box for user input can be displayed on the picture viewing display interface to allow the user to enter information about the monitoring indicator name, resource label and time period, as well as the number of line segments that can be displayed for each graph, where each line segment represents the change in the indicator data value of a resource within the expected time period.

[0076] The graph viewing module can respond to user input indicating the desired monitoring indicator, the desired resource tag (the indicated resource cluster is the desired monitoring object), and the desired time period, and after processing the indication information, display the indicator data value of the desired monitoring indicator for the resources in the resource cluster corresponding to the desired resource tag during the desired time period on the graph viewing display interface. For example, as shown in Figure 6, two line segments are shown indicating the change of the memory idle rate (monitoring indicator) of two resources (filtered by the resource tag) over time during the desired time period.

[0077] Similarly, the content on the image viewing display interface shown in FIG6 can switch with user operations, that is, not all the content shown in FIG6 can be displayed simultaneously. For example, an input box or selection box for user input can be displayed first in response to and after processing a user operation, and then, in response to and after processing a subsequent user operation, the indicator data value changes of each resource within a desired time period can be displayed, or the indicator data value changes can be displayed on another page when presented in the form of a web page.

[0078] In addition, the monitoring dashboard module in the display management module can also collaborate with the data acquisition module, user and resource management module to display the indicator data values ​​of various indicators of a certain monitoring object (for example, resources, resource clusters, services) and / or containers, middleware, etc.

[0079] FIG. 7 shows a display interface of indicator data values ​​of various indicators of a certain resource (server).

[0080] As shown in FIG7 , four graphs are used to respectively display the indicator data values ​​of various indicators (shown as the average system load and average load related to the system load indicator and the CPU usage rate and the user-mode CPU usage time ratio related to the CPU indicator, and other indicators may also be included in addition or in lieu thereof) of a predetermined resource (for example, specified by a resource identifier or IP address) within a predetermined time period. In this way, the various indicators of the resource can be displayed comprehensively. Optionally, before displaying the indicator data values ​​of the selected resource, it is possible to, for example, respond to a user operation for viewing the indicator data values ​​of the resource (the user operation is, for example, inputting or selecting the resource identifier of the desired resource (for example, the IP address of the server resource shown in the figure) on the interface shown in FIG7 , or inputting indication information for indicating the desired resource (for example, resource ID or resource name, etc.) on the interface of FIG2 , for example, clicking on the “Identifier 1” or “Resource 1” text icon in FIG2 ), and after data processing, the indicator data values ​​of one or more monitoring indicators of the desired resource are displayed on the display interface shown in FIG7 . In other examples, the display interface may also provide different ways for the user to input indication information of a desired monitoring object, so as to monitor various indicators of other types of monitoring objects (eg, resource clusters).

[0081] Similarly, the content on the display interface shown in FIG7 can switch in response to user operations (e.g., the user swipes up or down on the display area or drags a slider (not shown), etc.). For example, the first two charts can be displayed first, and then the remaining charts can be displayed in response to the user swiping down. Similarly, these contents can also be displayed on different pages (display interfaces).

[0082] In addition, the large-screen display management module in the display management module can also collaborate with the data acquisition module and the user and resource management module to count and display the indicator data values ​​of various indicators of multiple monitoring objects (for example, resources, resource clusters, services, containers, middleware, etc.) targeted by the monitoring platform and the resource-related quantities of the network system, thereby performing data statistics from the perspective of the entire monitoring platform.

[0083] Optionally, the large-screen display management module can display statistical data on a display screen that is different from the display screen used to display the aforementioned various display interfaces.

[0084] FIG8 shows a statistical data display interface, wherein the statistical data may include: the total number of projects in the last six months and the number of projects in each month, the total number of resource clusters in the last six months and the number of resource clusters in each month, the number of servers under each current project, the total number of servers, the current number of projects, the current number of clusters, the overall CPU usage at multiple recent time points, the overall memory usage at multiple recent time points, and the overall disk usage at multiple recent time points. Of course, the types of the multiple statistical data shown in FIG8 can be increased, decreased, or changed according to the content involved in the preset statistical data.

[0085] Likewise, the content on the display interface shown in FIG. 8 may switch in response to user operations (eg, the user slides the display area upward or downward or drags a slide bar (not shown) etc.).

[0086] Alternatively, the project-related quantities, resource cluster-related quantities, and server-related quantities can be obtained from the user and resource management module, and the various overall utilization rates can be the average of the corresponding utilization rates of all collected servers obtained by the data acquisition module. For example, as shown in FIG7 , the CPU utilization rate of a server can be obtained by obtaining the utilization rates of all servers at multiple time points and averaging them to obtain the overall CPU utilization rate shown in FIG8 .

[0087] By referring to the monitoring platform for monitoring network systems described in Figures 1-8, different resources of the network system can be grouped and monitored from different perspectives. For example, individual resources can be grouped according to tags or organizations to which they belong, so that each resource can correspond to multiple resource clusters based on the grouping method. This allows the resource clusters to be screened according to actual needs, meeting the personalized monitoring perspectives of different roles and organizations. In addition, different content is displayed through the display management module, and the content of the resources to be displayed can be screened. The concept of health is also proposed, which can intuitively provide an indication of the performance, operating status, or health of the system, resources, resource clusters, or services from the perspective of the end user, and realizes the monitoring of the network system from both local and global perspectives.

[0088] The following describes a method for monitoring a network system according to an embodiment of the present application. The method can be executed by the monitoring platform (eg, including a monitoring server) described with reference to FIG. 1 to FIG. 8 .

[0089] Figure 9 shows a flow chart of a method for monitoring a network system according to an embodiment of the present application. The network system may be a network system (eg, system A) shown in Figure 1 , and the network system may include at least one resource cluster.

[0090] Optionally, the at least one resource cluster is obtained by grouping resources included in the network system through a user and resource management module, and the resources included in each resource cluster have the same resource tag, or belong to the same organization, or have the same operating environment.

[0091] As shown in FIG9 , in step S910, in response to a health display trigger event, the health of a monitored object is displayed on a first display interface, where different health levels correspond to different tags. The monitored object may include at least one of the following: the network system, one or more resource clusters of the at least one resource cluster, one or more resources included in the network system, or one or more services provided by the network system.

[0092] For example, the relevant health can be calculated for each resource, resource cluster, network system, service, public network portal, etc., and as will be described later with respect to the health calculation method, the health of the resource cluster can be calculated based on the health of each resource included, the service health can be calculated from the health of the relevant resources or resource clusters that provide the service, and the comprehensive health of the system can be calculated based on the service health and the public network portal health.

[0093] That is to say, the embodiment of the present application can calculate a variety of different levels of health, and different health levels can be displayed on the first display interface according to actual needs.

[0094] For example, this step can be performed by the user and resource management module, the health management module, and the display management module shown in Figure 1. The health management module of the monitoring platform has calculated the health corresponding to the monitored object based on the indicator data of the monitored object obtained by the data acquisition module. Therefore, if the display management module determines that a health display trigger event occurs at this time, for example, receiving user input information (touching or selecting a display trigger-related button presented on the display interface through an input device, double-clicking or long pressing the current display area, sliding on the current display area, and other predetermined actions, etc.), in response to the startup of the monitoring platform or in response to the opening of a web page, etc., the processor corresponding to the display management module controls the operation of the UI component to present the calculated health on the display interface of the display screen.

[0095] Optionally, due to the limited size of the display interface, it may be possible to display only the health of some monitored objects at a time. Therefore, in response to an interface switching event (for example, detecting a user's sliding or clicking operation), the health of other monitored objects can be presented on the switched display interface.

[0096] In addition, since the health status of each monitored object may be different, different marks (for example, different colors, different font sizes, different shapes or different shape sizes, etc.) can be used to display different health status.

[0097] Optionally, the health management module can also display a reminder sign (for example, displaying a graphic next to it, or adjusting the font or graphic size, or adding a border, etc.) while displaying the health of the resource cluster in response to determining that the health of a monitored object does not meet the qualified threshold conditions.

[0098] Optionally, in step S920, while displaying the health of one or more resource clusters in the at least one resource cluster on the first display interface, the health of each resource included in the focus resource cluster in the one or more resource clusters is displayed on the first display interface or another display interface with the different marks.

[0099] Generally, as described in the health calculation method to be described later, when the health management module calculates the health of each resource cluster, it needs to be based on the calculated health of each resource in the resource cluster, that is, the health management module will also calculate the health of each resource in the resource cluster. In addition, the health of each resource in the resource cluster of interest can also be displayed with different marks (different colors) on the first display interface or another display interface. For example, the method may also include the display management module determining a corresponding mark from a plurality of candidate marks based on the health calculated by the health management module, and displaying the current health with the determined mark. For example, multiple threshold ranges can be set in advance, and different threshold ranges correspond to different marks, so that there are multiple candidate marks, and then the mark corresponding to the current health is determined based on the threshold range in which the current health is located.

[0100] When using different colors to display different health levels, RGB values ​​can be used to represent each color. In this case, the method may include: determining corresponding RGB values ​​based on the current health level of the monitored object to be displayed; and displaying the current health level using the color corresponding to the determined RGB value. RGB represents the three color channels of red, green, and blue. This standard encompasses nearly all colors perceptible by human vision and is one of the most widely used color systems. In addition to determining the color corresponding to the current health level based on the threshold range within which the current health level falls (each threshold range corresponds to a color with a preset RGB value), an alternative method is to determine the RGB value corresponding to the current health level using RGB(255*(1-SA), 255*SA, 0), where the health level SA is a value between 0 and 1, and displaying the current health level using the color corresponding to the determined RGB value. This eliminates the need to pre-set a correspondence between threshold ranges and colors; the corresponding RGB value for any health level can be calculated using a simple formula.

[0101] Optionally, the number of the focused resource clusters may be one or more.

[0102] Optionally, the resource cluster of interest is a resource cluster whose health does not meet a qualified threshold condition among the one or more resource clusters currently displayed, or a resource cluster with the lowest health among the one or more resource clusters currently displayed, or a resource cluster including more than a threshold number of resources whose health does not meet the qualified threshold condition, or a resource cluster determined in response to a user selection, etc. This application does not limit the method for determining the resource cluster of interest.

[0103] At this time, a schematic diagram of the first display interface can be shown in Figure 5, where the monitored object is a resource cluster as an example. As shown in Figure 5, the first display interface shows the health of four resource clusters. The four resource clusters can be grouped by belonging to the same organization and can respectively correspond to the resources included in Project Group 1 of Department 1, Project Group 1 of Department 2, Project Group 2 of Department 1, and Project Group 2 of Department 2. These four resource clusters have different health levels (for example, on a 100-point scale, they are 23, 93, 55, and 100 points, respectively) and can be displayed using different colors.

[0104] In addition, the right side of the first display interface also shows the real-time status of the resource cluster (focus resource cluster) corresponding to Project Group 1 of Department 1, that is, the health of each resource in the resource cluster is shown (for example, a numerical value plus a marker of a different color). In addition, more information about the server corresponding to each resource can also be shown, such as server ID, IP address, CPU count, memory, total disk size, etc. Of course, as mentioned above, the real-time status (for example, health) of the focus resource cluster can also be displayed on another display interface.

[0105] In addition, the network system may include a large number of resource clusters, and the first display interface may only be able to display the health of a portion of the resource clusters at a time. Therefore, the first display interface can also provide an interface for the user to input switching information. For example, the user can slide in the area displaying the health of the resource cluster or click on the interface element indicating the switch below the area.

[0106] In addition, in order to enable users to have a more comprehensive understanding of the various resources of the network system and their groupings, the total number of resource clusters of the network system (for example, 28 clusters) and the total number of resources (servers) (for example, 128 servers) can also be displayed on the first display interface or another display interface.

[0107] It can be seen that through this method, when monitoring the network system, the health of the monitored objects can be calculated and displayed, so that the user can intuitively see the health of each monitored object, and based on the health, the monitored objects with low health can be quickly located. When the monitored object is a resource cluster, the health and information of each resource in the resource cluster can also be displayed, and resources with low health can be quickly located.

[0108] Optionally, to more comprehensively display various indicator data collected by the resource cluster of interest, the method may further include step S930: while displaying the health of the one or more resource clusters, simultaneously displaying indicator data values ​​of one or more monitoring indicators of resources included in the resource cluster of interest on at least the first display interface or on another display interface. For example, this may be in response to a user selecting a resource cluster of interest.

[0109] For example, for each resource cluster, the performance, operating status or health of the resource cluster is associated with multiple influencing factors, among which closely related influencing factors are called key factors. For example, key factors may include: the connectivity and bandwidth utilization of the public network entrance to which the resource cluster is connected; the overall load and CPU utilization (for example, shown separately or as an average value) of each resource (server) of the resource cluster; the connectivity and role of the services provided by the resource cluster, etc. Therefore, these key factors can be given the highest priority, and after collecting the indicator data corresponding to these key factors, if the resource cluster is a focus resource cluster, the indicator data corresponding to these highest priority key factors can also be displayed. Or, more generally, all indicator data or part of the focus resource cluster can be displayed.

[0110] For example, compared with the first display interface shown in Figure 5, Figure 10 shows that while the health of multiple resource clusters is being displayed on the first display interface, the indicator data values ​​of multiple monitoring indicators (CPU usage, memory capacity and usage, real-time bandwidth / total bandwidth, etc.) of the resource cluster corresponding to project group 1 of department 1 (whose health is very low, as a resource cluster of concern) are further displayed on a portion of the display area corresponding to the health of the multiple resource clusters. Optionally, the hollow circles in the figure can refer to display marks when the health is low, and the solid circles can refer to the locations of each resource with low health in the resource cluster. In addition, in addition to covering at least the display area corresponding to the health on the first display interface, it can also be displayed on other display areas of the first display interface, and can be displayed in the form of a window. Optionally, the indicator data values ​​of the multiple monitoring indicators can also be displayed on another display interface. For example, when the first display interface is in the form of a web page, the indicator data values ​​of the multiple monitoring indicators can also be displayed on a new web page (considered as another display interface).

[0111] It can be seen that while intuitively seeing the health of each resource cluster, it is also possible to display the indicator data values ​​corresponding to the key factors of each resource in the resource cluster, so as to provide more reference information for operation and maintenance personnel and flexibly display user concerns.

[0112] In addition, as mentioned above, an alarm strategy can be set for certain monitoring objects, so that an alarm operation is performed when the indicator data of these monitoring objects meet the alarm strategy.

[0113] Therefore, the method may also include step S940: in response to the indicator data value of the monitoring indicator of a certain monitored object meeting the alarm trigger condition, displaying the alarm details on the second display interface, wherein the alarm details include at least the representation of the alarm trigger condition, the indicator data value that meets the alarm trigger condition and relevant information of the monitored object it targets.

[0114] For example, an alarm strategy can be configured on the configuration interface as shown in Figure 3 for important indicators of certain important resources. The second display interface for displaying alarm details can be as shown in Figure 4, and the monitoring object can be a resource cluster obtained by filtering through resource tags, which includes multiple resources. For example, when displaying the indicator data values ​​that meet the alarm trigger conditions in a list, the indicator data values ​​at multiple time points when the alarm event occurs within the statistical period can be displayed in the chronological order of when the alarm trigger conditions are met, that is, the alarm event occurs (for example, the overall CPU usage of the resource cluster is greater than 85% within the statistical period (for example, 60s)); when displaying the indicator data values ​​that meet the alarm trigger conditions in the form of a chart, the indicator data values ​​within multiple statistical periods can be displayed according to the timeline, and a reminder mark (for example, a vertical dotted line) can be displayed at the time point when the alarm event first occurs.

[0115] In addition, as described above, the indicator data values ​​of certain monitoring indicators of certain monitoring objects within a predetermined time period can be monitored, so the method can also include step S950: in response to the first interface switching event, a third display interface is displayed for the user to input the expected monitoring indicator, the expected monitoring object and the expected time period, and in response to the input information for the expected monitoring indicator, the expected monitoring object and the expected time period, the indicator data value associated with the expected monitoring indicator of the expected monitoring object within the expected time period is displayed on the third display interface or another display interface.

[0116] Optionally, this step can be performed collaboratively by the image viewing module, the user and resource management module, and the data acquisition module in the display management module described above. Optionally, the desired monitoring object can be a resource, a resource cluster, or a service.

[0117] For example, the first interface switching event may be receiving a user's selection operation for an icon related to viewing a picture. The third display interface may be the picture viewing display interface shown in FIG6 . For example, based on the input information for the expected monitoring indicator, the expected resource label and the expected time period, the indicator data value of the expected monitoring indicator of the resource corresponding to the expected resource label within the expected time period may be displayed on the picture viewing display interface. For example, as shown in FIG6 , two line segments are shown indicating how the memory idle rate (expected monitoring indicator) of two resources (filtered by the expected resource label) changes over time within the expected time period.

[0118] In addition, as previously described, the indicator data values ​​of the monitoring indicators of a certain monitoring object within a predetermined time period can be monitored, so the method can also include step S960: in response to the second interface switching event, displaying a fourth display interface for the user to input indication information of the desired monitoring object; in response to the user inputting indication information for the desired monitoring object, displaying the indicator data values ​​of one or more monitoring indicators of the desired monitoring object on the fourth display interface or another display interface. For example, the second interface switching event can be receiving a user selection operation for an icon used to guide the viewing of the indicator data values ​​of a certain monitoring object. Optionally, this step can be performed collaboratively by the monitoring dashboard module, the data acquisition module, and the user and resource management module in the display management module.

[0119] Furthermore, as previously described, statistics can be collected and displayed for various indicator data values ​​of multiple monitoring objects (e.g., resources, resource clusters, services, containers, middleware, etc.) of the network system. Therefore, the method can further include step S970: determining statistical data of the network system and displaying the statistical data on a fifth display interface, wherein the statistical data includes statistical data on quantities related to resources of the network system and statistical data on indicator data values ​​of monitoring indicators of the monitoring objects included in the network system. Optionally, this step can be collaboratively performed by the large-screen display management module, the user and resource management module, and the data acquisition module within the display management module described above.

[0120] For example, the fifth display interface can be the statistical data display interface shown in Figure 8. The statistical data may include: the total number of projects in the last six months and the number of projects in each month, the total number of resource clusters in the last six months and the number of resource clusters in each month, the number of servers under each current project, the total number of servers, the current number of projects, the current number of clusters, the overall CPU usage at multiple recent time points, the overall memory usage at multiple recent time points, and the overall disk usage at multiple recent time points. Of course, the types of the multiple statistical data shown in Figure 8 can be increased, decreased, or changed according to the content involved in the preset statistical data.

[0121] It can be seen that through the above additional steps, various indicator data or statistical data of each monitoring object of the network system can be intuitively displayed to the user at different angles according to user needs, and these data can be presented in the form of reports to improve readability.

[0122] In the aforementioned health management module, the health of the monitored object needs to be calculated. Therefore, the health calculation process will be described in detail below.

[0123] As previously described, embodiments of the present application may involve health at various levels (e.g., the overall health of the system, and further refined server cluster (resource cluster) health, server (resource) health, service health, public network portal health, etc.). The overall health of the network system can be determined based on service health and public network portal health, and service health can be determined based on the health of the servers providing the corresponding services.

[0124] Since the performance, operating status, and health of each monitored object (e.g., resources, resource clusters, systems, services, public network portals, etc.) depend on corresponding influencing factors, and these influencing factors can be measured by corresponding indicators, the performance, operating status, and health of the monitored object can be determined by collecting indicator data values.

[0125] Therefore, the monitoring indicators of different monitoring objects are introduced in conjunction with Figure 11. The various monitoring indicators shown in Figure 11 are all the monitoring indicators required when calculating the comprehensive health of the network system, and when calculating the health of other levels, the corresponding monitoring indicators required can also be determined based on Figure 11.

[0126] As shown in FIG11 , three types of influencing factors of the comprehensive health of the network system are shown: public network access, server, and service availability. For each type of influencing factor, a further secondary influencing factor is set, and each secondary influencing factor corresponds to a monitoring indicator.

[0127] For example, for public network access, we can consider two secondary influencing factors: connectivity and bandwidth utilization. We can then set two monitoring indicators for these two secondary influencing factors: network latency (unit: ms) and real-time bandwidth / total bandwidth. The monitoring indicators for server and service availability in Figure 11 are similarly set.

[0128] In addition, for each monitoring indicator, a health value and an unavailable value are set, where the health value can indicate that if the indicator data value of the monitoring indicator is less than or equal to the health value, the monitoring object for the monitoring indicator is healthy, and the unavailable value is used for subsequent health calculations. In addition, for the settings of monitoring indicators under the same type of influencing factors, the same trend is represented as the indicator data value increases or decreases. For example, the monitoring indicators under the influencing factors of the server class (for example, the percentage of occupied memory, the percentage of CPU computing time, etc.) decrease in health as the indicator data value increases. The range of the health value and unavailable value of each monitoring indicator is adjustable, and can be comprehensively considered based on the importance of the actual network system, the user's tolerance, cost, etc.

[0129] In addition, when it is necessary to determine the health of a single server, the various monitoring indicators corresponding to the secondary impact factors of the server in Figure 11 can be set; when it is necessary to determine the health of a server cluster, it can be obtained based on the health of each server included in the server cluster, and thus it is actually also obtained based on the various monitoring indicators corresponding to the secondary impact factors of the server in Figure 11; when it is necessary to determine the health of a service, it can be obtained based on the health of each server included in the server cluster that provides the service, and therefore it is also necessary to set the various monitoring indicators corresponding to the secondary impact factors of the server as shown in Figure 11 and the monitoring indicators related to connectivity and the number of service nodes respectively.

[0130] After explaining the monitoring indicators that need to be set when calculating different levels of health based on FIG11 , the specific calculation process is introduced below.

[0131] When calculating the comprehensive health of a network system, i.e., the monitored object is a network system, as shown in Figure 12 , method 900 may further include the following steps for calculating health. These steps may be performed by the health management module in Figure 1 . Furthermore, after the user and resource management modules group the network system, the health management module may also be aware of the resource grouping information.

[0132] In step S911, a first set of indicator data values ​​of public network portal-related indicators corresponding to the public network to which the network system is connected, a second set of indicator data values ​​of service monitoring-related indicators corresponding to the one or more services provided by the network system, and a third set of indicator data values ​​of server-related indicators corresponding to the servers included in the network system can be obtained.

[0133] In step S912, a first health of the public network portal may be determined based on the first set of indicator data values, and a second health of the one or more services may be determined based on the second set of indicator data values ​​and the third set of indicator data values.

[0134] In step S913 , the system health of the network system may be determined based on the first health and the second health.

[0135] For example, as described previously with reference to FIG11 , public network portal-related indicators can be defined based on network connectivity and / or bandwidth utilization, and exemplary indicators are shown in FIG11 as network delay and real-time bandwidth / total bandwidth, so that the health management module can determine the first health of the public network portal after obtaining the indicator data value of network delay and the indicator data value of real-time bandwidth and total bandwidth from the data acquisition module (their ratio can be further calculated in the health management module).

[0136] More specifically, the first health of the public network entrance can be determined by the following formula (1): min[(1-DET 1 / UNA 1), (1-DET 2 / UNA 2)] (1)

[0137] Where min() represents the minimum value, DET 1 represents the monitored network delay value 1, UNA 1 represents the unavailable value of the network delay shown in Figure 11, DET 2 represents the monitored ratio of real-time bandwidth to total bandwidth, and UNA 2 represents the unavailable value of the ratio of real-time bandwidth to total bandwidth. If DET 1 / UNA 1 or DET 2 / UNA 2 is greater than 1, the value is 1.

[0138] Alternatively, the first health level may be determined based on an average value of the two or other calculated values.

[0139] For another example, in conjunction with FIG11, the health of a single service among the one or more services provided by the network system can be determined based on the health of a server providing the service and an indicator data value of a service-related indicator, and the health of the server is determined based on the indicator data value of the server-related indicator. The service-related indicator is defined based on at least one of port connectivity, the number of nodes of a single service, and the role, and exemplary indicators are shown in FIG11.

[0140] More specifically, for example, the health SEV i of a single service i can be determined by the following formula (2): SEV i = sum (whether the port is connected * role value * server health s) / number of nodes (2)

[0141] Where sum() is a summation function, i is an integer greater than or equal to 1, s is an integer greater than or equal to 1 and less than or equal to p, p is the number of servers providing the service, and the value is 1 if the port is connected, otherwise it is 0. If the role type is master / none / all, the role value is 1, and if it is replica, the role value is 0. The number of nodes is the number of replicas of the service. Because a single service may be provided by multiple servers, after calculating the health of each of the p servers, we bring them into formula (2) to sum the multiple values ​​obtained to obtain the health of the single service i.

[0142] Therefore, after determining the health of a single service, the health SSEV (second health) of the service provided by the network system can be determined by formula (3): SSEV = min (SEV 1, SEV 2, ..., SEV N) (3)

[0143] Among them, min() is the minimum value, and N is the number of services provided by the network system. That is, the health of the service provided by the network system is the minimum value of the health of all services provided by the network system.

[0144] Optionally, an average value, a median value, or other calculated values ​​of the health of all services provided by the network system may be used as the second health.

[0145] In addition, server health reflects the adequacy of server resources, the busyness of the server, and the network connectivity between the server and related servers. It is designed to have a value range of 0 to 1, with the larger the value, the healthier it is.

[0146] Regarding the calculation process of the server health, for each server, a set of indicator data values ​​of server-related indicators corresponding to the server is obtained, and then the health of the resource cluster is determined based on the set of indicator data values.

[0147] In conjunction with Figure 11, server-related indicators can be defined based on at least one of the server's overall load, central processing unit (CPU), memory, storage capacity, operating system resources, mounted container CPU, and memory, and exemplary indicators are shown in Figure 11.

[0148] For example, the health of a single server i, SEVER i, can be determined by the following formula (4): SEVER i = min[(1-DET j / UNA j), 1≤j≤M] (4)

[0149] Where min() represents the minimum value, M represents the number of server-related indicators, DET j represents the monitored indicator data value, and UNA j represents the unavailable value of the corresponding indicator. If DET j / UNA j is greater than 1, the value is 1.

[0150] Optionally, an average value, a median value, or other calculated values ​​of the values ​​(1-DET j / UNA j) calculated for each indicator may be used as the health level SEVER i of a single server i.

[0151] In this way, after the health of each server is calculated, it can be substituted into formula (3) to obtain the second health of the service provided by the network system.

[0152] Therefore, after the first health and the second health are calculated, the comprehensive health of the system SC can be determined, for example, by formula (5): SC = min (first health, second health) (5)

[0153] Among them, min() is the minimum value.

[0154] On the other hand, when calculating the health of resource clusters included in the network system, that is, the monitoring object is the at least one resource cluster included in the network system, as shown in Figure 13, method 900 may also include the following steps of calculating the health. These steps may be performed by the health management module in Figure 1 and for each resource cluster.

[0155] In step S921, a set of indicator data values ​​of server-related indicators corresponding to the server providing the resources included in the resource cluster is obtained;

[0156] In step S922 , the health of the resource cluster is determined based on the set of indicator data values.

[0157] For example, the resource cluster may be a group of resources from the same organization (for example, the same department and the same project team), or a group of resources obtained by filtering resource tags (for example, used to provide the same service). The resources in the resource cluster are provided by multiple servers, so the health of the resource cluster can be regarded as the comprehensive health of all servers that provide resources in the resource cluster.

[0158] For example, when calculating the health of the resource cluster, the health of each server can be calculated according to the above formula (4), and then the minimum value, average value, or other calculated value of the health of all servers obtained can be used as the health of the resource cluster. When the resource cluster is used to provide the same service, the determination process is equivalent to determining the health of a single service as described above.

[0159] On the other hand, when calculating the health of one or more services provided by the network system, that is, the monitoring object is the one or more services provided by the network system, as shown in Figure 14, method 900 may also include the following steps of calculating the health. These steps may be performed by the health management module in Figure 1 and for each service.

[0160] In step S931, a set of indicator data values ​​of service-related indicators corresponding to the service and another set of indicator data values ​​of server-related indicators corresponding to the server providing the service in the network system are obtained; and

[0161] In step S932 , the health of the service is determined based on the set of indicator data values ​​and the other set of indicator data values.

[0162] For example, the health of a single service can be calculated by combining the above formula (2) and further combining it with formula (4). More details are similar to the above, so they will not be repeated here.

[0163] Similarly, when the monitoring object is one or more resources provided by the network system, the health of the server providing the resource can also be determined for each resource based on a set of indicator data values ​​of server-related indicators corresponding to the server providing the resource.

[0164] The following describes the process of calculating the comprehensive health of the network system in combination with the specific indicator data values ​​collected for each indicator.

[0165] FIG14 shows a table of specific indicator data values ​​collected for each indicator.

[0166] According to the indicator data values ​​in the table shown in Figure 14, according to formula (1), the health of the public network entrance can be calculated to be 0.4. According to formula (4), the health of each server (1-9) is calculated to be 0.7, 0, 0, 0.1, 0.4, 0, 0.1, 0 and 0.5 respectively. The network system provides three services. According to formulas (2) and (3) and combined with the health of each server, the service health provided by the network system is calculated to be min(0, 0.025, 0.25), that is, the service health is 0. Then, according to formula (5), the comprehensive health of the system is determined to be min(0.7, 0), that is, 0, indicating that the system is very unhealthy.

[0167] As mentioned above, the unavailable value of the corresponding indicator may be required when calculating the health level, and the unavailable value can be modified to adapt to various needs so that users can set it according to the needs of clusters, functions, server core concerns, etc. For example, the display management module can display a configuration interface on the display screen to receive user input information to configure or modify the health value and unavailable value of these monitoring indicators, as well as to set or modify the display priority of the secondary factors corresponding to each indicator. When the health level of a resource cluster does not meet the conditions, the indicator data values ​​of these indicators (as indicators corresponding to key factors) of the resource cluster are displayed on the display interface, as shown in Figure 10.

[0168] After calculating the health level, a corresponding RGB value can be determined based on the health level, and the color corresponding to the determined RGB value can be used to display the current health level. For example, if the health level SA obtained through the calculation process described above is a value between 0 and 1, it can be displayed using the color RGB(255*(1-SA),255*SA,0).

[0169] In summary, Figures 11-14 describe the calculation method for the health of different monitoring objects, which can present different levels of health, such as system health, resource cluster health, service health, etc., thereby realizing local and global perspectives for monitoring the system. Moreover, since health can reflect the availability and service quality of the network system and each resource, it can avoid excessive or delayed processing.

[0170] According to another aspect of the present application, a device for monitoring a network system is also provided.

[0171] FIG15 shows a block diagram of a device 1500 according to an embodiment of the present application. The device 1500 may be the monitoring platform shown in FIG1 .

[0172] As shown in FIG15 , apparatus 1500 may include a data collection module 1510 and a monitoring server 1520. Data collection module 1510 may be configured to collect various indicator data from a network system and the public network, such as indicator data associated with the health of a monitored object. The monitoring server may obtain various indicator data from data collection module 1520 and process, analyze, calculate, and display the data.

[0173] The data acquisition module 1510 may include at least a portion of a storage device (eg, memory), a processor, a plug-in, an interface, a program, and the like.

[0174] The monitoring server 1520 may be the monitoring server in the monitoring platform described in FIG. 1 to FIG. 8 .

[0175] Further, as shown in Figure 15, the monitoring server 1500 may include a display management module 1520-1, which is used to display the health of the monitored object on the first display interface in response to a health display trigger event, wherein different health levels correspond to different marks, and wherein the monitored object includes at least one of the following: the network system, one or more resource clusters in the at least one resource cluster, one or more resources included in the network system, or one or more services provided by the network system.

[0176] Optionally, the display management module 1520-1 is also used to display the health of one or more resource clusters in the at least one resource cluster on the first display interface, and to display the health of each resource included in the focus resource cluster in the one or more resource clusters with the different marks on the first display interface or another display interface.

[0177] Optionally, the monitoring server 1520 may further include a data acquisition module 1520 - 2 and a determination module 1520 - 3 (eg, the health management module mentioned above).

[0178] When the monitored object is the network system, the data acquisition module 1520-2 can be used to collect a first group of indicator data values ​​of public network portal-related indicators corresponding to the public network to which the network system is connected, a second group of indicator data values ​​of service monitoring-related indicators corresponding to the one or more services provided by the network system, and a third group of indicator data values ​​of server-related indicators corresponding to the servers included in the network system, and the determination module 1520-3 can be used to determine the first health of the public network portal based on the first group of indicator data values, and determine the second health of the one or more services based on the second group of indicator data values ​​and the third group of indicator data values; and determine the system health of the network system based on the first health and the second health.

[0179] In the case where the monitored object is the at least one resource cluster included in the network system, the data acquisition module 1520-2 can be used to obtain, for each resource cluster in the at least one resource cluster, a set of indicator data values ​​of server-related indicators corresponding to the server that provides the resources included in the resource cluster, and the determination module can be used to determine the health of each resource cluster based on a set of indicator data values ​​for each resource cluster in the at least one resource cluster.

[0180] In the case where the monitoring object is the one or more services provided by the network system, the data acquisition module 1520-2 can be used to collect, for each of the one or more services, a set of indicator data values ​​of service-related indicators corresponding to the service and another set of indicator data values ​​of server-related indicators corresponding to the server providing the service in the network system; and the determination module 1530 can be used to determine the health of each service based on a set of indicator data values ​​for each of the one or more services and the other set of indicator data values.

[0181] In the case where the monitored object is one or more resources provided by the network system, the data acquisition module 1520-2 can be used to obtain a set of indicator data values ​​of server-related indicators corresponding to the server providing the resource for each of the one or more resources; and the determination module 1520-3 can determine the health of the server providing the resource based on a set of indicator data values ​​for each resource.

[0182] Optionally, the public network portal-related indicators are defined based on network connectivity and / or bandwidth utilization; the server-related indicators are defined based on at least one of the server's overall load, central processing unit (CPU), memory, storage capacity, operating system resources, onboard container CPU, and memory; and / or the service-related indicators are defined based on at least one of port connectivity, the number of nodes of a single service, and the role.

[0183] As described above with reference to Figures 1 to 8, the device 1500 or the monitoring server 1520 may also include other modules, such as an alarm management module, for displaying alarm details on a second display interface in response to the indicator data value of a monitoring indicator of a monitored object satisfying an alarm trigger condition, wherein the alarm details include at least a representation of the alarm trigger condition, the indicator data value satisfying the alarm trigger condition, and relevant information of the monitored object to which it is targeted.

[0184] For more details of the device 1500, please refer to the above description of the monitoring platform, so it will not be repeated here.

[0185] By referring to the device for monitoring the network system described in Figure 15, different contents are displayed through the display management module, and the concept of different levels of health is also proposed. It can intuitively provide indications of the performance, operating status or health of the system, resources, resource clusters or services from the perspective of the end user, and realize monitoring of the network system from local and global perspectives.

[0186] According to another aspect of the present application, a computing device is further provided, which may reside on the monitoring platform shown in FIG1 to FIG8 , for example, a monitoring server on the monitoring platform.

[0187] FIG16 shows a schematic block diagram of a computing device 1600 according to the fourth aspect of the present application.

[0188] As shown in FIG16 , a computing device 1600 may include one or more processors, one or more memories connected via a system bus, and optionally a network interface, an input device, and a display screen. Each memory may include a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computing device may store an operating system and may also store a computer program that, when executed by the processor, enables the processor to perform the various operations described in the aforementioned steps. The internal memory may also store a computer program that, when executed by the processor, enables the processor to perform the various operations described in the same steps.

[0189] Each processor can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X84 architecture or an ARM architecture.

[0190] The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. It should be noted that the memory of the methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0191] The display screen of the computing device 1600 can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computing device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the terminal housing, or an external keyboard, touchpad or mouse, etc.

[0192] The computing device 1600 may be a server. The server may be a cloud server, i.e., a standalone server or a server cluster or distributed system consisting of multiple servers, which may provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0193] According to another aspect of the present application, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the processor executes the steps of the method described above.

[0194] According to another aspect of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the method described above are implemented.

[0195] Although the present disclosure has been described in detail with respect to various specific example embodiments of the present disclosure, each example is provided by way of explanation rather than limitation of the present disclosure. Those skilled in the art, after obtaining an understanding of the foregoing, can easily make changes, variations, and equivalents to such embodiments. Therefore, the present invention does not exclude such modifications, variations, and / or additions to the present disclosure that would be apparent to those of ordinary skill in the art. For example, a feature illustrated or described as part of one embodiment can be used together with another embodiment to produce yet another embodiment. Therefore, it is intended that the present disclosure cover such changes, variations, and equivalents.

[0196] Specifically, although the figures of the present disclosure describe steps performed in a specific order for the purposes of illustration and discussion, the method of the present disclosure is not limited to the order or arrangement of the specific illustrations. The various steps of the above method can be omitted, rearranged, combined and / or adjusted in various ways without departing from the scope of the present disclosure.

[0197] It will be appreciated by those skilled in the art that various aspects of the present application may be illustrated and described by a number of patentable categories or situations, including any new and useful process, machine, product or combination of substances, or any new and useful improvements thereto. Accordingly, various aspects of the present application may be performed entirely by hardware, entirely by software (including firmware, resident software, microcode, etc.), or by a combination of hardware and software. The above hardware or software may be referred to as "data blocks," "modules," "engines," "units," "components," or "systems." In addition, various aspects of the present application may be represented by a computer product located in one or more computer-readable media, the product including computer-readable program code.

[0198] The above is an illustration of the present disclosure and should not be considered as a limitation thereof. Although several exemplary embodiments of the present disclosure have been described, it will be readily understood by those skilled in the art that many modifications may be made to the exemplary embodiments without departing from the novel teachings and advantages of the present disclosure. Therefore, all such modifications are intended to be included within the scope of the present disclosure as defined by the claims. It should be understood that the above is an illustration of the present disclosure and should not be considered as limited to the specific embodiments disclosed, and modifications to the disclosed embodiments and other embodiments are intended to be included within the scope of the appended claims. The present disclosure is defined by the claims and their equivalents.

Claims

1. A method for monitoring a network system, wherein the network system includes at least one resource cluster, the method comprising: In response to a health display trigger event, the health of the monitored object is displayed on the first display interface, wherein different health levels correspond to different marks. The monitoring object includes at least one of the following: the network system, one or more resource clusters in the at least one resource cluster, one or more resources included in the network system, or one or more services provided by the network system.

2. The method according to claim 1, further comprising: While displaying the health of one or more resource clusters in the at least one resource cluster on the first display interface, the health of each resource included in the resource cluster of interest in the one or more resource clusters is displayed with the different marks on the first display interface or another display interface.

3. The method according to claim 2, wherein: The resource cluster of interest is a resource cluster whose health does not meet the qualified threshold conditions among the one or more resource clusters, or a resource cluster with the lowest health among the one or more resource clusters, or a resource cluster in which the number of resources whose health does not meet the qualified threshold conditions exceeds the quantity threshold, or a resource cluster determined in response to a user's selection.

4. The method according to claim 2, wherein: The health of each monitored object is associated with one or more monitoring indicators. The method further comprises: While displaying the health of the one or more resource clusters, indicator data values ​​of one or more monitoring indicators of resources included in the resource cluster of interest are displayed at least on the first display interface in an overlay manner or on another display interface.

5. The method according to claim 1, wherein The health of each monitored object is associated with one or more monitoring indicators. The method further comprises: In response to the indicator data value of a monitoring indicator of a monitoring object meeting the alarm trigger condition, the alarm details are displayed on the second display interface, The alarm details include at least the representation of the alarm triggering condition, the indicator data value that meets the alarm triggering condition, and relevant information of the monitoring object it targets.

6. The method according to claim 1, wherein The at least one resource cluster is obtained by grouping resources included in the network system, and the resources included in each resource cluster have the same resource tag, or belong to the same organization, or have the same operating environment.

7. The method according to claim 1, further comprising: In response to the first interface switching event, a third display interface is displayed for the user to input indication information for indicating a desired monitoring indicator, a desired monitoring object, and a desired time period. In response to the indication information input by the user, the indicator data value of the expected monitoring indicator of the expected monitoring object within the expected time period is displayed on the third display interface or on another display interface.

8. The method according to claim 1, further comprising: In response to the second interface switching event, displaying a fourth display interface for the user to input indication information for indicating a desired monitoring object; In response to the indication information input by the user, indicator data values ​​of one or more monitoring indicators of the desired monitoring object are displayed on the fourth display interface or on another display interface.

9. The method according to claim 1, further comprising: Determine statistical data of the network system and display the statistical data on a fifth display interface, wherein the statistical data include statistical data of resource-related quantities of the network system and statistical data of indicator data values ​​of monitoring indicators of monitoring objects included in the network system.

10. The method according to claim 1, wherein the monitoring object is the network system. The method further comprises: Obtaining a first set of indicator data values ​​of public network portal-related indicators corresponding to a public network to which the network system is connected, a second set of indicator data values ​​of service monitoring-related indicators corresponding to the one or more services provided by the network system, and a third set of indicator data values ​​of server-related indicators corresponding to servers included in the network system; Determining a first health of the public network portal based on the first set of indicator data values, and determining a second health of the one or more services based on the second set of indicator data values ​​and the third set of indicator data values; as well as The system health of the network system is determined based on the first health and the second health.

11. The method according to claim 1, wherein the monitoring object is the at least one resource cluster included in the network system. The method further comprises: For each resource cluster in the at least one resource cluster, Obtaining a set of indicator data values ​​of server-related indicators corresponding to servers that provide resources included in the resource cluster; A health of the resource cluster is determined based on the set of indicator data values.

12. The method according to claim 1, wherein the monitoring object is the one or more services provided by the network system. The method comprises: For each of the one or more services, Obtaining a set of indicator data values ​​of service-related indicators corresponding to the service and another set of indicator data values ​​of server-related indicators corresponding to the server providing the service in the network system; as well as A health of the service is determined based on the set of indicator data values ​​and the another set of indicator data values.

13. The method according to claim 1, wherein the monitoring object is the one or more resources provided by the network system. The method comprises: For each of the one or more resources, Obtaining a set of indicator data values ​​of server-related indicators corresponding to the server providing the resource; as well as Based on the set of indicator data values, a health of a server providing the resource is determined.

14. The method according to claim 11, wherein The public network access-related indicators are defined based on network connectivity and / or bandwidth usage; The server-related indicator is defined based on at least one of the server's overall load, central processing unit (CPU), memory, storage capacity, operating system resources, mounted container CPU, and memory; and / or The service-related indicator is defined based on at least one of port connectivity, number of nodes of a single service, and role.

15. The method according to claim 1, wherein The different marks are different colors, The method further comprises: Determine the corresponding RGB value based on the current health of the monitored object to be displayed; The current health level is displayed using a color corresponding to the determined RGB value.

16. A device for monitoring a network system, the network system comprising at least one resource cluster, the device comprising: A data collection module, configured to collect various indicator data related to the health of the monitored object from the network system and the public network to which the network system is connected; Monitoring servers, including: A display management module is used to display the health of the monitored object on the first display interface in response to a health display triggering event, wherein different health levels correspond to different marks. The monitoring object includes at least one of the following: the network system, one or more resource clusters in the at least one resource cluster, one or more resources included in the network system, or one or more services provided by the network system.

17. The device according to claim 16, wherein The display management module is also used to display the health of one or more resource clusters in the at least one resource cluster on the first display interface, and to display the health of each resource included in the focus resource cluster in the one or more resource clusters with the different marks on the first display interface or another display interface.

18. The device according to claim 16, wherein The health of each monitored object is associated with one or more monitoring indicators. The monitoring server further includes an alarm management module for displaying alarm details on a second display interface in response to an indicator data value of a monitoring indicator of a monitored object meeting an alarm trigger condition. The alarm details include at least the representation of the alarm triggering condition, the indicator data value that meets the alarm triggering condition, and relevant information of the monitoring object it targets.

19. The device according to claim 16, wherein the monitoring object is the network system. The monitoring server also includes: a data collection module, configured to collect a first set of indicator data values ​​of public network portal-related indicators corresponding to a public network to which the network system is connected, a second set of indicator data values ​​of service monitoring-related indicators corresponding to the one or more services provided by the network system, and a third set of indicator data values ​​of server-related indicators corresponding to servers included in the network system; A determination module is used to determine the first health of the public network portal based on the first set of indicator data values, and to determine the second health of the one or more services based on the second set of indicator data values ​​and the third set of indicator data values; and to determine the system health of the network system based on the first health and the second health.

20. The apparatus according to claim 16, wherein the monitoring object is the at least one resource cluster included in the network system. The monitoring server also includes: a data acquisition module, configured to obtain, for each resource cluster of the at least one resource cluster, a set of indicator data values ​​of server-related indicators corresponding to servers providing resources included in the resource cluster; The determining module is configured to determine the health of each resource cluster based on a set of indicator data values ​​for each resource cluster in the at least one resource cluster.

21. The apparatus according to claim 16, wherein the monitoring object is the one or more services provided by the network system. The monitoring server includes: a data collection module, configured to collect, for each of the one or more services, a set of indicator data values ​​of service-related indicators corresponding to the service and another set of indicator data values ​​of server-related indicators corresponding to the server providing the service in the network system; as well as A determination module is configured to determine a health of each service based on a set of indicator data values ​​for each service of the one or more services and the another set of indicator data values.

22. A computing device comprising: one or more processors, One or more memories having a computer program stored thereon, which, when executed by the one or more processors, enables the one or more processors to implement the steps of the method according to any one of claims 1 to 15.