Monitoring system and method for railway multi-k8s cluster

By deploying monitoring components and monitoring data management components in multiple K8s clusters of railway enterprises, the problem of excessive workload of monitoring components cannot be obtained before cluster exceptions is solved, and more effective troubleshooting and monitoring stability is achieved.

CN120034468AInactive Publication Date: 2025-05-23CHINA ACADEMY OF RAILWAY SCI CORP LTD +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510186816.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the multi-k8s cluster monitoring of railway enterprises, the operating status information before the cluster abnormality cannot be obtained, resulting in difficulty in troubleshooting. At the same time, the work burden of monitoring components is too heavy and the stability is poor.

Method used

Design a monitoring system for railway multi-k8s clusters, including deploying a monitoring component within each cluster to capture and store operating status information, and deploying a monitoring data management component outside all clusters, responsible for storing operating status information with storage durations exceeding preset durations.

Benefits of technology

By deploying monitoring components on the cluster and deploying monitoring data management components outside the cluster, the problem of not being able to obtain the operating status information before cluster exceptions is solved, and the workload of monitoring components is reduced, and the stability of monitoring components is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034468A_ABST
    Figure CN120034468A_ABST
Patent Text Reader

Abstract

The invention discloses a railway multi-k8s cluster monitoring system and method, and relates to the technical field of cluster monitoring, the monitoring system comprises a plurality of monitoring components and a monitoring data management component, the monitoring components are in one-to-one correspondence with clusters, the monitoring components are deployed in the clusters, the monitoring components are used for capturing and storing the running state information of the clusters, and the monitoring data management component is used for managing the running state information of the clusters. And the monitoring module is used for receiving and storing the running state information of which the storage duration exceeds a preset duration to the monitoring data management module, the monitoring data management module is deployed outside all clusters, the monitoring data management module is in communication connection with each monitoring module, and the monitoring data management module is used for receiving and storing the running state information. The problem that the stability of the monitoring component is poor due to the fact that the running state information before the cluster is abnormal cannot be obtained, troubleshooting cannot be performed, and the workload of the monitoring component itself is heavy can be solved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cluster monitoring technology, and in particular to a monitoring system and method for a railway multi-k8s cluster. Background Art

[0002] In the scenario of software development in railway enterprises, applications need to be deployed in containerized form. In this scenario, multiple k8s clusters need to be deployed. One k8s cluster is used to deploy related components of the railway development and testing platform (i.e., devops platform), which is called the management cluster. One k8s cluster is used for development, debugging, and testing before the software is officially launched, which is called the development and testing environment cluster. One or more k8s clusters are used to deploy production applications, which are called production clusters. Therefore, there will be scenarios where multiple k8s clusters coexist in railway enterprises.

[0003] Cluster monitoring is the best way to detect whether the cluster operation status is abnormal. Currently, the cluster monitoring methods used include the following two:

[0004] (1) Deploy monitoring components in each cluster separately, such as Prometheus+Grafana, to monitor the running status of each cluster separately. However, when using the above method, if a cluster fails and becomes inaccessible, the monitoring components on it will also be inaccessible. Therefore, it is impossible to obtain the running status information before the cluster abnormality, which will not be effective for troubleshooting.

[0005] (2) A unified monitoring component is deployed outside multiple clusters to monitor the operating status of each cluster. However, when the number of clusters is large, the workload of the monitoring component itself will be heavy, resulting in poor stability of the monitoring component. In other words, the stability of the monitoring component becomes the bottleneck of the entire system. Summary of the invention

[0006] The purpose of this application is to provide a monitoring system and method for a railway multi-k8s cluster, which can simultaneously solve the problem that it is impossible to obtain the operating status information before the cluster abnormality, which will not be effective for troubleshooting, and the workload of the monitoring component itself will be heavy, resulting in poor stability of the monitoring component.

[0007] To achieve the above objectives, this application provides the following solutions:

[0008] In a first aspect, the present application provides a monitoring system for a railway multi-k8s cluster, the monitoring system for a railway multi-k8s cluster comprising: a plurality of monitoring components and a monitoring data management component;

[0009] The monitoring component corresponds to the cluster one by one. The monitoring component is deployed inside the cluster. The monitoring component is used to capture and store the running status information of the cluster, and regularly transmit the running status information whose storage time exceeds the preset time to the monitoring data management component; wherein, the cluster is a k8s cluster;

[0010] The monitoring data management component is deployed outside all clusters. The monitoring data management component is communicated with each monitoring component respectively. The monitoring data management component is used to receive and store the running status information.

[0011] Optionally, the running status information includes: node data, Pod data and API Server data;

[0012] The node data includes: CPU occupancy, memory usage, number of CPU cores, running time, total memory, storage usage, storage IO, average storage waiting time, number of inbound / outbound UDP packets, and number of inbound / outbound TCP packets;

[0013] The Pod data includes: Pod name, Pod status, Node to which the Pod belongs, Pod IP address, IP address of the Node to which the Pod belongs, Pod container information, and memory usage;

[0014] The API Server data includes: request QPS, latency, running time, processing time, queue size, retry ratio and etcd access latency.

[0015] Optionally, the monitoring component includes: a monitoring data storage module, and the monitoring data storage module is used to store the running status information captured by the monitoring component in real time.

[0016] Optionally, multiple clusters are interconnected in a ring networking manner, and each cluster is connected to two adjacent clusters; the monitoring component is also used to capture the operating status information of two adjacent clusters of the cluster.

[0017] Optionally, the monitoring data management component includes: a monitoring aggregation module and an object storage module, the monitoring aggregation module is respectively connected to each monitoring component and the object storage module in communication;

[0018] The monitoring aggregation module is used to aggregate and deduplicate the running status information transmitted by each monitoring component, and transmit the aggregated and deduplicated running status information to the object storage module;

[0019] The object storage module is used to store the running status information after aggregation and deduplication.

[0020] Optionally, the monitoring data management component further includes: a cluster state topology management module, the cluster state topology management module is respectively connected to communicate with the API Server in the Master node in each cluster;

[0021] The cluster status topology management module is used to send health status polls to the health check endpoint of the API Server at set intervals. When the API Server returns a normal status, it indicates that the cluster status is normal. Otherwise, it indicates that the cluster status is faulty. The monitoring topology map is updated based on the judgment result of the cluster status; the monitoring topology map includes the connection relationship between all clusters and the fault status of each cluster.

[0022] Optionally, the monitoring data management component further includes: a monitoring query module, the monitoring query module is respectively connected to each monitoring component and the cluster state topology management module for communication;

[0023] The monitoring query module is used to determine whether the data query time corresponding to the query request is less than the preset duration when receiving a query request, and obtain a first judgment result; if the first judgment result is no, obtain the query result from the object storage module; if the first judgment result is yes, determine whether the data query cluster corresponding to the query request is faulty, and obtain a second judgment result; if the second judgment result is no, forward the query request to the monitoring component corresponding to the data query cluster, and receive the query result returned by the monitoring component corresponding to the data query cluster; if the second judgment result is yes, determine whether both adjacent clusters of the data query cluster are faulty; if not, forward the query request to the monitoring component corresponding to the target cluster that is not faulty among the two adjacent clusters of the data query cluster, and receive the query result returned by the monitoring component corresponding to the target cluster.

[0024] Optionally, the monitoring topology map is a topology map obtained by connecting nodes based on the connection relationship between each cluster and two adjacent clusters of the cluster, with each node having a fault attribute, which is used to indicate the fault status of the cluster corresponding to the node.

[0025] In the second aspect, the present application provides a monitoring method for a railway multi-k8s cluster, which is applied to the monitoring system for a railway multi-k8s cluster described in any one of the above, and the monitoring method for a railway multi-k8s cluster includes:

[0026] Acquire and store the operation status information transmitted by each monitoring component and whose storage time exceeds a preset time; the operation status information is the operation status information of the cluster captured by the monitoring component.

[0027] Optionally, the railway multi-k8s cluster monitoring method further includes:

[0028] When receiving a query request, determining whether the data query time corresponding to the query request is less than a preset time length, and obtaining a first determination result;

[0029] If the first judgment result is no, obtaining the query result from the object storage module;

[0030] If the first judgment result is yes, then determining whether the data query cluster corresponding to the query request is faulty, and obtaining a second judgment result;

[0031] If the second judgment result is no, forwarding the query request to the monitoring component corresponding to the data query cluster, and receiving the query result returned by the monitoring component corresponding to the data query cluster;

[0032] If the second judgment result is yes, then judging whether two adjacent clusters of the data query cluster are both faulty;

[0033] If not, the query request is forwarded to the monitoring component corresponding to the target cluster that is not faulty in the two adjacent clusters of the data query cluster, and the query result returned by the monitoring component corresponding to the target cluster is received.

[0034] According to the specific embodiments provided in this application, this application has the following technical effects:

[0035] The present application provides a monitoring system and method for a railway multi-k8s cluster, including: multiple monitoring components and a monitoring data management component, the monitoring components correspond to the clusters one by one, the monitoring components are deployed inside the cluster, the monitoring components are used to capture and store the running status information of the cluster, and regularly transmit the running status information whose storage time exceeds the preset time to the monitoring data management component, the monitoring data management component is deployed outside all clusters, the monitoring data management component is respectively communicated with each monitoring component, and the monitoring data management component is used to receive and store the running status information. The present application deploys a monitoring component inside each cluster to capture and store the running status information of the cluster, and deploys a monitoring data management component outside all clusters to store the running status information captured by each monitoring component and whose storage time exceeds the preset time. Compared with the method of deploying monitoring components separately only in each cluster, since a monitoring data management component is deployed outside all clusters, and the monitoring data management component is used to store the running status information whose storage time exceeds the preset time, even if the cluster fails and is inaccessible, and the monitoring components on it are also inaccessible, the running status information before the cluster abnormality can be obtained by accessing the monitoring data management component. It is beneficial to troubleshooting. Compared with the method of deploying a unified monitoring component outside multiple clusters, since a monitoring component is deployed inside each cluster and the monitoring component is used to capture and store operating status information, the monitoring data management components deployed outside the multiple clusters only need to store the operating status information whose storage time exceeds the preset time. There is no need to perform the capture work, which greatly reduces the workload of the monitoring data management components deployed outside the multiple clusters themselves, and further improves the stability of the monitoring components. This can simultaneously solve the problems of being unable to obtain the operating status information before the cluster abnormality, which will not be effective for troubleshooting, and the monitoring component itself will have a heavy workload, resulting in poor stability of the monitoring component. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0037] Figure 1 A structural diagram of a railway multi-k8s cluster monitoring system provided in Example 1 of the present application.

[0038] Figure 2 A schematic diagram of a ring network of four clusters provided in Example 1 of the present application.

[0039] Figure 3A schematic diagram of a ring network of five clusters provided in Example 1 of the present application.

[0040] Figure 4 A schematic diagram of the structure of a computer device provided in Example 3 of the present application. DETAILED DESCRIPTION

[0041] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0042] Example 1

[0043] In order to realize the monitoring of multiple k8s clusters in the environment, this embodiment provides a monitoring system for multiple k8s clusters in railways, such as Figure 1 As shown, the railway multi-k8s cluster monitoring system includes: multiple monitoring components and a monitoring data management component.

[0044] The monitoring component corresponds to the cluster one by one. The monitoring component is deployed inside the cluster. The monitoring component is used to capture and store the running status information of the cluster, and regularly transmit the running status information whose storage time exceeds the preset time to the monitoring data management component. Among them, the cluster is a k8s cluster, which can be a management cluster, a development and testing environment cluster, or a production cluster.

[0045] The monitoring data management component is deployed outside all clusters. The monitoring data management component is communicated with each monitoring component respectively. The monitoring data management component is used to receive and store the running status information.

[0046] In this embodiment, a monitoring component is deployed inside each cluster to capture and store the running status information of the cluster, and a monitoring data management component is deployed outside all clusters to store the running status information captured by each monitoring component and whose storage time exceeds the preset time. Compared with the method of deploying monitoring components only in each cluster, since a monitoring data management component is deployed outside all clusters and the monitoring data management component is used to store the running status information whose storage time exceeds the preset time, even if the cluster fails and is inaccessible, and the monitoring components thereon are also inaccessible, the running status information before the cluster abnormality can be obtained by accessing the monitoring data management component, which is helpful for subsequent troubleshooting. Compared with the method of deploying a unified monitoring component only outside multiple clusters, since a monitoring component is deployed inside each cluster and the monitoring component is used to capture and store the running status information, the monitoring data management components deployed outside the multiple clusters only need to complete the storage of the running status information whose storage time exceeds the preset time, and there is no need to perform the capturing work, which greatly reduces the workload of the monitoring data management components deployed outside the multiple clusters themselves and further improves the stability of the monitoring components.

[0047] In this embodiment, the running status information is the indicator parameter information of the cluster, and the running status information may include: node data, Pod data and API Server data. The node data is the relevant data of each node in the cluster. The cluster includes a Master node and multiple Node nodes. The Master node is used for management and control. The Node node is a workload node, which stores specific containers inside. The Pod data is the relevant data of each Pod in each Node node in the cluster. The Pod represents a process running in the cluster. The API Server data is the relevant data of the API Server in the Master node in the cluster. The API Server is the only entry for resource operations and provides authentication, authorization, access control and other mechanisms.

[0048] Node data may include: CPU occupancy, memory usage, number of CPU cores, running time, total memory, storage usage, storage IO, average storage waiting time, number of inbound / outbound UDP (User Datagram Protocol) packets, and number of inbound / outbound TCP (Transmission Control Protocol) packets.

[0049] Pod data may include: Pod name, Pod status, Node to which the Pod belongs, Pod's IP address, IP address of the Node to which the Pod belongs, Pod's container information, and memory usage.

[0050] API Server data can include: query per second (QPS), latency, run time, processing time, queue size, retry ratio, and etcd access latency.

[0051] A monitoring component is deployed in each cluster, and a monitoring data storage module is deployed in the monitoring component. The monitoring component is used to capture the running status information of the cluster from the cluster and store the captured running status information in the monitoring data storage module. At this time, the monitoring component of this embodiment includes: a monitoring data storage module, which is used to store the running status information captured by the monitoring component in real time.

[0052] Since the monitoring component will periodically transmit the operating status information whose storage time exceeds the preset time to the monitoring data management component, the monitoring data storage module of this embodiment can adopt an overwriting storage method, that is, after it is full, it will automatically overwrite the data with the longest storage time during the next storage. Of course, it is necessary to set the storage capacity of the monitoring data storage module to avoid overwriting the operating status information before the operating status information whose storage time exceeds the preset time is transmitted.

[0053] The monitoring data management component is deployed outside all clusters and is connected to all clusters separately. Specifically, the monitoring data management component is communicated with each monitoring component respectively, and the monitoring data management component is used to receive and store the operation status information.

[0054] This embodiment introduces multi-cluster networking, where multiple clusters are interconnected in a ring networking manner, such as three clusters forming a triangle. Figure 1 As shown, the four clusters form a quadrilateral, such as Figure 2 As shown, the five clusters form a pentagon, such as Figure 3 As shown, by analogy, each cluster is connected to two adjacent clusters. At this time, the monitoring component is used not only to capture the running status information of the cluster, but also to capture the running status information of the two adjacent clusters of the cluster.

[0055] In this embodiment, the monitoring data management component includes: a monitoring aggregation module and an object storage module, and the monitoring aggregation module is respectively connected to each monitoring component and the object storage module for communication.

[0056] The monitoring aggregation module aggregates and stores the operating status information of all clusters, that is, the monitoring aggregation module is used to aggregate and deduplicate the operating status information transmitted by each monitoring component, and transmit the aggregated and deduplicated operating status information to the object storage module.

[0057] The object storage module is used to store the running status information after aggregation and deduplication. Since the storage time of the running status information sent by the monitoring component to the monitoring data management component has exceeded the preset time, the object storage module stores the running status information whose storage time exceeds the preset time. The preset time can be two hours or three hours, which depends on the actual needs of the user.

[0058] At this point, the data collection and storage process of the monitoring system of the railway multi-k8s cluster in this embodiment is as follows:

[0059] (1) Cluster data collection

[0060] In addition to capturing the running status information of the cluster itself, the monitoring component of each cluster also captures the running status information of the two adjacent clusters of the cluster. Every five minutes, the running status information that has been stored for more than two hours is transmitted to the monitoring data management component deployed outside all clusters, and the monitoring data management component archives the historical data. It should be noted that five minutes and two hours are just examples and can be adjusted according to user needs.

[0061] (2) Cluster data screening, deduplication and storage

[0062] The monitoring aggregation module of the monitoring data management component aggregates and deduplicates the operating status information of all clusters received. Deduplication is the process of aggregating the operating status information obtained from monitoring the same cluster from different monitoring components and retaining only one copy. At the same time, the storage duration of the received operating status information will be determined. Operating status information within two hours will be directly removed to complete the screening. The object storage module of the monitoring data management component is used to store the operating status information after aggregation and deduplication.

[0063] In this embodiment, the monitoring data management component also includes: a cluster state topology management module, and the cluster state topology management module is respectively connected to the API Server in the Master node in each cluster.

[0064] The cluster status topology management module is used to configure the API Server information of all clusters. It can obtain the status information of whether the cluster is accessible by sending heartbeats at regular intervals. That is, the cluster status topology management module is used to send health status polls to the health check endpoint of the API Server at set intervals. When the API Server returns a normal status, it means that the cluster status is normal. Otherwise, it means that the cluster status is faulty. The monitoring topology map is updated based on the judgment result of the cluster status. The monitoring topology map includes the connection relationship between all clusters and the fault status of each cluster.

[0065] Specifically, the set time can be 30 seconds. At this time, a health status poll is sent to the health check endpoint of each cluster's API Server every 30 seconds to confirm whether the cluster is accessible. When the cluster's API Server returns a status of OK (i.e., normal), it indicates that the cluster is in a normal state. Otherwise, it is marked as a fault, which means that the cluster is in a faulty state.

[0066] Among them, a monitoring topology graph is formed by describing the topological relationship between all clusters, each cluster is a node on the monitoring topology graph, and the nodes corresponding to each cluster are connected to the nodes corresponding to the two adjacent clusters of the cluster, that is, the monitoring topology graph is a topology graph obtained by connecting the nodes based on the connection relationship between each cluster and the two adjacent clusters of the cluster, with each cluster as a node, and each node has a fault attribute, which is used to indicate the fault state of the cluster corresponding to the node, that is, whether the cluster corresponding to the node is faulty.

[0067] A monitoring data management component is deployed outside all clusters. The configuration file of the cluster state topology management module in the monitoring data management component stores the topological relationship between all clusters, describes the topological relationship between all clusters, and forms a monitoring topology map. After the topological relationship between all clusters is described as a specific monitoring topology map, the cluster state topology management module in the monitoring data management component updates the available nodes in the monitoring topology map according to the heartbeat data packets returned by each cluster. If the fault attribute of the node is that the cluster corresponding to the node is not faulty, the node is available; otherwise, the node is unavailable.

[0068] In this embodiment, the monitoring data management component also includes: a monitoring query module, the monitoring query module is respectively communicated with each monitoring component and the cluster state topology management module, and the monitoring query module is used to determine whether the data query time corresponding to the query request (that is, the difference between the current time point and the time point of the data to be queried) is less than the preset time length when receiving the query request, and obtain a first judgment result; if the first judgment result is no, obtain the query result from the object storage module; if the first judgment result is yes, determine whether the data query cluster corresponding to the query request (that is, which cluster's data is to be queried) is faulty, and obtain a second judgment result; if the second judgment result is no, forward the query request to the monitoring component corresponding to the data query cluster, and receive the query result returned by the monitoring component corresponding to the data query cluster; if the second judgment result is yes, determine whether the two adjacent clusters of the data query cluster are both faulty; if not, forward the query request to the monitoring component corresponding to the target cluster that is not faulty in the two adjacent clusters of the data query cluster, and receive the query result returned by the monitoring component corresponding to the target cluster.

[0069] The monitoring data management component deployed outside all clusters in this embodiment is responsible for monitoring data (i.e., running status information) aggregation, query and status perception of all clusters. At this time, the cluster data query and health status feedback process of the monitoring system of the railway multi-k8s cluster in this embodiment is as follows:

[0070] The monitoring data query of multiple clusters is performed by accessing the monitoring query module in the monitoring data management component. When querying the monitoring data within two hours, the monitoring query module will forward the query request. For different fault points, the specific situation is as follows:

[0071] (1) When the data query cluster is in a healthy state, the query request is directly forwarded to the monitoring component of the data query cluster itself and the query result is returned.

[0072] (2) When the data query cluster is inaccessible due to a fault, the query request is forwarded to the monitoring component of any one of the two adjacent clusters of the data query cluster according to the monitoring topology information maintained by the cluster status topology management module, and the query result is returned.

[0073] (3) When the data query cluster and an adjacent cluster of the data query cluster fail at the same time, the query request is forwarded to the monitoring component of another healthy cluster of the two adjacent clusters of the data query cluster according to the monitoring topology information maintained by the cluster status topology management module, and the query result is returned.

[0074] When the monitoring data management component itself fails, the monitoring component in the cluster is directly accessed to obtain the query results.

[0075] In this embodiment, when performing data query, it is not completely completed by the monitoring data management component, but is completed in combination with the monitoring data management component and the monitoring component, which further reduces the burden of the monitoring data management component and improves the stability of the monitoring data management component.

[0076] In addition to monitoring the running status information of the cluster itself, the monitoring component of each cluster in this embodiment also monitors the running status information of other clusters. The specific rules are as follows: the topological relationship between multiple clusters is recorded in the form of a monitoring topology map. Each cluster serves as a node on the monitoring topology map. The nodes corresponding to each cluster are interconnected with the nodes corresponding to the two adjacent clusters of the cluster. In addition to monitoring its own running status information, the monitoring component of each cluster also monitors the running status information of the two adjacent clusters at the same time. The monitoring data management component describes the topological relationship between all clusters to form a monitoring topology map. The monitoring data management component can find the adjacent cluster that monitors the faulty cluster through the monitoring topology map, and send the monitoring data query request for the faulty cluster to the adjacent cluster. The running status information of the cluster can be obtained even if a cluster failure occurs, thereby further improving the stability of monitoring.

[0077] Example 2

[0078] This embodiment provides a monitoring method for a railway multi-k8s cluster, which is applied to the monitoring system for a railway multi-k8s cluster described in Embodiment 1. The monitoring method for a railway multi-k8s cluster includes: obtaining and storing operating status information transmitted by each monitoring component whose storage time exceeds a preset time, and the operating status information is the operating status information of the cluster captured by the monitoring component.

[0079] The monitoring method of the railway multi-k8s cluster of this embodiment also includes:

[0080] (1) When a query request is received, it is determined whether the data query time corresponding to the query request is less than a preset time length, and a first determination result is obtained.

[0081] (2) If the first judgment result is no, the query result is obtained from the object storage module.

[0082] (3) If the first judgment result is yes, determine whether the data query cluster corresponding to the query request is faulty to obtain a second judgment result.

[0083] (4) If the second judgment result is no, forward the query request to the monitoring component corresponding to the data query cluster, and receive the query result returned by the monitoring component corresponding to the data query cluster.

[0084] (5) If the second judgment result is yes, determine whether both adjacent clusters of the data query cluster are faulty.

[0085] (6) If not, forward the query request to the monitoring component corresponding to the target cluster that is not faulty in the two adjacent clusters of the data query cluster, and receive the query result returned by the monitoring component corresponding to the target cluster.

[0086] If both adjacent clusters of the data query cluster are not faulty, one of them is randomly selected as the target cluster.

[0087] Example 3

[0088] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a monitoring method for a railway multi-k8s cluster is implemented.

[0089] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0090] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the monitoring method of a railway multi-k8s cluster in Example 2 is implemented.

[0091] Example 4

[0092] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the monitoring method of a railway multi-k8s cluster in Example 2.

[0093] Example 5

[0094] In an exemplary embodiment, a computer program product is provided, including a computer program, which, when executed by a processor, implements the monitoring method of a railway multi-k8s cluster in Example 2.

[0095] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0096] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A railway multi-k8s cluster monitoring system, characterized in that: The railway multi-k8s cluster monitoring system includes: multiple monitoring components and a monitoring data management component; The monitoring component corresponds to the cluster one by one. The monitoring component is deployed inside the cluster. The monitoring component is used to capture and store the running status information of the cluster, and regularly transmit the running status information whose storage time exceeds the preset time to the monitoring data management component; wherein, the cluster is a k8s cluster; The monitoring data management component is deployed outside all clusters. The monitoring data management component is communicated with each monitoring component respectively. The monitoring data management component is used to receive and store the running status information.

2. The railway multi-k8s cluster monitoring system according to claim 1 is characterized in that: The running status information includes: node data, Pod data and API Server data; The node data includes: CPU occupancy, memory usage, number of CPU cores, running time, total memory, storage usage, storage IO, average storage waiting time, number of inbound / outbound UDP packets, and number of inbound / outbound TCP packets; The Pod data includes: Pod name, Pod status, Node to which the Pod belongs, Pod IP address, IP address of the Node to which the Pod belongs, Pod container information, and memory usage; The API Server data includes: request QPS, latency, running time, processing time, queue size, retry ratio and etcd access latency.

3. The railway multi-k8s cluster monitoring system according to claim 1 is characterized in that: The monitoring component includes: a monitoring data storage module, which is used to store the running status information captured by the monitoring component in real time.

4. The monitoring system for railway multi-k8s clusters according to claim 1 is characterized in that: Multiple clusters are interconnected in a ring networking manner, and each cluster is connected to two adjacent clusters; the monitoring component is also used to capture the operating status information of the two adjacent clusters of the cluster.

5. The railway multi-k8s cluster monitoring system according to claim 4 is characterized in that: The monitoring data management component includes: a monitoring aggregation module and an object storage module, and the monitoring aggregation module is respectively connected to each monitoring component and the object storage module in communication; The monitoring aggregation module is used to aggregate and deduplicate the running status information transmitted by each monitoring component, and transmit the aggregated and deduplicated running status information to the object storage module; The object storage module is used to store the running status information after aggregation and deduplication.

6. The railway multi-k8s cluster monitoring system according to claim 5 is characterized in that: The monitoring data management component also includes: a cluster state topology management module, which is connected to the API Server in the Master node of each cluster; The cluster status topology management module is used to send health status polls to the health check endpoint of the API Server at set intervals. When the API Server returns a normal status, it indicates that the cluster status is normal. Otherwise, it indicates that the cluster status is faulty. The monitoring topology map is updated based on the judgment result of the cluster status; the monitoring topology map includes the connection relationship between all clusters and the fault status of each cluster.

7. The railway multi-k8s cluster monitoring system according to claim 6 is characterized in that: The monitoring data management component also includes: a monitoring query module, which is respectively connected to each monitoring component and the cluster state topology management module in communication; The monitoring query module is used to determine whether the data query time corresponding to the query request is less than the preset duration when receiving a query request, and obtain a first judgment result; if the first judgment result is no, obtain the query result from the object storage module; if the first judgment result is yes, determine whether the data query cluster corresponding to the query request is faulty, and obtain a second judgment result; if the second judgment result is no, forward the query request to the monitoring component corresponding to the data query cluster, and receive the query result returned by the monitoring component corresponding to the data query cluster; if the second judgment result is yes, determine whether both adjacent clusters of the data query cluster are faulty; if not, forward the query request to the monitoring component corresponding to the target cluster that is not faulty among the two adjacent clusters of the data query cluster, and receive the query result returned by the monitoring component corresponding to the target cluster.

8. The railway multi-k8s cluster monitoring system according to claim 6 is characterized in that: The monitoring topology is a topology obtained by connecting nodes based on the connection relationship between each cluster and its two adjacent clusters, with each node having a fault attribute, which is used to indicate the fault status of the cluster corresponding to the node.

9. A monitoring method for a railway multi-k8s cluster, applied to a monitoring system for a railway multi-k8s cluster as claimed in any one of claims 1 to 8, characterized in that: The monitoring method of the railway multi-k8s cluster includes: Acquire and store the operation status information transmitted by each monitoring component whose storage time exceeds a preset time; the operation status information is the operation status information of the cluster captured by the monitoring component.

10. The method for monitoring railway multi-k8s clusters according to claim 9 is characterized in that: The monitoring method of the railway multi-k8s cluster also includes: When receiving a query request, determining whether the data query time corresponding to the query request is less than a preset time length, and obtaining a first determination result; If the first judgment result is no, obtaining the query result from the object storage module; If the first judgment result is yes, then determining whether the data query cluster corresponding to the query request is faulty, and obtaining a second judgment result; If the second judgment result is no, forwarding the query request to the monitoring component corresponding to the data query cluster, and receiving the query result returned by the monitoring component corresponding to the data query cluster; If the second judgment result is yes, then judging whether two adjacent clusters of the data query cluster are both faulty; If not, the query request is forwarded to the monitoring component corresponding to the target cluster that is not faulty in the two adjacent clusters of the data query cluster, and the query result returned by the monitoring component corresponding to the target cluster is received.

Citation Information

Patent Citations

  • Private cloud monitoring early-warning method and system

    CN109302324A

  • Cross-kubernetes cluster monitoring system and cross-kubernetes cluster monitoring method

    CN111459763A

  • Multi-cluster operation monitoring method, device and system and readable storage medium

    CN112667465A

  • Wide area cluster control system

    JP2002251384A

  • Cathode active materials and fabrication method of the same

    KR1020250063139A