Method, system, device, medium and product for state monitoring of data nodes
By introducing a control manager to actively monitor the target port status of data nodes in the Kubernetes system, the problem of false alarms and missed alarms caused by relying on heartbeat signal monitoring in existing technologies is solved, and the accuracy of data node status monitoring is improved.
Patent Information
- Application Number
- CN202411223127.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-02
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-09-02
AI Technical Summary
In existing technologies, data node fault monitoring methods in Kubernetes systems rely on heartbeat signals periodically sent by data nodes for monitoring. This is prone to false alarms or missed alarms due to network problems, resulting in the inability to monitor data node health information in a timely manner and compromising accuracy.
The control manager of the management node obtains the node status indicator data actively reported by the data node, determines whether the preset abnormal conditions are met, and determines that the data node is in an abnormal state. The control manager actively monitors the status of the target port to determine if the data node is in an abnormal state.
The accuracy of data node status monitoring has been improved. By actively monitoring the target through the control manager, the accuracy of data nodes in the existing technology has been solved, and the accuracy of data nodes has been improved. The accuracy of data node status monitoring has been improved by actively monitoring the target.
Smart Images

Figure CN118869543B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of communication, and particularly relates to a state monitoring method, system, device, medium and product of a data node. BACKGROUND
[0002] With containerized applications on the cloud, a containerized application management system (Kubernetes, K8S) is usually selected for cluster environment containerized deployment of business applications. The K8S architecture includes a management node and multiple data nodes, and the management node needs to monitor the faults of the data nodes to determine the health status of the data nodes.
[0003] In the related art, the way of monitoring the faults of the data nodes mainly relies on monitoring the health information in the heartbeat signals periodically reported and sent by the data nodes. However, when the heartbeat signals are periodically sent, false positives or false negatives may occur due to network problems, which may result in the health information of the data nodes not being monitored in a timely manner. In addition, the heartbeat signal monitoring component may also malfunction, resulting in inaccurate monitoring. Therefore, only relying on the data nodes to actively report the faults of the data nodes for monitoring cannot guarantee the accuracy.
[0004] In summary, how to design a more accurate method for monitoring the data nodes in the K8S system has become a technical problem to be solved by those skilled in the art. SUMMARY
[0005] The embodiments of the present application provide a state monitoring method, system, device and computer storage medium of a data node, which can not only rely on the data nodes to actively report the state of the data nodes for monitoring.
[0006] In one aspect, the embodiments of the present application provide a state monitoring method of a data node, applied to a management node, the management node including an application programming interface (API) service and a control manager, and the method includes:
[0007] acquiring, by the control manager, node state indicator data of a data node sent by the API service, wherein the state indicator data is reported by the data node to the API service;
[0008] determining, by the control manager, whether the node state indicator data meets a preset abnormal indicator condition to obtain a node state preliminary monitoring result of the data node;
[0009] In a case where the node state preliminary monitoring result indicates that the state of the data node is abnormal, the control manager monitors whether the state of a target port of the data node is abnormal, to obtain a port state monitoring result of the target port.
[0010] In a case where the port state monitoring result indicates that the state of the target port is abnormal, the control manager determines that the state of the data node is abnormal.
[0011] In another aspect, an embodiment of the present application provides a state monitoring system of a data node, which comprises:
[0012] a data node configured to report node state index data of the data node to a management node;
[0013] the management node comprises an application program interface (API) service and a control manager,
[0014] the API service is configured to receive the node state index data reported by the data node and send the node state index data to the control manager,
[0015] the control manager is configured to receive the node state index data sent by the API service, determine whether the node state index data satisfies a preset abnormal index condition, and obtain a node state preliminary monitoring result of the data node; in a case where the node state preliminary monitoring result indicates that the state of the data node is abnormal, monitor whether the state of a target port of the data node is abnormal, to obtain a port state monitoring result of the target port; and in a case where it is determined that the port state monitoring result indicates that the state of the target port is abnormal, determine that the state of the data node is abnormal.
[0016] In another aspect, an embodiment of the present application provides a state monitoring device of a data node, which comprises a processor and a memory storing computer program instructions.
[0017] The processor executes the computer program instructions to implement the above-described state monitoring method of a data node.
[0018] In another aspect, an embodiment of the present application provides a computer storage medium, which stores computer program instructions. The computer program instructions are executed by a processor to implement the above-described state monitoring method of a data node.
[0019] In another aspect, an embodiment of the present application provides a computer program product. Instructions in the computer program product are executed by a processor of an electronic device, so that the electronic device executes the above-described state monitoring method of a data node.
[0020] The state monitoring method of the data node provided by the embodiment of the present application is applied to a management node, the management node comprising an API service and a control manager, the control manager being used to acquire the node state index data of the data node sent by the API service, the node state index data being actively reported by the data node to the API service, the control manager actively monitoring the state of the target port of the data node in the case that the control manager preliminarily monitors the state of the data node according to the data node state abnormality indication data actively reported by the data node, and determining that the state of the data node is abnormal in the case that the port state of the target port is determined to be abnormal. In this way, the scheme of the embodiment of the present application no longer simply depends on the node state index data actively sent by the data node to determine the state abnormality, but can also actively monitor the state of the target port by the control manager to determine whether the data node is abnormal, thereby improving the state monitoring accuracy of the data node. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments of the present application will be briefly introduced as follows, and other drawings can also be obtained by those of ordinary skill in the art without creative labor on the premise of not paying creative labor.
[0022] Figure 1 is a specific structure schematic diagram of a containerized application management system;
[0023] Figure 2 is a flow schematic diagram of the state monitoring method of the data node provided by an embodiment of the present application;
[0024] Figure 3 is a flow schematic diagram of the state monitoring method of the data node provided by an embodiment of the present application;
[0025] Figure 4 is a schematic diagram of the interaction between the data node, the management node and the user end to implement the state monitoring method of the data node provided by an embodiment of the present application;
[0026] Figure 5 is a structure schematic diagram of the state device of the data node provided by another embodiment of the present application. DETAILED DESCRIPTION
[0027] The features and exemplary embodiments of various aspects of the present application will be described below in detail, in order to make the purposes, technical solutions and advantages of the present application more clear and apparent, the present application will be further described in detail below in combination with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, but not to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.
[0028] It should be noted that, in this paper, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or equipment. Without more limitations, the elements defined by the statement "include" do not exclude the presence of other identical elements in the process, method, article or equipment including the above elements.
[0029] With the trend of containerized applications moving to the cloud, more and more enterprises choose K8S for containerized deployment of business applications in cluster environment. The goal of K8S is to make it simple and efficient to deploy containerized applications, and K8S provides a complete set of functions such as resource scheduling, deployment management, service discovery, scaling, monitoring, maintenance, etc. It strives to become a platform for automatically deploying, expanding and running application containers across host clusters. Containerization technology is the mainstream trend of future technology development, and K8S, as the mainstream platform of containerization technology, business reliability is a problem that must be solved by each manufacturer, so this patent mainly optimizes the community K8S heartbeat abnormal isolation mechanism through reverse probing to improve business reliability.
[0030] As Figure 1 The specific structure diagram of the containerized application management system 100 is shown, which can include a management node (Master Node) 110 and a data node (Worker Node) 120, wherein the management node 110 can also be referred to as a master node, and the data node 120 can also be referred to as a business node.
[0031] It should be noted that the data node 120 can have multiple, Figure 1 Two data nodes are taken as an example in the above description, and the number of specific data nodes 120 can be set by the user as needed, which is not limited in the embodiments of the present application.
[0032] In the Kubernetes cluster, the connection relationship between the management node 110 and the data node 120 is crucial. As shown in FIG. 1, the management node 110 can include an API service (Kube-apiserver) 112, a scheduler (Kube-scheduler) 113, a control manager (Controller-Manager) 114 and a database (etcd) 111, and the management node 110 is responsible for managing and coordinating the state and operation of the entire cluster. The data node 120 runs actual application programs and workloads, and each data node 120 can have an executor (Kubelet) 121, a network proxy (Kube-proxy) 122 and a business application 123 corresponding to the data node, wherein the number of business applications 123 can be determined according to the actual function of the business node. Figure 1
[0033] The K8S architecture is a layered and highly scalable system architecture, which aims to provide high availability, scalability and fault tolerance to simplify the management and deployment of containerized applications. Users can interact with K8S through an API server, and describe the required state through declarative configuration. The K8S core component management node manages multiple data nodes, and the executor, controller, scheduler and network proxy all request data from the API gateway; the controller manages the state of all nodes and applications, and the executor manages the start and stop of multiple business applications.
[0034] In the related art, the way of monitoring the fault of the data node mainly relies on monitoring the health information in the heartbeat signal periodically sent by the data node. However, when the heartbeat signal is periodically sent, false positives or false negatives may occur due to network problems, which may result in the health information of the data node not being monitored in time. In addition, the heartbeat signal monitoring component may also fail, resulting in inaccurate monitoring. As can be seen, only relying on the heartbeat signal to monitor the fault of the data node cannot guarantee the accuracy.
[0035] To solve the problems in the prior art, the embodiments of the present application provide a data node state monitoring method, system, device, medium and product. The data node state monitoring method provided by the embodiments of the present application is applied to a management node, the management node includes an API service and a control manager, the control manager is used to acquire node state indicator data of a data node sent by the API service. The node state indicator data is actively reported by the data node to the API service. In a case where the control manager preliminarily monitors the node state and determines that the data node is abnormal according to the node state indicator data actively reported by the data node, the control manager actively monitors whether the state of a target port of the data node is abnormal. In a case where it is determined that the state of the target port is abnormal, it is determined that the state of the data node is abnormal. In this way, the embodiments of the present application no longer simply rely on the node state indicator data actively sent by the data node to determine the state abnormality, but also actively monitor the state of the target port by the control manager to determine whether the data node is abnormal, thereby improving the state monitoring accuracy of the data node.
[0036] Before introducing the data node state monitoring method provided by the embodiments of the present application, first introduce the data node state monitoring system for implementing the data node state monitoring method.
[0037] As shown in the above Figure 1 The data node state monitoring system provided by the embodiments of the present application can include a management node 110 and a data node 120.
[0038] The data node 120 is configured to report node state indicator data of the data node to the management node 110.
[0039] The management node 110 can include an API service 112 and a control manager 114. The API service 112 is configured to receive the node state indicator data reported by the data node 120 and send the node state indicator data to the control manager 114. The control manager 114 is configured to receive the node state indicator data sent by the API service 112, determine whether the node state indicator data meets a preset abnormal indicator condition, and obtain a node state preliminary monitoring result of the data node. In a case where the node state preliminary monitoring result indicates that the state of the data node is abnormal, the control manager 114 is configured to monitor whether the state of a target port of the data node is abnormal and obtain a port state monitoring result of the target port. In a case where it is determined that the port state monitoring result indicates that the state of the target port is abnormal, the control manager 114 is configured to determine that the state of the data node is abnormal.
[0040] The node state indicator data can be data reported by the data node to represent the state indicators of the data node. The node state indicator data can include, but is not limited to, the CPU usage rate, the memory usage rate, the disk usage rate, and the network delay of the data node. The specific node state indicator data can be selected according to user requirements, and is not limited in the embodiments of the present application.
[0041] The preset abnormal indicator condition can be a condition for determining the abnormality of the node state indicator data. The preset abnormal indicator condition can include, but is not limited to, the following: the CPU or memory usage rate is continuously too high, the disk usage rate is close to full load, significant network delay and error, node unreachable, Pod state continuously abnormal, disk I / O performance degradation or a large number of errors, and system load continuously too high.
[0042] It should be noted that the Pod is the smallest deployable unit in Kubernetes, which encapsulates one or more containers, and the containers share storage, network, and configuration information. The state of the Pod can include hard indicators, soft indicators, and node call volumes of the Pod.
[0043] It should be noted that the node state indicator data corresponds to the preset abnormal indicator condition, that is, the node state indicator data is the CPU usage rate, and the preset abnormal indicator condition is that the CPU usage rate is continuously too high. If the node state indicator data is the memory usage rate, the preset abnormal indicator condition is that the memory usage rate is continuously too high.
[0044] The node state preliminary monitoring result can be a preliminary monitoring result of the data node obtained according to the node state indicator data.
[0045] The target port can be a port of the data node. The target port can be a connection port between the control manager and the data node when the control manager actively monitors the data node.
[0046] In some embodiments of the present application, the data node and the management node can be connected in communication through a local port of the data node. Specifically, the API service in the management node can be connected with the local port of the data node to realize the communication connection between the API service and the data node.
[0047] The port state monitoring result can be a state monitoring result of the target port.
[0048] In some embodiments of the present application, the data node reports the node state indicator data to the API service through an executor in the data node.
[0049] In the embodiments of the present application, the data node actively reports the node state indicator data corresponding to the data node to the API service in the management node, and then the API service sends the node state indicator data to the control manager in the management node. The node state indicator data is obtained through the control manager. In the case that the control manager determines that the data node is in an abnormal state according to the preliminary monitoring result of the node state actively reported by the data node, the control manager actively monitors whether the target port of the data node is in an abnormal state. In the case that it is determined that the port state of the target port is abnormal, it is determined that the data node is in an abnormal state. In this way, the scheme of the embodiments of the present application no longer simply relies on the node state indicator data actively sent by the data node to determine the state abnormality, but also actively monitors the state of the target port through the control manager to determine whether the data node is abnormal, thereby improving the state monitoring accuracy of the data node.
[0050] First, the state monitoring method of the data node provided by the embodiments of the present application is introduced.
[0051] Figure 2 A flowchart of the state monitoring method of the data node provided by an embodiment of the present application is shown. The state monitoring method of the data node can be applied to the management node 110 in the above Figure 1 As shown in Figure 2 , the state monitoring method of the data node can include steps 210-240.
[0052] Step 210: Obtain the node state indicator data of the data node sent by the API service through the control manager.
[0053] The state indicator data can be reported by the data node to the API service.
[0054] Step 220: Determine whether the node state indicator data meets a preset abnormal indicator condition through the control manager to obtain a preliminary monitoring result of the node state of the data node.
[0055] In some embodiments of the present application, the node state indicator data described above can be the node state indicator data reported by the data node in at least two reporting periods in a preset time period. The preset time period can be a period for judging the state of the data node, which can be set by the user according to the user's demand, and is not limited in the embodiments of the present application.
[0056] The reporting period described above can be a period for reporting the node state indicator data by the data node.
[0057] To improve the accuracy of the data node monitoring, step 220 can specifically include:
[0058] determining, by the control manager, whether the node state indicator data obtained in each of the at least two reporting periods within the preset time period satisfies the preset abnormal indicator condition respectively;
[0059] determining, by the control manager, that the node state preliminary monitoring result is abnormal in a case where it is determined that the node state indicator data obtained in each of the at least two reporting periods within the preset time period satisfies the preset abnormal indicator condition.
[0060] In some embodiments of the present application, the control manager can determine whether the node state indicator data reported by the data node in each reporting period within the preset time period satisfies the preset abnormal indicator condition respectively, and determine that the node state preliminary monitoring result is abnormal in a case where it is determined that the node state indicator data reported by the data node in each reporting period within the preset time period satisfies the preset abnormal indicator condition.
[0061] In the embodiments of the present application, the node state preliminary monitoring result is determined to be abnormal in a case where it is determined that the node state indicator data reported by the data node in each reporting period within the preset time period satisfies the preset abnormal indicator condition, and the node state preliminary monitoring result thus determined is more accurate, thereby improving the accuracy of the data node monitoring.
[0062] Step 230: In a case where the node state preliminary monitoring result indicates that the state of the data node is abnormal, monitoring, by the control manager, whether the state of the target port of the data node is abnormal to obtain a port state monitoring result of the target port.
[0063] In some embodiments of the present application, in a case where it is determined that the node state preliminary monitoring result indicates that the state of the data node is abnormal, the control manager can automatically monitor whether the state of the target port of the data node is abnormal, rather than letting the data node actively report, so as to obtain the port state monitoring result of the target port.
[0064] In some embodiments of the present application, to further improve the accuracy of the port state monitoring result, step 230 can specifically include:
[0065] monitoring, by the control manager, the data node at least twice to obtain at least two port state monitoring results of the target port;
[0066] determining, by the control manager, the proportion of the first monitoring result and the second monitoring result in the at least two port state monitoring results;
[0067] The control manager determines the proportion of the monitoring result indicating the normal port state of the target port and the monitoring result indicating the abnormal port state of the target port in the first monitoring result and the second monitoring result.
[0068] The first monitoring result can be a monitoring result indicating a normal port state of the target port.
[0069] The second monitoring result can be a monitoring result indicating an abnormal port state of the target port.
[0070] In some embodiments of the present application, the control manager can monitor the data node multiple times to obtain the port state monitoring result of the target port in each monitoring, and then determine the proportion of the monitoring result indicating the normal port state of the target port and the monitoring result indicating the abnormal port state of the target port in the multiple port state monitoring results, and then take the monitoring result indicating the normal port state of the target port and the monitoring result indicating the abnormal port state of the target port as the port state monitoring result of the target port, thereby improving the accuracy of the port state monitoring result of the target port.
[0071] In some embodiments of the present application, in order to further improve the accuracy of the port state monitoring result, step 230 can specifically include:
[0072] The control manager sends a port monitoring signal of the target port to the data node;
[0073] The control manager receives port monitoring state index data fed back by the data node in response to the port monitoring signal;
[0074] The control manager determines whether the port monitoring state index data meets a preset abnormal index condition to obtain the port state monitoring result of the target port.
[0075] The port monitoring information can be a monitoring signal of the target port of the data node sent by the control manager to the data node.
[0076] The port monitoring state index data can be state index data of the target port fed back by the data node in response to the port monitoring signal. The port monitoring state index data can include disk I / O usage and network delay.
[0077] In some embodiments of the present application, when the control manager actively monitors the state of the data node, the control manager can first send a port monitoring signal of the target port to the data node, then receive port monitoring state index data fed back by the data node in response to the port monitoring signal, determine whether the port monitoring state index data meets a preset abnormal index condition, and then obtain the port state monitoring result of the target port.
[0078] In the embodiments of the present application, the control manager determines the port state monitoring result of the target port by first sending a port monitoring signal of the target port to the data node, then receiving the port monitoring state index data fed back by the data node in response to the port monitoring signal, and finally determining the port state monitoring result of the target port. Thus, the control manager must determine the port state monitoring result of the target port after receiving the port monitoring state index data fed back by the data node in response to the port monitoring signal, thereby avoiding monitoring the data node privately without the feedback of the data node, causing data leakage in the data node, and improving the security of the data in the data node.
[0079] In some embodiments of the present application, in order to improve the accuracy of the state monitoring of the target port, the sending of the port monitoring signal of the target port to the data node by the control manager can specifically include:
[0080] periodically and asynchronously sending the port monitoring signal of the target port to the data node by the control manager;
[0081] The receiving of the port monitoring state index data fed back by the data node in response to the port monitoring signal by the control manager can specifically include:
[0082] periodically and asynchronously receiving the port monitoring state index data fed back by the data node in response to the port monitoring signal by the control manager;
[0083] The monitoring of whether the state of the target port of the data node is abnormal by the control manager can specifically include:
[0084] periodically and asynchronously monitoring whether the state of the target port of the data node is abnormal by the control manager.
[0085] In some embodiments of the present application, the control manager can periodically send the port monitoring signal of the target port to the data node, then periodically receive the port monitoring state index data fed back by the data node in response to the port monitoring signal, and further periodically monitor whether the state of the target port of the data node is abnormal, thereby improving the accuracy of the state monitoring of the target port, and the sending of the port monitoring signal of the target port to the data node, the receiving of the port monitoring state index data, and the monitoring of whether the state of the target port of the data node is abnormal by the control manager can be asynchronously processed, thereby improving the efficiency of the state monitoring of the target port.
[0086] In some embodiments of the present application, in order to further determine the state of the data node, after step 230, the above-mentioned method can further include:
[0087] determining that the state of the data node is normal by the control manager in the case that the port state monitoring result of the target port indicates that the port state of the target port is normal.
[0088] In some embodiments of the present application, in a case where it is determined that the port state monitoring result of the target port indicates that the port state of the target port is normal, it is determined that the state of the data node is normal, that is, even if the node state indicator data actively reported by the data node determines that the state of the data node is abnormal, because it is determined that the state of the data node is suspected to be abnormal, not the final conclusion, the state of the target port of the data node needs to be actively monitored by the control manager, and if the state of the target port of the data node actively monitored by the control manager is normal, it is determined that the state of the data node is normal, that is, the actively monitored state of the control manager is used as the criterion.
[0089] In embodiments of the present application, in a case where it is determined that the port state monitoring result of the target port indicates that the port state of the target port is normal, it is determined that the state of the data node is normal, so that the state abnormality is no longer simply determined by relying on the node state indicator data actively sent by the data node, but the state of the target port is actively monitored by the control manager to determine whether the data node is abnormal, thereby improving the state monitoring accuracy of the data node.
[0090] Step 240, determining that the state of the data node is abnormal by the control manager in a case where it is determined that the port state monitoring result indicates that the port state of the target port is abnormal.
[0091] In some embodiments of the present application, in a case where it is determined that the port state monitoring result indicates that the port state of the target port is abnormal, it is finally determined that the state of the data node is abnormal.
[0092] In some embodiments of the present application, after step 240, the above-mentioned method can further include:
[0093] Sending the node state preliminary monitoring result of the data node to an application program interface service in the management node by the control manager for storage.
[0094] In some embodiments of the present application, the node state preliminary monitoring result of the data node can be sent to the API interface service for storage.
[0095] In embodiments of the present application, the node state preliminary monitoring result of the data node is sent to the API interface service for storage, so as to facilitate subsequent tracing of the node state preliminary monitoring result.
[0096] In some embodiments of the present application, after the node state preliminary monitoring result of the data node is sent to the application program interface service in the management node by the control manager for storage, the above-mentioned method can further include:
[0097] receiving, by the API service, an information acquisition request sent by the data node;
[0098] sending, by the API service according to the information acquisition request, the node state preliminary monitoring result stored in the API service to the data node, so that the data node determines a control signal for controlling the local port of the data node based on the node state preliminary monitoring result, and controls the opening and closing of the local port according to the control signal.
[0099] The information acquisition request can be a request sent by the data node to the API service to acquire the node state preliminary monitoring result.
[0100] The control signal can be a signal for controlling the opening and closing of the target port.
[0101] In some embodiments of the present application, the data node can send an information acquisition request to the API service to acquire the node state preliminary monitoring result, and then the API service sends the node state preliminary monitoring result stored therein to the data node based on the information acquisition request, so that the data node can determine a control signal for controlling the local port of the data node based on the node state preliminary monitoring result, and control the opening and closing of the local port according to the control signal. Specifically, if it is determined that the node state preliminary monitoring result indicates that the state of the data node is abnormal, the local port is closed to prevent the data node from continuing to send node state indicator data to the API service, thereby preventing data leakage of the data node. If it is determined that the node state preliminary monitoring result indicates that the state of the data node is normal, the local port is opened so that the data node can send node state indicator data to the API service.
[0102] In the embodiments of the present application, the API service receives the information acquisition request sent by the data node, and then sends the node state preliminary monitoring result stored in the API service to the data node according to the information acquisition request, so that the data node determines a control signal for controlling the local port of the data node based on the node state preliminary monitoring result, and controls the opening and closing of the local port according to the control signal. In this way, the security of the data of the data node can be ensured.
[0103] In some embodiments of the present application, after step 240, the above-mentioned method can further include:
[0104] The control manager stops receiving the node state indicator data within a preset time length.
[0105] The preset time length can be a time length set in advance after it is determined that the state of the data node is abnormal. The specific value of the preset time length can be set according to user demand, and is not limited in the embodiments of the present application.
[0106] In some embodiments of the present application, after determining that the state of the data node is abnormal, the control manager can stop receiving node state indicator data within a preset time period, thus avoiding the control manager from continuously receiving incorrect node state indicator data, thereby obtaining incorrect detection results and affecting the accuracy of data node monitoring.
[0107] In some embodiments of the present application, in order to better understand the data node state monitoring method provided by the embodiments of the present application, the data node state monitoring method provided by the embodiments of the present application is described in detail below with specific examples, with reference to Figure 3 , Figure 3 A flowchart of a data node state monitoring method provided by an embodiment of the present application is shown in the figure. The execution subject of the data node state monitoring method is the data node state monitoring system in the above Figure 1 , as shown in Figure 3 , the data node state monitoring method can specifically include steps 301-307.
[0108] Step 301: The executor in the data node sends node state indicator data of the data node to the API service.
[0109] Step 302: The API service sends the node state indicator data to the control manager.
[0110] Step 303: The control manager determines whether the node state indicator data meets the preset abnormal indicator condition to obtain a preliminary monitoring result of the node state of the data node.
[0111] Step 304: In the case that the preliminary monitoring result indicates that the state is suspected to be abnormal, the control manager sends a port monitoring signal of the target port to the data node.
[0112] Step 305: The data node feeds back port monitoring state indicator data to the control manager in response to the port monitoring signal.
[0113] Step 306: The control manager determines whether the port monitoring state indicator data meets the preset abnormal indicator condition to obtain a port state monitoring result.
[0114] Step 307: In the case that the port state monitoring result indicates that the port state of the target port is abnormal, it is determined that the state of the data node is abnormal.
[0115] It should be noted that after the API service stores the preliminary monitoring result of the node state, the network agent in the data node synchronously updates the preliminary monitoring result of the node state.
[0116] In some embodiments of this application, outside of data nodes and management nodes, if a user on a client side wants to access certain data in a data node, the user can obtain the data through the client side. Specifically, the client side can connect to the service port in the data node and obtain the required data through the service port. If the data node is in an abnormal state, the service port is closed, and the user cannot obtain the required data. If the data node is in a normal state, the service port is open, and the user can obtain the required data. The client side can display alarm information for service abnormality. Specifically, the reason for the service abnormality and the alarm information for service abnormality can be fed back to the client side through a network proxy.
[0117] refer to Figure 4 , Figure 4 A schematic diagram illustrating a data node status monitoring method for interaction between data nodes, management nodes, and user terminals, as shown below. Figure 4 As shown, the actuator 121 in data node 120 actively reports the node status indicator data of the data node to the API service 112 in management node 110 through the local port. Then, the API service 112 feeds back the node status indicator data to the control manager 114. The control manager 114 determines the preliminary monitoring result of the node status of the data node based on the node status indicator data. If the preliminary monitoring result of the node status indicates that the data node status is abnormal, the control manager actively monitors the port monitoring status indicator data of the target port and determines whether the data node is really abnormal based on the port monitoring status indicator data.
[0118] Data node 120 can obtain the preliminary monitoring results of its stored node status from API service 112, and then control the opening and closing of the local port of data node 120 based on the preliminary monitoring results of the node status.
[0119] Users can access data in data node 120 through user terminal 130. Specifically, they can access data stored in network proxy 122. Network proxy 122 can send access data corresponding to the access request to user terminal 130 based on the access request sent by user terminal 130.
[0120] When the actuator 121 reports the node status index data of the data node to the API service 112, it can also report the service status of the local port to the API service 112. The control manager 114 can modify the service status of the local port stored in the API service 112.
[0121] Figure 5 A schematic diagram of the hardware structure for status monitoring of data nodes provided in an embodiment of this application is shown.
[0122] The state monitoring device of the data node can include a processor 501 and a memory 502 storing computer program instructions.
[0123] In particular, the processor 501 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application.
[0124] The memory 502 can include a mass storage for data or instructions. By way of example and not limitation, the memory 502 can include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. The memory 502 can include removable or non-removable (or fixed) media, where appropriate. The memory 502 can be internal or external to the integrated gateway disaster recovery device, as appropriate. In particular embodiments, the memory 502 is non-volatile, solid-state memory.
[0125] The memory can include read-only memory (ROM), random-access memory (RAM), magnetic disk storage mediums, optical storage mediums, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software that, when executed (by one or more processors), is operable to access the data nodes' state monitoring method described with reference to the method according to an aspect of the present disclosure.
[0126] The processor 501 implements any one of the above-mentioned data nodes' state monitoring methods by reading and executing the computer program instructions stored in the memory 502.
[0127] In one example, the data nodes' state monitoring device can further include a communication interface 503 and a bus 510. As shown, the processor 501, the memory 502, and the communication interface 503 are connected through the bus 510 and complete communication with each other. Figure 5
[0128] The communication interface 503 is mainly used to realize the communication between the modules, devices, units and / or equipment in the embodiments of the present application.
[0129] Bus 510 includes hardware, software, or both, to couple components of the online data traffic metering device to each other and to couple components to other components within the online data traffic metering device. While bus 510 is shown for the sake of clarity as a single bus, it can comprise one or more buses operating together. Bus 510 can be implemented using any suitable type of bus or buses, including, but not limited to, an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infmiband interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or any other suitable bus or interconnect, or a combination of two or more of these. Where appropriate, bus 510 can be implemented as a system-wide interconnect or a combination of busses.
[0130] The state monitoring device of the data node can perform the state monitoring method of the data node in the embodiments of the present application based on adding an active monitoring step.
[0131] In addition, in combination with the method for monitoring the state of the data node in the above embodiments, the embodiments of the present application can provide a computer storage medium for implementation. The computer storage medium has computer program instructions stored thereon; the computer program instructions are executed by a processor to implement any one of the methods for monitoring the state of the data node in the above embodiments.
[0132] The embodiments of the present application also provide a computer program product comprising a computer program, which, when executed by a processor, implements any one of the methods for monitoring the state of the data node in the above embodiments.
[0133] It needs to be made clear that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of well-known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method processes of the present application are not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the present application.
[0134] The functional blocks shown in the above described block diagrams can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium, or transmitted through a carrier signal in a data signal on a transmission medium or communication link. A "machine-readable medium" includes any medium that can store or transport information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROMs, flash memories, erasable ROMs (EROMs), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and the like. The code segments can be downloaded via computer networks such as the Internet, Intranet, and the like.
[0135] It is also important to note that the examples mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the steps mentioned in the examples, that is, the steps can be performed in the order mentioned in the examples, or in an order different from the examples, or several steps can be performed simultaneously.
[0136] The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer program instructions can also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other processing devices to operate in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other processing devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other processing devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer program instructions can also be stored in a computer readable medium that can direct a computer, other programmable data processing apparatus, or other processing devices to operate in a particular manner, such that the instructions stored in the computer readable medium produce an article of manufacture including instructions which implement the function / act specified in the flowchart and / or block diagram block or blocks.
[0137] The above merely illustrates the specific implementation of the present application. Those skilled in the art can clearly understand the specific working process of the system, module and unit described above for the convenience and brevity of description, and the corresponding process in the foregoing method embodiments can be referred to, which will not be described herein again. It should be understood that the protection scope of the present application is not limited to this. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed in the present application, and these modifications or replacements should be covered in the protection scope of the present application.
Claims
1. A method for monitoring the status of a data node, characterized in that, Applied to a management node, which includes an application programming interface (API) service and a control manager, the method includes: The control manager obtains node status indicator data of the data nodes sent by the API service, wherein the status indicator data is reported by the data nodes to the API service; The control manager determines whether the node status indicator data meets the preset abnormal indicator conditions, and obtains the preliminary monitoring results of the node status of the data node. If the preliminary monitoring results of the node status indicate that the data node's status is abnormal, the control manager monitors whether the status of the target port of the data node is abnormal, and obtains the port status monitoring results of the target port. If the control manager determines that the port status of the target port is abnormal based on the port status monitoring results, the data node's status is also abnormal.
2. The method according to claim 1, characterized in that, After obtaining the port status monitoring results of the target port, the method further includes: If the control manager determines that the port status monitoring results of the target port indicate that the port status of the target port is normal, then the status of the data node is determined to be normal.
3. The method according to claim 1, characterized in that, The step of monitoring the status of the target port of the data node through the control manager to obtain the port status monitoring result of the target port includes: The data node is monitored multiple times by the control manager to obtain the port status monitoring results of the target port for each monitoring. The control manager determines the ratio of monitoring results indicating that the target port's port status is normal to those indicating that the target port's port status is abnormal in multiple port status monitoring results. The control manager uses the monitoring results that indicate the target port's normal port status, or the monitoring results that indicate the target port's abnormal port status, which constitute the larger proportion, as the target port's port status monitoring result.
4. The method according to claim 1, characterized in that, The node status indicator data refers to the node status indicator data reported by the data node within at least two reporting cycles within a preset time period. The step of determining whether the node status indicator data meets preset abnormal indicator conditions through the control manager to obtain the preliminary monitoring results of the node status of the data node includes: The control manager determines whether the node status indicator data obtained in at least two reporting cycles within the preset time period meets the preset abnormal indicator conditions. If the control manager determines that the node status indicator data obtained in at least two reporting cycles within the preset time period all meet the preset abnormal indicator conditions, the preliminary monitoring result of the node status is abnormal.
5. The method according to claim 1, characterized in that, If the preliminary monitoring results of the node status indicate that the data node's status is abnormal, the method further includes: The control manager sends the preliminary monitoring results of the node status of the data node to the API service for storage.
6. The method according to claim 5, characterized in that, After the control manager sends the preliminary monitoring results of the data node's status to the application programming interface service in the management node for storage, the method further includes: The API service receives information retrieval requests sent by the data nodes. According to the information acquisition request, the API service sends the preliminary monitoring results of the node status stored in the API service to the data node, so that the data node determines the control signal to control the target port based on the preliminary monitoring results of the node status, and controls the opening and closing of the target port according to the control signal.
7. The method according to claim 1, characterized in that, The step of monitoring the status of the target port of the data node through the control manager to obtain the port status monitoring result of the target port specifically includes: The control manager sends the port monitoring signal of the target port to the data node; The control manager receives port monitoring status indicator data from the data node in response to the port monitoring signal. The control manager determines whether the port monitoring status indicator data meets the preset abnormal indicator conditions, and obtains the port status monitoring result of the target port.
8. The method according to claim 7, characterized in that, The node status indicator data reported by the data node and received by the API service are sent periodically by the data node. The step of sending the port monitoring signal of the target port to the data node through the control manager includes: The control manager periodically and asynchronously sends port monitoring signals for the target port to the data node; The process of receiving port monitoring status indicator data from the data node in response to the port monitoring signal via the control manager includes: The control manager periodically and asynchronously receives port monitoring status indicator data from the data nodes in response to the port monitoring signals. The step of monitoring the status of the target port of the data node for abnormalities through the control manager includes: The control manager periodically and asynchronously monitors whether the status of the target port of the data node is abnormal.
9. The method according to claim 1, characterized in that, After determining that the data node is in an abnormal state, the method further includes: The control manager stops receiving node status indicator data within a preset time period.
10. A data node status monitoring system, characterized in that, The system includes: Data nodes are used to report node status indicator data of the data nodes to the management node; The management node includes an application programming interface (API) service and a control manager. The API service receives node status indicator data reported by the data nodes and sends the node status indicator data to the control manager. The control manager receives the node status indicator data sent by the API service and determines whether the node status indicator data meets preset abnormal indicator conditions to obtain a preliminary monitoring result of the node status of the data nodes. If the preliminary monitoring result indicates that the data node's status is abnormal, the control manager monitors whether the status of the target port of the data node is abnormal to obtain a port status monitoring result of the target port. If the port status monitoring result indicates that the target port's port status is abnormal, the control manager determines that the data node's status is abnormal.
11. A data node status monitoring device, characterized in that, The device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the data node status monitoring method as described in any one of claims 1-9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the data node status monitoring method as described in any one of claims 1-9.
13. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the data node status monitoring method as described in any one of claims 1-9.
Citation Information
Patent Citations
Service publishing method and device, electronic equipment and storage medium
CN114461229A
Resource monitoring and dynamic scheduling method and device based on k8s cluster and medium
CN117785382A