Monitoring system, agent instance deployment method and device and electronic equipment
By adopting the master-slave agent instance combination and coordination services in the monitoring system, data monitoring and business interruption problems caused by proxy instance failure are solved, and a monitoring system with high availability and data integrity is realized.
Patent Information
- Application Number
- CN202510196333.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, failure of proxy instances on the server may lead to the failure of data monitoring, data acquisition and security guard functions, which in turn causes security risks, data loss and discontinuous business operations.
A monitoring system is designed, adopting the method of combining master-slave agent instances and coordination services. The master agent instance obtains monitoring data through remote connection and uploads it, receives data synchronized by the master agent instance from the proxy instance and continues to transmit data when the master agent instance fails. The coordination service elects the slave agent instance as a new master agent instance when the master agent instance fails.
The high availability of proxy instance functions is realized. Even if the main proxy instance fails, the system can continue to monitor, collect data and execute plug-ins to ensure the correctness and integrity of the data, and avoid security risks and business interruptions caused by proxy instance failure.
Smart Images

Figure CN120075233A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technologies, and in particular, to a monitoring system, a method and device for deploying agent instances, and an electronic device. Background Art
[0002] Enterprises have a large number of servers, which are divided into virtualized clusters, containerized clusters, storage clusters, CDN clusters, etc. according to different usage methods, services, and technical implementation methods of the servers, in order to more efficiently manage and utilize server resources. In all scenarios of resource management and utilization, without exception, each server needs to be configured with a comprehensive-function agent (e.g., security agent instance, data collection agent instance, monitoring agent instance, etc.) to play roles such as data monitoring, data collection, and security protection for the nodes.
[0003] If the security agent instance fails, it will not be able to detect potential intrusions in a timely manner, thus leaving security risks; once the data collection agent instance is abnormal, the data of the current node cannot be effectively collected, which will in turn cause a series of subsequent problems; and the malfunction of the monitoring agent instance will make the node status unknown, and it is difficult to immediately detect a machine failure, which is extremely likely to cause a chain reaction and adverse effects on business operations. Summary of the Invention
[0004] To solve the above technical problems or at least partially solve the above technical problems, the present application provides a monitoring system, a method and device for deploying agent instances, and an electronic device.
[0005] In a first aspect, the present application provides a monitoring system, including: a plurality of agent instance groups, each agent instance group including a main agent instance, a plurality of slave agent instances, and a coordination service, where the main agent instance and the slave agent instances in each agent instance group are respectively deployed on different host machines;
[0006] The main agent instance is configured to obtain monitoring data of a monitored node group through a remote connection and upload the monitoring data;
[0007] The slave agent instances are configured to receive the monitoring data synchronized by the main agent instance and continue to transmit the monitoring data being transmitted by the main agent instance when elected as a new main agent instance;
[0008] The coordination service is configured to elect any one of the slave agent instances as a new main agent instance when the main agent instance fails.
[0009] In a second aspect, the present application provides a method for deploying agent instances, which is applied to the coordination service as described in the first aspect, and the method includes:
[0010] Obtain the number of main agents of the main agent instances on multiple host machines;
[0011] Determine a target node group for the new agent instance combination to be deployed among multiple host machines according to the number of main agents;
[0012] If the number of host machines in the target node group is greater than or equal to a preset quantity threshold, obtain the total resource information and resource distribution information of each host machine in the target node group;
[0013] Deploy the new agent instance combination on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group.
[0014] Optionally, determining a target node group for the new agent instance combination to be deployed among multiple host machines according to the number of main agents includes:
[0015] Determine whether there are host machines with the number of main agents less than the main agent quantity threshold;
[0016] If there are host machines with the number of main agents less than the main agent quantity threshold, add the host machines with the number of main agents less than the average value to the target node group.
[0017] Optionally, deploying the new agent instance combination on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group includes:
[0018] For each host machine, determine the difference between the total resource information of each host machine and the resource distribution information as the remaining resource information of the host machine;
[0019] If there are multiple host machines in the target node group with the remaining resource information greater than or equal to the resource requirements of the new agent instance combination, deploy the new agent instance combination on multiple host machines in the target node group.
[0020] Optionally, deploying the new agent instance combination on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group further includes:
[0021] If there are multiple host machines in the target node group with the remaining resources less than the resource requirements of the new agent instance combination, increase the number of host machines.
[0022] Optionally, obtaining the resource distribution information of each host machine in the target node group includes:
[0023] For each host machine, obtain the first resource information allocated to all main agent instances on the host machine;
[0024] Obtain the second resource information allocated to all slave proxy instances on the host;
[0025] Obtain the third resource information reserved by the host for a slave proxy instance to switch to a master proxy instance, where the third resource information is determined according to the slave proxy instance that requires the most resources for the switch;
[0026] Obtain the fourth resource information reserved by the host for the system;
[0027] Determine the sum of the first resource information, the second resource information, the third resource information, and the fourth resource information as the resource distribution information of the host.
[0028] Optionally, obtain the first resource information allocated to all master proxy instances on the host, including:
[0029] Obtain the number of plugins consumed by the host to monitor a monitored node;
[0030] Obtain the number of metrics consumed by the host to monitor a monitored node;
[0031] Determine the plugin resource share according to the number of plugins and the resource share consumed by each plugin;
[0032] Determine the metric resource share according to the number of metrics and the resource share consumed by each metric;
[0033] Determine the sum of the plugin resource share and the metric resource share as the first resource information.
[0034] Optionally, obtain the total resource information of each host in the target node group, including:
[0035] For each host, obtain the CPU resources, storage resources, and bandwidth resources included in the host;
[0036] Determine the first resource share corresponding to the CPU resources according to the CPU resources and the CPU resource conversion coefficient;
[0037] Determine the second resource share corresponding to the storage resources according to the storage resources and the storage resource conversion coefficient;
[0038] Determine the third resource share corresponding to the bandwidth resources according to the bandwidth resources and the bandwidth resource conversion coefficient;
[0039] Determine the sum of the first resource share, the second resource share, and the third resource share as the total resource information.
[0040] In a third aspect, the present application provides a proxy instance deployment device, which is applied to the coordination service as described in the first aspect. The device includes:
[0041] A first acquisition module, configured to acquire the number of main proxies of the main proxy instances on multiple host machines;
[0042] A first determination module, configured to determine a target node group for the new proxy instance combination to be deployed among multiple host machines according to the number of main proxies;
[0043] A second acquisition module, configured to, if the number of host machines in the target node group is greater than or equal to a preset quantity threshold, acquire the total resource information and resource distribution information of each host machine in the target node group;
[0044] A deployment module, configured to deploy the new proxy instance combination on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group.
[0045] Optionally, the first determination module includes:
[0046] A first determination sub-module, configured to determine whether there is a host machine with the number of main proxies less than the main proxy quantity threshold;
[0047] An addition sub-module, configured to, if there is a host machine with the number of main proxies less than the main proxy quantity threshold, add the host machine with the number of main proxies less than the average value to the target node group.
[0048] Optionally, the deployment module includes:
[0049] A second determination sub-module, configured to, for each host machine, determine the difference between the total resource information and the resource distribution information of each host machine as the remaining resource information of the host machine;
[0050] A deployment sub-module, configured to, if there are multiple host machines in the target node group whose remaining resources are greater than or equal to the resource requirements of the new proxy instance combination, deploy the new proxy instance combination on multiple host machines in the target node group.
[0051] Optionally, the deployment sub-module is further configured to:
[0052] If there are multiple host machines in the target node group whose remaining resources are less than the resource requirements of the new proxy instance combination, increase the number of host machines.
[0053] Optionally, the second acquisition module includes:
[0054] A first acquisition sub-module, configured to, for each host machine, acquire the first resource information allocated to all main proxy instances on the host machine;
[0055] A second acquisition sub-module, configured to acquire second resource information allocated to all slave proxy instances on the host machine;
[0056] A third acquisition sub-module, configured to acquire third resource information reserved by the host machine for a slave proxy instance to be switched to a master proxy instance, where the third resource information is determined according to the slave proxy instance that requires the most resources for the switch;
[0057] A fourth acquisition sub-module, configured to acquire fourth resource information reserved by the host machine for the system;
[0058] A first summation sub-module, configured to determine the sum of the first resource information, the second resource information, the third resource information, and the fourth resource information as the resource distribution information of the host machine.
[0059] Optionally, the first acquisition sub-module includes:
[0060] A first acquisition unit, configured to acquire the number of plugins consumed by the host machine for monitoring a monitored node;
[0061] A second acquisition unit, configured to acquire the number of metrics consumed by the host machine for monitoring a monitored node;
[0062] A third acquisition unit, configured to determine the plugin resource share according to the number of plugins and the resource share consumed by each plugin;
[0063] A determination unit, configured to determine the metric resource share according to the number of metrics and the resource share consumed by each metric;
[0064] A summation unit, configured to determine the sum of the plugin resource share and the metric resource share as the first resource information.
[0065] Optionally, the second acquisition module includes:
[0066] A fifth acquisition sub-module, configured to, for each host machine, acquire the CPU resources, storage resources, and bandwidth resources included in the host machine;
[0067] A third determination sub-module, configured to determine a first resource share corresponding to the CPU resources according to the CPU resources and the CPU resource conversion coefficient;
[0068] A fourth determination sub-module, configured to determine a second resource share corresponding to the storage resources according to the storage resources and the storage resource conversion coefficient;
[0069] A fifth determination sub-module, configured to determine a third resource share corresponding to the bandwidth resources according to the bandwidth resources and the bandwidth resource conversion coefficient;
[0070] The first summing sub-module is configured to determine the sum of the first resource share, the second resource share, and the third resource share as the total resource information.
[0071] In a fourth aspect, the present application provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0072] The memory is used to store a computer program;
[0073] When the processor is configured to execute the program stored in the memory, it implements the proxy instance deployment method according to any one of the second aspects.
[0074] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art:
[0075] In the embodiments of the present application, by decoupling each proxy instance in the proxy instance combination from the monitored node and not setting them in the same physical space, the main proxy instance and the slave proxy instance are deployed separately. While the proxy instance obtains node monitoring data by remotely connecting to observe the monitored node, it can avoid direct interaction or dependence with the hardware and software environment of the node. Even if the node is affected by hardware or software, the monitoring ability is not damaged; the proxy instance combination adopts a multi-copy method (the main proxy instance with multiple slave proxy instances, selecting the master in real time) to achieve high availability of the proxy instance function. If a proxy instance fails, other proxy instances can still continue to monitor, collect data, and execute plugins; for the case of missing data, the replica can continue to transmit the data, thus ensuring the correctness and integrity of the data. Description of the Drawings
[0076] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0077] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0078] Figure 1 It is an overall structure diagram of a monitoring system provided by an embodiment of the present application;
[0079] Figure 2 It is a flowchart of a proxy instance deployment method provided by an embodiment of the present application;
[0080] Figure 3 Schematic diagram of the deployment status of a newly added proxy instance combination (Group1) provided by an embodiment of the present application;
[0081] Figure 4 Schematic diagram of the deployment status of a newly added proxy instance combination (Group2) provided by an embodiment of the present application;
[0082] Figure 5 Schematic diagram of the deployment status of a newly added proxy instance combination (Group3) provided by an embodiment of the present application;
[0083] Figure 6 Structural diagram of a proxy instance deployment device provided by an embodiment of the present application;
[0084] Figure 7 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0085] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.
[0086] In the prior art, if a security proxy instance fails, it will not be able to detect potential intrusions in a timely manner, thus leaving a security risk; once a data collection proxy instance is abnormal, the data of the current node cannot be effectively collected, thereby triggering a series of subsequent problems; and the malfunction of the monitoring proxy instance will make the node status unknown, and it is difficult to immediately detect a machine failure, which is extremely likely to cause a chain reaction and adverse effects on business operations. For this reason, the embodiments of the present application provide a monitoring system, a proxy instance deployment method, device, and electronic device.
[0087] The embodiments of the present application provide a monitoring system, as Figure 1 shown, including: a plurality of proxy instance combinations, each of the proxy instance combinations includes a main proxy instance (agent leader), a plurality of slave proxy instances (agent slave), and a coordination service (such as: Zookeeper or Etcd), and each proxy instance combination can monitor a plurality of monitored nodes (nodes) in a monitored node group (cluster), that is: The main proxy instance and the slave proxy instances in each proxy instance combination are respectively deployed on different host machines (hosts);
[0088] When starting up, each agent instance in an agent instance combination will first register with a coordination service (such as Zookeeper or Etcd). This can be achieved by creating a node in a specific directory of the coordination service. For example, in Etcd, key-value pairs are used to store path registration information:
[0089] Basic information of the agent instance combination: / agent / group / {agent_group_id};
[0090] Metrics associated with the agent instance combination: / agent / group / {agent_group_id} / metric / , information about each metric;
[0091] Plugins associated with the agent instance combination: / agent / group / {agent_group_id} / plugin / ;
[0092] Monitored nodes (list of observed nodes) associated with the agent instance combination: / agent / group / {agent_group_id} / node / ;
[0093] Basic information of each agent instance in the agent instance combination:
[0094] / agent / group / {agent_group_id} / agent / {agent_id}, information about each plugin;
[0095] Deployment information of each agent instance in the agent instance combination:
[0096] / agent / group / {agent_group_id} / agent / {agent_id} / available_resource, proportion of resources already allocated on the node where it is located;
[0097] Among them, {agent_group_id} and {agent_id} represent the agent instance combination ID and the agent instance ID respectively.
[0098] The main agent instance is used to obtain monitoring data of the monitored node combination through remote connection and upload the monitoring data;
[0099] In the embodiments of the present application, the main agent instance can remotely access nodes through a secure network connection and use technologies such as SSH, API, or other remote management technologies to obtain necessary monitoring data. In addition, a proxy mechanism can be considered to run a lightweight proxy program on the local node for data preprocessing and improve the efficiency of data transmission. The main agent instance can be pre-set with mechanisms such as periodic retries, data caching, and fallback to cope with network unreliability and ensure the continuity of monitoring.
[0100] The monitored nodes do not need to install agent instances or plugins. The relationship between the agent instance and the monitored nodes is completed by the monitoring system. The monitored nodes only need to reserve the ssh port in the conventional manner (usually 22, which can be used according to the ssh port opened by the system); the public key corresponding to the agent instance can be added to the monitored nodes to ensure that the agent instance has the permission to obtain monitoring data from the monitored nodes, etc.
[0101] The main agent instance obtains the monitoring data of the combined monitored nodes through a remote connection. This operation can be defined as a task, and timeout parameters (such as async and poll) can be set for the task to ensure that the task is completed within the expected time to avoid excessive execution time.
[0102] When the task execution times out or is completed, the task process will be recycled, and related processes and side effects will be automatically cleaned up to avoid leaving hanging processes or other unnecessary resource occupations on the remote nodes.
[0103] After a certain task is initiated, it is executed in parallel for each monitored node in the monitored node group. At the same time, because the functions, timeouts, etc. of the task all belong to the same task, and the monitored nodes in the monitored node group are also in a homogeneous environment, it is easier to identify abnormal nodes for the returned data, which is beneficial to diagnosing abnormal situations through the differences in indicators.
[0104] The slave agent instance is used to receive the monitoring data synchronized by the main agent instance and continue to transmit the monitoring data being transmitted by the main agent instance when it is elected as the new main agent instance;
[0105] The coordination service is used to elect any slave agent instance as the new main agent instance when the main agent instance fails.
[0106] At the first registration, the agent with min_agent_id in the agent instance combination min_agen_id can be designated as the main agent instance (leader agent) and recorded in the / agent / group / {agent_group_id} / leader directory in the basic information of the agent instance combination; at this time The slave agent in the agent instance combination is in the standby state, listening for changes under / agent / {agent_group_id} / leader. If the existing master agent instance fails or the connection is interrupted, the coordination service will trigger a new round of leader election. A new master agent instance will be selected from the slave agent instances (based on the resource allocation of the nodes where the agents are located, with priorities. The agent with less allocated resources on the node is preferred), so as to ensure the normal function of this group of agents.
[0107] In the embodiment of the present application, when the master agent instance fails, the leader election algorithm can be used to elect any slave agent instance as the new master agent instance, which can ensure that there is always a master agent instance responsible for monitoring tasks at any time, and at the same time ensure that the election process is carried out quickly when one of the master agent instances fails.
[0108] In the embodiment of the present application, by decoupling each agent instance in the agent instance combination from the monitored node and not setting them in the same physical space, the master agent instance and the slave agent instances are deployed separately. While the agent instances obtain node monitoring data by remotely connecting to observe the monitored node, it is possible to avoid direct interaction or dependence with the hardware and software environment of the node. Even if the node is affected by hardware or software, the monitoring ability is not damaged; the agent instance combination adopts a multi-copy method (the master agent instance has multiple slave agent instances, and the master is selected in real time) to achieve high availability in the function of the agent instance. If one agent instance fails, other agent instances can still continue to monitor, collect data, and execute plugins; for the case of missing data, the replica can continue to transmit the data, thus ensuring the correctness and integrity of the data.
[0109] The embodiment of the present application provides a method for deploying agent instances, which is applied to the coordination service as described in the foregoing embodiments, as Figure 2 shown, the method includes:
[0110] Step S101, obtain the number of master agents of the master agent instances on multiple said host machines;
[0111] Based on the foregoing embodiments, it can be seen that in the monitoring system, the master agent instance and the slave agent instance in any agent instance combination are deployed on different host machines. The master agent instance undertakes the observation and monitoring tasks of the corresponding node group and bears the load, while the slave agent instance is in a standby and low-load state. In order to make the resource distribution on multiple host machines more balanced, the master agent instances in different agent instance combinations should be deployed as dispersedly as possible to evenly distribute the load of the monitoring tasks, and at the same time, the nodes should reserve redundant resources; when deploying, the sequential circular polling method can be used. In order to understand the deployment situation of the master agent instances, in this step, the number of master agent instances deployed on multiple host machines, that is, the number of master agents, can be obtained.
[0112] Step S102: Determine a target node group for the new agent instance combination to be deployed among the multiple host machines according to the number of master agents.
[0113] In this step, the host machines with the number of master agents less than a preset threshold can be used as target nodes, and finally the target node group is obtained.
[0114] In an implementation manner of the present application, determining a target node group for the new agent instance combination to be deployed among the multiple host machines according to the number of master agents includes: determining whether there are host machines with the number of master agents less than the master agent number threshold; if there are host machines with the number of master agents less than the master agent number threshold, adding the host machines with the number of master agents less than the average value to the target node group.
[0115] That is, the target nodes can be selected with reference to the conditions of the following formula to construct the target node group:
[0116]
[0117] node total_agent_leader The number of master agents of the master agent instance on any host machine, cluster total_agent_leader represents the total number of master agent instances in all agent instance combinations. a is an adjustment coefficient used to adjust the deployment density. Generally, a is 1. If it is necessary to increase the scheduling density, for example, a can be increased to 1.5.
[0118] For example: Assume a = 1, and there is no master agent instance deployed on a certain host machine, then node total_agent_leader is 0. Assume that there are 5 host machines in the cluster and 2 master agent instances have been deployed on two of them. In this case, the remaining 3 host machines meet the conditions, so the remaining 3 host machines can be used as target nodes;
[0119] For another example: Assume a = 1, there are 5 host machines in the cluster, and a total of 6 main agent instances have been deployed. If 1 main agent instance is deployed on some host machines and 2 main agent instances are deployed on some other host machines, then for any host machine, there exists or At this time, select the host machine with only 1 main agent instance deployed as the target node, and do not select the host machine with 2 main agent instances deployed as the target node.
[0120] Step S103: If the number of host machines in the target node group is greater than or equal to the preset quantity threshold, obtain the total resource information and resource distribution information of each host machine in the target node group;
[0121] If the number of host machines in the target node group is greater than or equal to the preset quantity threshold, it indicates that there are enough host machines to deploy the newly added proxy instance combination, and the total resource information and resource distribution information of each host machine in the target node group can be obtained.
[0122] Among them, the total resource information includes: CPU resources, storage (memory) resources, network card bandwidth resources, and the resource distribution information includes: resources allocated to the main agent instance, resources allocated to the slave agent instance, resources reserved for the slave agent to switch to the main agent, and system reserved resources, etc.
[0123] If the number of host machines in the target node group is less than the preset quantity threshold, it is prompted that the newly added proxy instance combination cannot be deployed and expansion is required.
[0124] Step S104: Deploy the newly added proxy instance combination on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group.
[0125] In this step, the remaining resource information of each host machine in the target node group can be determined according to the total resource information and the resource distribution information, and the main agent instance and the slave agent instance in the newly added proxy instance combination are successively deployed on different host machines in the order of the remaining resource information from more to less.
[0126] After the newly added proxy instance group is deployed on the host machines in the target node group, the host machines can be associated with the proxy instance group. The host machines need to be entered, associated, and delivered to the monitoring system, and the monitoring system will synchronize a node list to the node directory of a proxy instance group in the coordination service ( / agent / group / {agent_group_id} / node / {$node}); similar to the grouping logic used in traditional monitoring && alerting logic, generally a batch of machines with the same business and purpose are delivered as the same proxy instance group; the newly added proxy instance group will be associated with metrics and plugins, such as the number of the aforementioned metrics, the number and types of plugins, which will be reflected at the resource level; at the same time, the target node group associated with the proxy instance group supports expansion.
[0127] In the embodiment of the present application, by determining the target node group of the newly added proxy instance group to be deployed among multiple host machines according to the number of main proxies, when the number of host machines in the target node group is greater than or equal to the preset quantity threshold, the newly added proxy instance group can be deployed on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group, realizing the decentralized deployment of the main proxy instances in multiple proxy instance groups and the decentralized deployment of the main proxy instances and slave proxy instances in the newly added proxy instance group, realizing the equal distribution of the monitoring tasks for monitoring the monitored nodes among multiple host machines, and at the same time ensuring the resource redundancy reserved for the host machines.
[0128] In another embodiment of the present application, deploying the newly added proxy instance group on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group includes:
[0129] Step S201, for each host machine, determine the difference between the total resource information of each host machine and the resource distribution information as the remaining resource information of the host machine;
[0130] On each host machine in the target node group, there may be multiple main proxy instances of different proxy instance groups at the same time. Therefore, it is necessary to calculate the remaining resource information of each host machine in the target node group.
[0131] Step S202, if there are multiple host machines in the target node group whose remaining resource information is greater than or equal to the resource requirements of the newly added proxy instance group, deploy the newly added proxy instance group on multiple host machines in the target node group.
[0132] In this step, if the remaining resource information of multiple host machines in the target node group is greater than or equal to the resource requirements of the newly added agent instance combination, it indicates that the remaining resource information of multiple host machines in the target node group can bear the resource requirements of the newly added agent instance combination, and the newly added agent instance combination can be deployed on multiple host machines in the target node group.
[0133] The resource requirements of the newly added agent instance combination are also the resource consumption of the newly added agent instance combination. The resource consumption of the newly added agent instance combination depends on the scale (node n ) of the associated monitored node group, the metrics (metric i ), and the number of plugins (plugin j ).
[0134] group id The resources required for the main agent instance of the newly added agent instance combination can be calculated according to the following formula:
[0135]
[0136] The minimum resource consumption of a plugin can be expressed as 0.1 rs (the main consumption type is CPU). The resource consumption will vary depending on the different functions of the plugin. Indicating the resource consumption of j plugins required to observe one machine can be expressed as:
[0137]
[0138] Where represents the resource share of plugin j;
[0139] The resource consumption for collecting one metric can be expressed as 0.05 rs (the main consumption type is memory). Indicating the resource consumption of i metrics required to observe one machine can be expressed as:
[0140]
[0141] Then, for a monitored node group with n monitored nodes, the resource overhead of its associated main agent instance is:
[0142]
[0143] Where j is the plugin id and i represents the number of metrics.
[0144] For example, for a group of 50 observed nodes, each node needs to collect 200 metrics, and 30 plugins need to be executed (assuming the same resources), then the resource requirements of the main agent instance can be obtained as follows:
[0145]
[0146] Step S203, if the remaining resources of multiple host machines in the target node group are less than the resource requirements of the newly added agent instance combination, increase the number of host machines.
[0147] In the embodiment of the present application, the remaining resources of multiple host machines in the target node group can be sorted from more to less. If the most idle host machine does not meet the resource requirements of the main agent instance in the newly added agent instance combination, it can be determined that the remaining resource information of multiple host machines in the target node group is less than the resource requirements of the newly added agent instance combination.
[0148] Alternatively, the resource shares required by the newly added agent instance combination can be respectively accumulated to the host machines where the main agent instance and the slave agent instance are located, and the node resource distribution of their respective nodes can be judged whether it can meet If it meets the requirements, directly expand the capacity, that is, increase the number of host machines.
[0149] In this step, in the case where the remaining resource information of multiple host machines in the target node group is less than the resource requirements of the newly added agent instance combination, it indicates that the remaining resource information of multiple host machines in the target node group cannot bear the resource requirements of the newly added agent instance combination. At this time, it is necessary to expand the number of host machines to share the pressure on the existing host machines in the target node group.
[0150] After increasing the number of host machines, the target node group can be re-determined to re-deploy the newly added agent instance combination.
[0151] The embodiment of the present application can automatically deploy the main agent instance and the slave agent instance in the newly added agent instance combination according to the remaining resource conditions of the host machines in the target node group, realize the decentralized deployment of the main agent instance and the slave agent instance in the newly added agent instance combination, realize the even distribution of the monitoring tasks for monitoring the monitored nodes on multiple host machines, and at the same time ensure the redundancy of the reserved resources of the host machines.
[0152] In another embodiment of the present application, step S103 of obtaining the resource distribution information of each host machine in the target node group includes:
[0153] Step S301, for each host machine, obtain the first resource information allocated to all the main agent instances on the host machine;
[0154] In an implementation of the present application, obtaining the first resource information allocated to all the master agent instances on the host includes: obtaining the number of plugins consumed by the host for monitoring one monitored node; obtaining the number of metrics consumed by the host for monitoring one monitored node; determining the plugin resource share according to the number of plugins and the resource share consumed by each plugin; determining the metric resource share according to the number of metrics and the resource share consumed by each metric; and determining the sum of the plugin resource share and the metric resource share as the first resource information.
[0155] Step S302: Obtain the second resource information allocated to all the slave agent instances on the host;
[0156] Step S303: Obtain the third resource information reserved by the host for a slave agent instance to be switched to a master agent instance, where the third resource information is determined according to the slave agent instance that requires the most resources for switching;
[0157] Step S304: Obtain the fourth resource information reserved by the host for the system;
[0158] Step S305: Determine the sum of the first resource information, the second resource information, the third resource information, and the fourth resource information as the resource distribution information of the host.
[0159] That is to say, for the host t the node resource distribution (i.e., the resource distribution information) is as follows:
[0160]
[0161] Among them, represents the first resource allocated to all the master agent instances; represents the second resource allocated to all the slave agent instances;
[0162] represents the third resource reserved for a slave agent instance to be switched to a master agent instance; RS_system represents the fourth resource reserved for the system.
[0163] Furthermore, refer to the formula in the foregoing embodiment:
[0164]
[0165] Calculate the resource overhead of each master agent instance (i.e., the resources allocated to the master agent instance), and then sum up the resource overheads of multiple master agent instances.
[0166] For example: If there are 5 hosts with the foregoing configuration in the current target node group, and a total of 3 agent instance combinations are deployed =>
[0167] Cluster 1 (200 nodes, 30 plugins, 200 metrics), resource overhead of its main agent instance:
[0168]
[0169] Cluster 2 (100 nodes, 20 plugins, 100 metrics), resource overhead of its main agent instance:
[0170]
[0171] Cluster 3 (50 nodes, 30 plugins, 200 metrics), resource overhead of its main agent instance:
[0172]
[0173] For the resource overhead in the standby state, a fixed resource share can be taken.
[0174] RS_system is the system reservation, that is, 5% of the nodes resource , according to the above model, RS_system = 5% * 9600 = 480 rs;
[0175] Indicates the reserved resources for the identity conversion of the slave agent instance to the main agent instance, that is, the corresponding group needs to be reserved of the resource overhead:
[0176] host t If the number of agent_slaves scheduled on it is n, then the reserved share is
[0177] where A(n) represents the number of agents determined according to the number n; n represents the total number; The symbol represents the ceiling function, meaning taking the smallest integer greater than or equal to the given number. When (n) is 1, 2, or 3, When (n) is 4, 5, or 6, And so on.
[0178]
[0179] Represents taking the corresponding largest top n;
[0180] Furthermore, on the basis of the foregoing embodiments, there will be such as Figure 3 , Figure 4, and Figure 5 as shown in the deployment distribution, taking Figure 3 as an example, the node resource distribution RS of each host host is:
[0181]
[0182] In the embodiment of the present application, the resource distribution information of the host can be determined according to the sum of the first resource information allocated to all the main proxy instances on the host, the second resource information allocated to all the slave proxy instances, the third resource information reserved for switching the slave proxy instance to the main proxy instance, and the fourth resource information reserved for the system. The allocated resources on the host are quantified, which is convenient for subsequent calculation of the remaining resource information of the host. Furthermore, it is convenient to realize the decentralized deployment of the main proxy instance and the slave proxy instance in the newly added proxy instance combination, so as to evenly distribute the monitoring tasks for monitoring the monitored nodes on multiple hosts, and at the same time ensure the redundancy of the reserved resources of the host.
[0183] In another embodiment of the present application, step S103 of obtaining the total resource information of each host in the target node group includes:
[0184] Step S401, for each host, obtain the CPU resources, storage resources, and bandwidth resources included in the host;
[0185] Step S402, determine the first resource share corresponding to the CPU resources according to the CPU resources and the CPU resource conversion coefficient;
[0186] Step S403, determine the second resource share corresponding to the storage resources according to the storage resources and the storage resource conversion coefficient;
[0187] Step S404, determine the third resource share corresponding to the bandwidth resources according to the bandwidth resources and the bandwidth resource conversion coefficient;
[0188] Step S405, determine the sum of the first resource share, the second resource share, and the third resource share as the total resource information.
[0189] The main resources of a host are CPU, memory, and network card bandwidth. The total resource amount can be expressed as node resource = Min(cpu_limit core , mem_limit GB , Net_limit speed ). For the current scenario, it is abstracted as a resource share, and one share is one rs, which can be expressed as:
[0190]
[0191] where "milli" represents millicore, and 1 core = 1000 milli cores.
[0192] For a host with 96 cores, 512 GB of memory, and a 50 Gbps bandwidth, since the storage resources and network card bandwidth resources are much larger than the CPU resources, taking the minimum value among the three, the total resource amount of this host can be abstracted as 9600 rs, that is, node resource = 9600 rs.
[0193] The embodiments of the present application can quantify the CPU resources, storage resources, and bandwidth resources included in the host and convert them into the same dimension, so as to use the remaining resource information of the host for subsequent calculations. Furthermore, it is convenient to implement the decentralized deployment of the master proxy instance and the slave proxy instance in the newly added proxy instance combination, and to evenly distribute the monitoring tasks for monitoring the monitored nodes on multiple hosts, while ensuring the redundancy of the reserved resources of the host.
[0194] In another embodiment of the present application, a proxy instance deployment device is further provided, which is applied to the coordination service as described in the foregoing embodiment, as Figure 6 shown, the device includes:
[0195] A first acquisition module 11, configured to acquire the number of master proxies of the master proxy instances on multiple hosts;
[0196] A first determination module 12, configured to determine a target node group for deploying the newly added proxy instance combination among multiple hosts according to the number of master proxies;
[0197] A second acquisition module 13, configured to, if the number of hosts in the target node group is greater than or equal to a preset number threshold, acquire the total resource information and resource distribution information of each host in the target node group;
[0198] A deployment module 14, configured to deploy the newly added proxy instance combination on the hosts in the target node group according to the total resource information and the resource distribution information of each host in the target node group.
[0199] Optionally, the first determination module includes:
[0200] A first determination sub-module, configured to determine whether there is a host with the number of master proxies less than the master proxy number threshold;
[0201] An addition sub-module, configured to, if there is a host with the number of master proxies less than the master proxy number threshold, add the host with the number of master proxies less than the average value to the target node group.
[0202] Optionally, the deployment module includes:
[0203] A second determination sub-module, configured to, for each of the host machines, determine the difference between the total resource information of each host machine and the resource distribution information as the remaining resource information of the host machine;
[0204] A deployment sub-module, configured to, if there are multiple host machines in the target node group whose remaining resources are greater than or equal to the resource requirements of the newly added proxy instance combination, deploy the newly added proxy instance combination on the multiple host machines in the target node group.
[0205] Optionally, the deployment sub-module is further configured to:
[0206] If there are multiple host machines in the target node group whose remaining resources are less than the resource requirements of the newly added proxy instance combination, increase the number of host machines.
[0207] Optionally, the second acquisition module includes:
[0208] A first acquisition sub-module, configured to, for each host machine, acquire first resource information allocated to all the primary proxy instances on the host machine;
[0209] A second acquisition sub-module, configured to acquire second resource information allocated to all the secondary proxy instances on the host machine;
[0210] A third acquisition sub-module, configured to acquire third resource information reserved by the host machine for a secondary proxy instance to be switched to a primary proxy instance, where the third resource information is determined according to the secondary proxy instance that requires the most resources for the switch;
[0211] A fourth acquisition sub-module, configured to acquire fourth resource information reserved by the host machine for the system;
[0212] A first summation sub-module, configured to determine the sum of the first resource information, the second resource information, the third resource information, and the fourth resource information as the resource distribution information of the host machine.
[0213] Optionally, the first acquisition sub-module includes:
[0214] A first acquisition unit, configured to acquire the number of plugins required for the host machine to monitor a monitored node;
[0215] A second acquisition unit, configured to acquire the number of metrics required for the host machine to monitor a monitored node;
[0216] A third acquisition unit, configured to determine the plugin resource share according to the number of plugins and the resource share consumed by each plugin;
[0217] A determination unit, configured to determine the metric resource share according to the number of metrics and the resource share consumed by each metric;
[0218] A summation unit for determining the sum of the plugin resource share and the metric resource share as the first resource information.
[0219] Optionally, the second acquisition module includes:
[0220] A fifth acquisition sub-module for, for each of the host machines, acquiring the CPU resources, storage resources, and bandwidth resources included in the host machine;
[0221] A third determination sub-module for determining, according to the CPU resources and the CPU resource conversion coefficient, the first resource share corresponding to the CPU resources;
[0222] A fourth determination sub-module for determining, according to the storage resources and the storage resource conversion coefficient, the second resource share corresponding to the storage resources;
[0223] A fifth determination sub-module for determining, according to the bandwidth resources and the bandwidth resource conversion coefficient, the third resource share corresponding to the bandwidth resources;
[0224] A first summation sub-module for determining the sum of the first resource share, the second resource share, and the third resource share as the total resource information.
[0225] In another embodiment of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory, and a communication bus, where the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0226] The memory is used for storing a computer program;
[0227] The processor, when executing the program stored on the memory, implements the proxy instance deployment method described in any of the foregoing method embodiments.
[0228] In the electronic device provided in the embodiments of the present invention, the processor decouples each proxy instance in the proxy instance combination from the monitored node by executing the program stored on the memory, and does not set them in the same physical space. The main proxy instance and the slave proxy instances adopt a separated deployment method. While the proxy instance observes the monitored node through a remote connection to obtain node monitoring data, it can avoid direct interaction or dependence with the hardware and software environments of the node. Even if the node is affected by hardware or software, the monitoring ability is not damaged; the proxy instance combination adopts a multi-copy method (the main proxy instance with multiple slave proxy instances, selecting the master in real time) to achieve high availability in the function of the proxy instance. If one proxy instance fails, other proxy instances can still continue to monitor, collect data, and execute plugins; for the case of missing data, the replica can perform data continuation transmission, thereby ensuring the correctness and integrity of the data.
[0229] The communication bus 1140 mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The communication bus 1140 can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 7 it is only represented by a thick line in the figure, but it does not mean that there is only one bus or one type of bus.
[0230] The communication interface 1120 is used for communication between the above electronic device and other devices.
[0231] The memory 1130 may include a Random Access Memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0232] The above-mentioned processor 1110 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware instances.
[0233] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
[0234] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A monitoring system, characterized in that: include: A plurality of proxy instance combinations, wherein the proxy instance combination includes a master proxy instance, a plurality of slave proxy instances and a coordination service, wherein the master proxy instance and the slave proxy instance in each proxy instance combination are respectively deployed on different host machines; The master agent instance is used to obtain monitoring data of the monitored node combination through a remote connection and upload the monitoring data; The slave proxy instance is used to receive the monitoring data synchronized by the master proxy instance, and when elected as a new master proxy instance, continue to transmit the monitoring data being transmitted by the master proxy instance; The coordination service is used to select any slave proxy instance as a new master proxy instance when the master proxy instance fails.
2. A proxy instance deployment method, characterized in that: Applied to the coordination service as claimed in claim 1, the method comprises: Obtain the number of master agents of the master agent instances on the multiple host machines; Determining a target node group for deploying a newly added agent instance combination in the plurality of host machines according to the number of the master agents; If the number of host machines in the target node group is greater than or equal to a preset number threshold, obtaining total resource information and resource distribution information of each host machine in the target node group; The newly added proxy instance combination is deployed on the host machines in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group.
3. The proxy instance deployment method according to claim 2, characterized in that: Determining a target node group for deploying a newly added agent instance combination in the plurality of host machines according to the number of the master agents includes: Determine whether there is a host machine whose number of master agents is less than the master agent number threshold; If there are hosts whose number of master agents is less than the master agent number threshold, the hosts whose number of master agents is less than the average value are added to the target node group.
4. The agent instance deployment method according to claim 2, characterized in that: Deploying the newly added proxy instance combination on the host machine in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group includes: For each of the host machines, determining the difference between the total resource information of each of the host machines and the resource distribution information as the remaining resource information of the host machine; If there are multiple host machines in the target node group whose remaining resources are greater than or equal to the resource demand of the newly added proxy instance combination, the newly added proxy instance combination is deployed on the multiple host machines in the target node group.
5. The agent instance deployment method according to claim 4, characterized in that: The method further comprises: deploying a newly added proxy instance combination on a host machine in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group; The remaining resources of multiple hosts in the target node group are less than the resource requirements of the newly added proxy instance combination, so the number of hosts is increased.
6. The proxy instance deployment method according to claim 2, characterized in that: Obtaining resource distribution information of each host in the target node group includes: For each host machine, obtaining first resource information allocated to all master agent instances on the host machine; Obtaining all second resource information allocated from the proxy instance on the host machine; Acquire third resource information reserved by the host machine for switching from a slave proxy instance to a master proxy instance, wherein the third resource information is determined based on the slave proxy instance with the most resources required for switching; Acquire fourth resource information reserved by the host machine for the system; The sum of the first resource information, the second resource information, the third resource information and the fourth resource information is determined as the resource distribution information of the host machine.
7. The agent instance deployment method according to claim 6, characterized in that: Obtaining first resource information allocated to all master agent instances on the host machine, including: Obtaining the number of plug-ins required by the host machine to monitor a monitored node; Obtaining the number of indicators consumed by the host machine to monitor a monitored node; Determining a plug-in resource share according to the number of plug-ins and the resource share consumed by each plug-in; Determining the indicator resource share according to the number of indicators and the resource share consumed by each indicator; The sum of the plug-in resource share and the indicator resource share is determined as the first resource information.
8. The proxy instance deployment method according to claim 2, characterized in that: Obtaining total resource information of each host in the target node group includes: For each of the host machines, obtain the CPU resources, storage resources, and bandwidth resources included in the host machine; Determine a first resource share corresponding to the CPU resource according to the CPU resource and the CPU resource conversion coefficient; Determine a second resource share corresponding to the storage resource according to the storage resource and the storage resource conversion coefficient; Determine a third resource share corresponding to the bandwidth resource according to the bandwidth resource and the bandwidth resource conversion coefficient; The sum of the first resource share, the second resource share and the third resource share is determined as the total resource information.
9. A proxy instance deployment device, characterized in that: Applied to the coordination service as claimed in claim 1, the device comprises: A first acquisition module is used to acquire the number of master agents of the master agent instances on the multiple host machines; A first determination module, configured to determine a target node group for deploying a newly added agent instance combination in the plurality of host machines according to the number of the master agents; A second acquisition module is used to acquire total resource information and resource distribution information of each host machine in the target node group if the number of host machines in the target node group is greater than or equal to a preset number threshold; A deployment module is used to deploy the newly added proxy instance combination on the host machine in the target node group according to the total resource information and the resource distribution information of each host machine in the target node group.
10. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; The processor is used to implement the proxy instance deployment method described in any one of claims 2 to 8 when executing the program stored in the memory.