Large-scale scene monitoring method and device, equipment, medium and program product

By splitting the monitoring objects of the business cluster into multiple triples and creating multiple monitoring instances based on these triples, the problem that a single monitoring instance cannot handle monitoring metrics in large-scale scenarios is solved, and stable large-scale monitoring is achieved.

CN119938451APending Publication Date: 2025-05-06CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122413.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In large-scale scenarios, a single monitoring instance cannot effectively process the monitoring indicators of the business cluster, resulting in memory resource explosion and monitoring function failure.

Method used

By splitting the business cluster size and hierarchical differences and category differences of monitoring objects, triplets (monitoring targets, custom rules, remote writing information) are obtained, and multiple monitoring instances are created in the business cluster according to the triplets, and monitoring indicator data is collected, processed and stored separately.

Benefits of technology

Ensure that the metric processing volume of each monitoring instance is within the peak value that a single monitoring instance can bear, avoiding monitoring function failure, and achieving stable monitoring in large-scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938451A_ABST
    Figure CN119938451A_ABST
Patent Text Reader

Abstract

The invention provides a large-scale scene monitoring method and device, equipment, a medium and a program product. According to the business cluster scale, splitting based on the hierarchy difference and the category difference of the monitoring object to obtain a triple; wherein the triple comprises a monitoring target, a custom rule and remote write-in information; creating a monitoring instance in the service cluster according to the triad; and running the monitoring instance to obtain monitoring index data according to the monitoring target, performing aggregation calculation on the monitoring index data according to a user-defined rule, and storing the monitoring index data subjected to aggregation calculation into a time sequence database in the management and control cluster according to the remote write-in information, and obtaining the monitoring index data from the time sequence database and displaying the monitoring index data to a visual interface. According to the scheme, the monitoring object is split to obtain the triad, the monitoring instance is created according to the triad, it is guaranteed that index processing of each monitoring instance is within the peak value capable of being borne by a single monitoring instance, and monitoring function failure is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a monitoring method, device, equipment, medium and program product for large-scale scenarios. Background Art

[0002] Currently, cloud native technology has become the mainstream technology. A financial enterprise-level container platform has been launched on the cloud native platform to serve the monitoring of a large number of business systems.

[0003] Currently, in the monitoring service provided by the cloud native platform, a platform consisting of a management cluster and multiple business clusters is used for monitoring; a business cluster is used as a monitoring object, and only one monitoring instance is deployed to collect and store the monitoring data of the business cluster. However, when the cluster scale of the business cluster is large-scale, the indicator processing of a single monitoring instance will exceed the peak value that a single monitoring instance can carry, resulting in memory resource explosion and monitoring function failure.

[0004] Therefore, this method cannot achieve monitoring in large-scale scenarios. Summary of the invention

[0005] The embodiments of the present application provide a large-scale scene monitoring method, device, equipment, medium and program product to achieve the effect of monitoring large-scale scenes.

[0006] In the first aspect, an embodiment of the present application provides a monitoring method for a large-scale scenario, including: splitting the monitored objects based on the hierarchical differences and category differences according to the scale of the business cluster to obtain triplets; wherein the triplets include monitoring targets, custom rules, and remote write information; creating a monitoring instance in the business cluster according to the triplets; and running the monitoring instance to obtain monitoring indicator data according to the monitoring targets, aggregate the monitoring indicator data according to the custom rules, and store the aggregated monitoring indicator data in a time series database in the management and control cluster according to the remote write information; obtaining the monitoring indicator data from the time series database and displaying it on a visualization interface.

[0007] In one possible implementation, according to the cluster scale, the monitoring object is split based on the hierarchical differences and category differences to obtain a triple, specifically including: multiplying the business cluster scale coefficient and the level indicator coefficient of the monitoring object to obtain the indicator coefficient of the level indicator of the monitoring object under the cluster scale; calculating according to the indicator coefficient, constant coefficient, category indicator value, and grouping amplitude value obtained by the calculation to obtain the indicator processing value of the triple of different indicator levels obtained by splitting according to the grouping amplitude under each category, and selecting the maximum value of the indicator processing value; subtracting the maximum value of the indicator processing value from the indicator processing peak value; if the result of the subtraction is a negative value or a zero value, splitting according to the grouping amplitude value to obtain a triple; wherein the indicator processing peak value represents the indicator peak value that a single monitoring instance can carry.

[0008] In one possible implementation, if the result of the subtraction is a positive value, the group amplitude value is halved; based on the indicator coefficient, constant coefficient, category indicator value, and halved group amplitude value, the maximum value of the indicator processing value is calculated and selected again, and the result value of subtracting the maximum value of the indicator processing value from the indicator processing peak value is calculated until a triplet is obtained.

[0009] In one possible implementation, a monitoring instance is created in a business cluster according to a triplet, specifically including: creating a monitoring instance resource according to the triplet; the monitoring instance resource includes a monitoring object of the monitoring instance and a proxy service of the monitoring instance; wherein the proxy service of the monitoring instance provides a monitoring instance access port; if the monitoring instance resource is successfully created, then periodically checking whether the monitoring instance access port can be accessed normally; if the monitoring instance access port can be accessed normally, then creating a monitoring instance in the business cluster according to the successfully created monitoring instance resource.

[0010] In a possible implementation manner, if the monitoring instance resource fails to be created, the monitoring instance resource that failed to be created is deleted; and the monitoring instance resource is recreated.

[0011] In a possible implementation manner, before recreating the monitoring instance resource, the process further includes: determining whether the historical reconstruction times of the monitoring instance resource has reached a preset maximum reconstruction times; if the maximum reconstruction times has been reached, determining that the monitoring instance creation has failed.

[0012] In one possible implementation, it is determined whether the reconstruction time of the monitoring instance resource reaches a preset retry time threshold; wherein the reconstruction time is the time elapsed from the first time the monitoring instance resource is recreated to the current time; if the retry time threshold is reached, it is determined that the monitoring instance creation has failed.

[0013] In a possible implementation, after periodically checking whether the monitoring instance access port can be accessed normally, it also includes: if the monitoring instance access port cannot be accessed normally, determining whether the inspection time reaches a preset inspection time threshold; the inspection time is the time from the first time the monitoring instance access port is checked to the current time; if the inspection time threshold is not reached, checking the monitoring instance access port again; if the inspection time threshold is reached, determining that the monitoring instance creation has failed.

[0014] In a possible implementation, before creating a monitoring instance in a business cluster according to a triplet, it also includes: selecting a node in the business cluster to be set as a monitoring node, adding a node resource flag to the monitoring node, and marking the monitoring node as a node of a first attribute; wherein the node resource flag indicates the monitoring object of the monitoring node; adding a tolerance corresponding to the first attribute to the monitoring instance resource to support the creation of a monitoring instance resource in the monitoring node marked as the first attribute.

[0015] In a possible implementation, after the monitoring instance ends, the monitoring node, the node resource flag, the tag of the monitoring node, and the tolerance of the first attribute are cleared.

[0016] In a possible implementation, after creating a monitoring instance in a business cluster according to a triplet, it also includes: during the operation of the monitoring instance, checking whether the operation of the monitoring instance is normal; if the monitoring instance is running abnormally, repairing the monitoring instance according to the triplet, and displaying the abnormal operation status of the monitoring instance to a visualization interface; if the monitoring instance is running normally, continuing to run the monitoring instance, and displaying the successful operation status to the visualization interface.

[0017] In the second aspect, an embodiment of the present application provides a monitoring device for a large-scale scenario, including: a processing module, which is used to split the monitored objects according to the scale of the business cluster based on the hierarchical differences and category differences to obtain triplets; wherein the triplets include monitoring targets, custom rules, and remote write information; an operation module, which is used to create a monitoring instance in the business cluster according to the triplets; and, to run the monitoring instance to obtain monitoring indicator data according to the monitoring target, aggregate the monitoring indicator data according to the custom rules, and store the aggregated monitoring indicator data in a time series database in the management and control cluster according to the remote write information; a display module, which is used to obtain the monitoring indicator data from the time series database and display it on a visual interface.

[0018] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0019] The memory stores computer-executable instructions;

[0020] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementations of the first aspect.

[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.

[0022] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.

[0023] The embodiments of the present application provide a monitoring method, device, equipment, medium and program product for large-scale scenarios, which obtains triples by splitting the monitored objects according to the scale of the business cluster based on the hierarchical differences and category differences; wherein the triples include monitoring targets, custom rules, and remote write information; a monitoring instance is created in the business cluster according to the triples; and the monitoring instance is run to obtain monitoring indicator data according to the monitoring target, aggregate the monitoring indicator data according to the custom rules, and store the aggregated monitoring indicator data in a time series database in the management and control cluster according to the remote write information; the monitoring indicator data is obtained from the time series database and displayed on a visual interface. This solution uses the triples obtained after splitting the monitoring object and the monitoring instances created according to the triples to ensure that the indicator processing of each monitoring instance is within the peak value that a single monitoring instance can carry, thereby avoiding failure of the monitoring function. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0025] Figure 1 A schematic diagram of a large-scale monitoring method provided in this application Figure 1 ;

[0026] Figure 2 A schematic diagram of a large-scale monitoring method provided in this application Figure 2 ;

[0027] Figure 3 A schematic diagram of a large-scale monitoring method provided in this application Figure 3 ;

[0028] Figure 4 A schematic diagram of a process for creating a monitoring instance based on a state machine provided in this application;

[0029] Figure 5 A schematic diagram of the structure of a large-scale monitoring device provided by the present application;

[0030] Figure 6 A schematic diagram of the structure of the electronic device provided in this application.

[0031] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0032] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0033] Currently, cloud native technology has become the mainstream technology. A financial enterprise-level container platform has been launched on the cloud native platform to serve the monitoring of a large number of business systems.

[0034] Currently, in the monitoring service provided by the cloud native platform, a platform consisting of a management cluster and multiple business clusters is used for monitoring; a business cluster is used as a monitoring object, and only one monitoring instance is deployed to collect and store the monitoring data of the business cluster. However, when the cluster scale of the business cluster is large-scale, the indicator processing of a single monitoring instance will exceed the peak value that a single monitoring instance can carry, resulting in memory resource explosion and monitoring function failure.

[0035] Therefore, this method cannot achieve monitoring in large-scale scenarios.

[0036] The present application provides a monitoring method, device, equipment, medium and program product for a large-scale scenario provided by the embodiment of the present application, which obtains a triple by splitting the monitored object based on the hierarchical differences and category differences according to the scale of the business cluster; wherein the triple includes a monitoring target, a custom rule, and remote write information; a monitoring instance is created in the business cluster according to the triple; and the monitoring instance is run to obtain monitoring indicator data according to the monitoring target, aggregate the monitoring indicator data according to the custom rule, and store the aggregated monitoring indicator data in a time series database in the management and control cluster according to the remote write information; the monitoring indicator data is obtained from the time series database and displayed on a visual interface. This scheme uses the triple obtained after splitting the monitoring object and the monitoring instance created according to the triple to ensure that the indicator processing of each monitoring instance is within the peak value that a single monitoring instance can carry, thereby avoiding failure of the monitoring function.

[0037] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0038] Embodiment 1

[0039] Figure 1 A schematic diagram of a large-scale monitoring method provided in this application Figure 1 ,like Figure 1 As shown, the method includes:

[0040] S101, splitting the monitored objects based on the scale of the business cluster and the hierarchical and category differences to obtain triples; wherein the triples include the monitored targets, the custom rules, and the remote writing information.

[0041] S102, creating a monitoring instance in the business cluster according to the triplet; and running the monitoring instance to obtain monitoring indicator data according to the monitoring target, aggregate the monitoring indicator data according to the custom rules, and store the aggregated monitoring indicator data in the time series database in the management and control cluster according to the remote write information.

[0042] S103, obtaining the monitoring indicator data from the time series database and displaying it on a visualization interface.

[0043] Among them, time series databases such as Influx clusters; by setting up a time series database to store monitoring indicator data, the monitoring instance does not need to store monitoring indicator data, but only needs to collect and process data; when the monitoring indicator data needs to be queried, it is queried from the time series database, which avoids the situation in large-scale scenarios where the monitoring instance has to be responsible for both collecting and processing data and reading data, which leads to confusion in monitoring tasks and causes the loss of monitoring indicator data.

[0044] Specifically, the monitoring indicator data that has been aggregated and calculated will be stored in the time series database in the management and control cluster according to the remote write information; the indicator name will be filtered according to the remote write information; the indicators required in each split monitoring instance will be filtered by whitelist fuzzy matching, and uniformly forwarded and stored in the storage time series database; after filtering, it can avoid duplicate and redundant data from polluting the time series database.

[0045] The levels of the monitoring object include Container / Pod level, Node level, Namespace level and Cluster level. The indicators at different levels have different magnitudes. For example, the number of indicators at a single Node level will be dozens to hundreds. In a medium-sized business cluster, the number of Node level indicators is dozens to hundreds, so the number of indicators is thousands to tens of thousands. The number of Cluster level indicators is dozens to hundreds. The number of Cluster level indicators is one, so the number of indicators is dozens to hundreds. It can be seen that there are differences in the magnitude of indicators at different levels, that is, differences in levels. The categories of the monitoring object include CPU, memory, network and file system.

[0046] In actual applications, the magnitudes of different indicators of a business cluster can be divided into multiple magnitudes according to the Cluster level, Namespace level, Node level, Container / Pod level, and service level.

[0047] Specifically, according to the scale of the business cluster, the splitting is carried out based on the hierarchical differences and category differences of the monitored objects. For example, in a business cluster of 50 nodes, the splitting is carried out at the Node level: the number of cpu indicators is 8-12, with an average of 10 cpu indicators on each node, and a total of 500 cpu indicators, which is in the order of 10² - 10³; the number of memory indicators is 6-10, with an average of 8 memory indicators on each node, and a total of 400 indicators, which is in the order of 10² - 10³; the number of network indicators is 8-12, with 10 network indicators on each node, and a total of 500 network indicators for 50 nodes, which is in the order of 10² - 10³; the number of file system indicators is 7-10. If each node has 8 file system indicators, 50 nodes have 400 indicators, which is in the order of 10² - 10³; similarly, after splitting at other levels, triples are combined according to indicators and levels.

[0048] For example, the business cluster scale for scenarios with 1 to 1000 nodes can be divided into small scale, medium scale, and large scale; among them, the small scale node scale is 1 to 300 nodes, the medium scale node scale is 300 to 600 nodes, and the large scale node scale is 600 to 1000 nodes. Among them, for small-scale scenarios, a single monitoring instance (Prometheus instance) can achieve monitoring, but in medium-scale and large-scale scenarios, a single monitoring instance will cause memory resource explosion and frequent OOM (Out Of Memory).

[0049] In actual applications, for example, in a business cluster of 300 nodes, the monitoring objects are split based on the hierarchical and category differences to obtain triplets, and then a monitoring instance is created based on the triplets and the monitoring instance is run to obtain monitoring indicator data based on the monitoring target, and the monitoring indicator data is aggregated and calculated based on custom rules. The aggregated monitoring indicator data is stored in the time series database in the management and control cluster based on remote write information; the monitoring indicator data is obtained from the time series database and displayed on the visual interface. By using the triplets obtained after splitting the monitoring objects and creating monitoring instances based on the triplets, it is ensured that the indicator processing of each monitoring instance is within the peak value that a single monitoring instance can carry, thus avoiding the failure of the monitoring function.

[0050] Among them, splitting the monitoring object and monitoring it by multiple monitoring instances can ensure that the resource consumption of the monitoring instance is controllable and runs smoothly; the following Table 1 shows the average resource consumption of multiple monitoring instances.

[0051] Table 1

[0052]

[0053] Among them, prometheus-cpu-mem-0 indicates the resource consumption during CPU monitoring; prometheus-net-fs-sts-0 indicates the resource consumption during network (net), file system (fs) and status (sts) monitoring. The status includes the running status and process status of the monitoring instance; prometheus-core-0 indicates the resource consumption when aggregating the collected data.

[0054] In some examples, Figure 2 A schematic diagram of a large-scale monitoring method provided in this application Figure 2 ,like Figure 2 As shown in the figure, according to the cluster size, the hierarchical differences and category differences of the monitored objects are split to obtain triples, including:

[0055] S201, multiplying the business cluster scale coefficient and the level index coefficient of the monitored object to obtain the index coefficient of the level index of the monitored object under the cluster scale.

[0056] S202, calculate according to the calculated indicator coefficient, constant coefficient, category indicator value, and grouping amplitude value to obtain indicator processing values ​​of triplets of different indicator levels obtained by splitting according to the grouping amplitude under each category, and select the maximum value of the indicator processing value.

[0057] S203, subtract the maximum value of the indicator processing value from the indicator processing peak value; if the result of the subtraction is a negative value or a zero value, split it according to the group amplitude value to obtain a triple; wherein the indicator processing peak value represents the indicator peak value that a single monitoring instance can carry.

[0058] Among them, the indicator coefficient of the level indicator of the monitoring object under the cluster scale can define a vector set of indicator coefficients, as shown below:

[0059] in, Represents the Cluster level index coefficient, Indicates the Namespace level index coefficient, represents the node level index coefficient, Indicates the Container / Pod level indicator coefficient. represents the service level index coefficient, Represents the coefficient ratio of different scales.

[0060] In addition, for indicators of the same category, multi-level aggregation calculations are required at the Container / Pod level, Namespace level, Node level, and Cluster level. High-level indicators rely on low-level indicators for aggregation. The category vector set is defined as follows:

[0061]

[0062] in, Indicates the CPU category indicator value. Indicates the memory category indicator value, Indicates the network category indicator value, Indicates the file system category index value. Indicates the status category indicator value, Indicates other category indicator values.

[0063] After the monitoring object is split, several different triplets are generated (scrape jobs, recording rules, remote write) to ensure that each instance can collect necessary and non-repetitive indicators, aggregate and calculate basic indicators and aggregate indicators, and transfer the required indicators to the time series database as needed. The atomic triplets of the monitoring instance are defined as follows:

[0064]

[0065] in, Indicates the __name__ indicator name regular matching rule set of the scrape jobs to capture the monitoring target. Indicates the record rules custom aggregation indicator set. Indicates the regular matching rule set of the __name__ indicator name of the remote write forwarding address.

[0066] The monitoring object splitting processing function is as follows:

[0067]

[0068] in, It indicates the peak value of the indicator that a single monitoring instance can handle. For example, in a specific version and specific software and hardware environment, the monitoring instance can handle a peak value of 100,000 indicators, which is a fixed constant value. Indicates the grouping range after classification, used to indicate the number of indicators in the group after grouping, and its minimum value is 1; Indicates the result value.

[0069] Among them, according to the calculated indicator coefficient, constant coefficient, category indicator value, and grouping amplitude value, the indicator processing values ​​of the triplets of different indicator levels obtained by splitting according to the grouping amplitude in each category are calculated, and the maximum value of the indicator processing value is selected. For example, the triplets obtained after splitting the cpu according to different levels are: split into cpu1, cpu2 at the Node level, and split into cpu3, ​​cpu4 at the Cluster level; calculate the processing values ​​of the processing indicators of cpu1, cpu2, cpu3, ​​and cpu4 respectively; similarly, calculate the triplets obtained by splitting each category according to different levels; then select the maximum value among the processing values ​​of all processing indicators.

[0070] Subtract the selected maximum value from the peak value that a single monitoring instance can carry. If the result is a negative value or zero, it means that the indicator processing capacity of the monitoring instance created based on the split triples is within the peak upper limit, ensuring that the monitoring instance will not experience memory resource explosion and ensuring the stability of the monitoring function.

[0071] In some examples, if the result of the subtraction is a positive value, the group amplitude value is halved;

[0072] According to the indicator coefficient, constant coefficient, category indicator value, and halved group amplitude value, the maximum value of the indicator processing value is calculated again and selected, and the result value of subtracting the maximum value of the indicator processing value from the indicator processing peak value is calculated until a triplet is obtained.

[0073] Among them, if the result value is a positive value, it means that the indicator processing volume of the monitoring instance created based on the split triplet exceeds the peak upper limit, which will cause the monitoring instance of the monitoring instance to fail; split it according to the grouping amplitude value and split it into more triplets to ensure that the indicator processing volume of the monitoring instance created by each triplet does not exceed the upper limit.

[0074] Therefore, in large-scale scenarios, by splitting the monitoring objects, the indicator processing volume of each monitoring instance will not exceed the peak value, thus ensuring the stability of the monitoring instance. In addition, monitoring instances can be flexibly created for changes in the scale of the business cluster to ensure that the memory resources of the monitoring instance can support the monitoring of the current business cluster scale.

[0075] In some examples, a monitoring instance is created in a business cluster according to a triple, specifically including:

[0076] Create monitoring instance resources according to the triplet; the monitoring instance resources include the monitoring object of the monitoring instance and the proxy service of the monitoring instance; wherein the proxy service of the monitoring instance provides the monitoring instance access port;

[0077] If the monitoring instance resource is created successfully, a check is performed periodically to see whether the monitoring instance access port can be accessed normally. If the monitoring instance access port can be accessed normally, a monitoring instance is created in the business cluster based on the successfully created monitoring instance resource.

[0078] Among them, the proxy service of the monitoring instance is an intermediate layer service located between the monitoring instance and the monitoring object. In actual applications, the monitoring instance may be in a secure internal network environment, which cannot be directly accessed by the external network; the proxy service can be deployed at the network boundary, providing an externally accessible port to authenticate and authorize external requests, and only verified requests will be forwarded to the monitoring instance. In addition, by using the port of the proxy service, external users and services do not need to know the real IP address and port of the monitoring instance, thereby protecting the monitoring instance from direct attacks. The proxy service can use the internal IP address and port to communicate with the monitoring instance, isolating external requests from the internal monitoring instance, and enhancing the security of the monitoring instance.

[0079] In actual applications, the creation of a monitoring instance in a business cluster according to a triplet is based on a state machine; the state machine can display all state changes of the monitoring instance from creation to operation in real time to the user, so that the user can observe the monitoring process in real time.

[0080] When you create a monitoring instance resource for the first time, it will be displayed in the initialization state; after the monitoring resource is successfully created, it will be displayed in the checking state.

[0081] After the monitoring instance resource is successfully created, a check is made periodically to see whether the monitoring instance access port can be accessed normally. For example, a check is made every 10 seconds to see whether the monitoring instance port can be accessed normally. If it can be accessed normally, a monitoring instance is created based on the monitoring instance resource, and the monitoring instance is run and the running status is displayed.

[0082] In some examples, if the monitoring instance resource fails to be created, the monitoring instance resource whose creation failed is deleted; and the monitoring instance resource is recreated.

[0083] If the monitoring resource creation fails, it will be converted to the retry initialization state to remind the user that the monitoring instance resource creation failed.

[0084] In actual applications, the deletion of the monitoring instance resources will be divided into multiple batches. The operation of deleting the monitoring instance resources in the state machine is to delete them in batches every 60 seconds. Among them, it will be determined every 15 seconds whether the monitoring instance resources of the current batch are completely cleared. After the monitoring instance resources are completely deleted, the deletion stop status will be displayed.

[0085] When the creation of monitoring instance resources fails, some resources have been allocated or occupied, but the creation process has not been completed, and they are in an unstable or incomplete state. If these failed resources are not deleted, conflicts may occur with these residual resources when they are recreated. For example, some configuration files, database tables, or network ports may have been created, but the corresponding services failed to start normally. When creating again, the new creation process may fail because the file already exists or the port is already occupied.

[0086] Therefore, before each re-creation, the monitoring instance resources that failed to be created are deleted to prevent useless monitoring instance resources from interfering with the creation of subsequent monitoring instances.

[0087] In some examples, before recreating the monitoring instance resources, the following steps are also included:

[0088] It is determined whether the historical reconstruction times of the monitoring instance resource have reached the preset maximum reconstruction times; if the maximum reconstruction times have been reached, it is determined that the monitoring instance creation has failed.

[0089] The preset maximum number of reconstruction times is set according to demand; for example, the maximum number is set to 5 times. When the monitoring instance resource still fails to be created after the fifth reconstruction, it is determined that the monitoring instance creation has failed.

[0090] Creating monitoring instance resources may consume various system resources, such as CPU, memory, storage, network bandwidth, etc. If you try to rebuild them unlimitedly, these resources may be over-consumed, causing other key services or system functions to be affected. For example, in a cloud computing environment, each time you create a monitoring instance, a certain amount of computing resources and storage resources may be allocated. If you rebuild them unlimitedly, you may exhaust the user's quota or exceed the budget.

[0091] In addition, setting a maximum number of rebuild times encourages developers or operation and maintenance personnel to conduct more in-depth analysis and resolution of the problem after the maximum number of times is reached, in order to find the real cause of the rebuild failure.

[0092] In some examples, it is determined whether the reconstruction time of the monitoring instance resource reaches a preset retry time threshold; wherein the reconstruction time is the time elapsed from the moment when the monitoring instance resource is first recreated to the current moment;

[0093] If the retry time threshold is reached, it is determined that the monitoring instance creation has failed.

[0094] Among them, the preset reconstruction time threshold can be set according to needs; for example, the reconstruction time threshold is 10 minutes. If the time from the first time the monitoring instance resource is recreated to the current moment exceeds 10 minutes, even if the maximum number of reconstructions is not exceeded, the monitoring instance creation is still deemed to have failed.

[0095] Setting a reconstruction time threshold can avoid resource exhaustion and resource locking; the creation of monitoring instance resources involves multiple system resources, such as CPU, memory, storage, and network bandwidth. If there is no time threshold, long-term creation attempts will continue to occupy these resources and affect the normal operation of the system; in addition, long-term creation attempts will hold system resource locks, preventing other services or processes from using these resources.

[0096] In some examples, after periodically checking whether the access port of the monitoring instance is normally accessible, the following is also included:

[0097] If the monitoring instance access port cannot be accessed normally, it is determined whether the inspection time reaches the preset inspection time threshold; the inspection time is the time from the first inspection of the monitoring instance access port to the current time;

[0098] If the check time threshold is not reached, the monitoring instance access port is checked again; if the check time threshold is reached, it is determined that the monitoring instance creation has failed.

[0099] The check time threshold is set according to the requirements; for example, if the check time threshold is set to 10 minutes, and the monitoring instance access port is still inaccessible after the check time of 10 minutes, it means that the monitoring instance still failed to be created. In addition, after the cause of the creation failure is found and corrected, it can be rebuilt again.

[0100] If the process of creating a monitoring instance exceeds the inspection time, it may indicate that there are some difficult-to-solve problems, such as network problems, configuration errors, and dependent service failures. If you promptly determine that it has failed, you can quickly locate the problem, avoid long waits, and improve maintenance efficiency.

[0101] In some examples, before creating a monitoring instance in a service cluster according to the triplet, the following steps are further included:

[0102] Selecting a node in the service cluster to be set as a monitoring node, adding a node resource flag to the monitoring node, and marking the monitoring node as a node of a first attribute; wherein the node resource flag indicates a monitoring object of the monitoring node;

[0103] A tolerance corresponding to the first attribute is added to the monitoring instance resource to support creation of the monitoring instance resource in the monitoring node marked as the first attribute.

[0104] In the scenario of large-scale business clusters, monitoring instances are resource-consuming components that consume a lot of CPU and memory resources, which in turn affects the management and control services or businesses on the same node and causes resource preemption. Therefore, monitoring nodes are created and marked to achieve node exclusivity.

[0105] In practical applications, marking a monitored node as a node with the first attribute means marking the node with a taint. A taint is a resource marking mechanism that allows certain features or restrictions to be added to a node. Only nodes with the corresponding tolerance can be on the node. Tolerance refers to the ability to tolerate taints on a node.

[0106] For example, the node resource flag is such as label node-roles.kubernetes.io / monitor=; the node taint is such as node-role.kubernetes.io / monitor:NoSchedule; and the tolerance is such as node-role.kubernetes.io / monitor:NoSchedule.

[0107] During the creation of monitoring instance resources, they will be scheduled to the monitoring node. The node will reject scheduling requests from other management and control components and services, so that the monitoring instance exclusively occupies the monitoring node, avoiding resource preemption and mutual impact on management and control components or services.

[0108] In some examples, after the monitoring instance ends, the monitoring node, the node resource flag, the monitoring node tag, and the tolerance of the first attribute are cleared.

[0109] When deleting a monitoring node, make sure that no other services depend on the node to avoid monitoring failures caused by incorrect deletion.

[0110] In addition, after the monitoring instance ends, deleting the monitoring node, node resource flag, monitoring node tag, and tolerance of the first attribute is crucial for resource recovery, system maintenance, security, and privacy protection. Through reasonable deletion operations, we can ensure the effective use of system resources, the clarity and security of the system status, and avoid interference and risks to subsequent system operations.

[0111] In some examples, Figure 3 A schematic diagram of a large-scale monitoring method provided in this application Figure 3 ,like Figure 3 As shown in the figure, after creating a monitoring instance in the business cluster according to the triplet, it also includes:

[0112] S301, during the operation of the monitoring instance, check whether the operation of the monitoring instance is normal.

[0113] S302: If the monitoring instance runs abnormally, the monitoring instance is repaired according to the triplet, and the abnormal running state of the monitoring instance is displayed on a visual interface.

[0114] S303: If the monitoring instance runs normally, continue to run the monitoring instance and display the successful running status on the visual interface.

[0115] The monitoring instance will continuously check the running status of the monitoring instance at a preset time interval during operation. Specifically, the normal operation of the monitoring instance is checked by checking whether the proxy service of the monitoring instance is normal.

[0116] In actual applications, for example, the time interval is set to 1 minute; then it will check whether the proxy service of the monitoring instance is normal. If it is normal, it will be displayed on the visualization interface. If it is abnormal, the monitoring instance will be repaired according to the triplet, and the abnormal operation status will be displayed on the visualization interface.

[0117] Then, in the next round of inspection, all monitoring instances will still be checked, regardless of whether the previous inspection was normal or abnormal; after checking that the proxy service of the monitoring instance is normal, it will also be determined whether the operating status displayed on the visualization interface in the previous inspection is consistent with this inspection; if the previous inspection showed a normal operating status, it will continue to display the normal operating status; if the previous inspection showed an abnormal operating status, it will be updated to a normal operating status.

[0118] By providing real-time feedback on the running status of the monitoring instance, it is ensured that the monitoring instance can avoid long-term abnormal operation but normal operation, which may lead to the loss of monitoring data. In addition, there is also the situation where the monitoring instance runs normally for a long time but shows abnormal operation, which causes the operation and maintenance personnel to waste time to troubleshoot non-existent faults.

[0119] Figure 4 A flow chart of creating a monitoring instance based on a state machine provided in this application, such as Figure 4 As shown in the figure, in actual application, in the initialization state, the monitoring instance resource is created. If the monitoring instance resource is successfully created, the state is updated to the check state; if the monitoring instance resource fails to be created, the state is updated to the retry initialization state. In the retry initialization state, the monitoring instance resource that failed to be created is deleted, and the reconstruction is started only after judging whether the maximum number of reconstructions or the reconstruction time threshold is exceeded; if the reconstruction is successful, the state is updated to the check state; if the creation fails, the reconstruction continues within the maximum number of reconstructions or the reconstruction time threshold; until the maximum number of reconstructions or the reconstruction time threshold is exceeded, it is directly determined that the monitoring instance creation failed. In the check state, the port access of the monitoring instance is checked regularly to see if it is abnormal; if the port access is normal, the state is updated to the running state; if the port access is abnormal, it can be checked again within the check time threshold. If the access is still abnormal after the check time threshold, it is determined that the monitoring instance creation failed. In the running state, the monitoring instance is checked in a round-robin manner to see if it can run normally, and the running abnormality is repaired according to the triplet.

[0120] Embodiment 2

[0121] Figure 5A schematic diagram of the structure of a large-scale monitoring device provided in this application, such as Figure 5 As shown, including:

[0122] The processing module 11 is used to split the monitoring objects according to the scale of the business cluster and based on the hierarchical differences and category differences to obtain triples; wherein the triples include monitoring targets, custom rules, and remote writing information.

[0123] Running module 12 is used to create a monitoring instance in the business cluster according to the triplet; and, to run the monitoring instance to obtain monitoring indicator data according to the monitoring target, aggregate the monitoring indicator data according to the custom rules, and store the aggregated monitoring indicator data in the time series database in the management and control cluster according to the remote write information.

[0124] The display module 13 is used to obtain the monitoring indicator data from the time series database and display it on a visual interface.

[0125] Among them, time series databases such as Influx clusters; by setting up a time series database to store monitoring indicator data, the monitoring instance does not need to store monitoring indicator data, but only needs to collect and process data; when the monitoring indicator data needs to be queried, it is queried from the time series database, which avoids the situation in large-scale scenarios where the monitoring instance has to be responsible for both collecting and processing data and reading data, which leads to confusion in monitoring tasks and causes the loss of monitoring indicator data.

[0126] Specifically, the monitoring indicator data that has been aggregated and calculated will be stored in the time series database in the management and control cluster according to the remote write information; the indicator name will be filtered according to the remote write information; the indicators required in each split monitoring instance will be filtered by whitelist fuzzy matching, and uniformly forwarded and stored in the storage time series database; after filtering, it can avoid duplicate and redundant data from polluting the time series database.

[0127] The levels of the monitoring object include Container / Pod level, Node level, Namespace level and Cluster level. The indicators at different levels have different magnitudes. For example, the number of indicators at a single Node level will be dozens to hundreds. In a medium-sized business cluster, the number of Node level indicators is dozens to hundreds, so the number of indicators is thousands to tens of thousands. The number of Cluster level indicators is dozens to hundreds. The number of Cluster level indicators is one, so the number of indicators is dozens to hundreds. It can be seen that there are differences in the magnitude of indicators at different levels, that is, differences in levels. The categories of the monitoring object include CPU, memory, network and file system.

[0128] In actual applications, the magnitudes of different indicators of a business cluster can be divided into multiple magnitudes according to the Cluster level, Namespace level, Node level, Container / Pod level, and service level.

[0129] In actual applications, for example, in a business cluster of 300 nodes, the monitoring objects are split based on the hierarchical and category differences to obtain triplets, and then a monitoring instance is created based on the triplets and the monitoring instance is run to obtain monitoring indicator data based on the monitoring target, and the monitoring indicator data is aggregated and calculated based on custom rules. The aggregated monitoring indicator data is stored in the time series database in the management and control cluster based on remote write information; the monitoring indicator data is obtained from the time series database and displayed on the visual interface. By using the triplets obtained after splitting the monitoring objects and creating monitoring instances based on the triplets, it is ensured that the indicator processing of each monitoring instance is within the peak value that a single monitoring instance can carry, thus avoiding the failure of the monitoring function.

[0130] Among them, splitting the monitoring object and monitoring it by multiple monitoring instances can ensure that the resource consumption of the monitoring instance is controllable and runs smoothly; the following Table 1 shows the average resource consumption of multiple monitoring instances.

[0131] Table 1

[0132]

[0133] Among them, prometheus-cpu-mem-0 indicates the resource consumption during CPU monitoring; prometheus-net-fs-sts-0 indicates the resource consumption during network (net), file system (fs) and status (sts) monitoring. The status includes the running status and process status of the monitoring instance; prometheus-core-0 indicates the resource consumption when aggregating the collected data.

[0134] In some examples, the processing module is further used to split according to the cluster size based on the hierarchical differences and category differences of the monitored objects to obtain triples, specifically including:

[0135] The business cluster scale coefficient and the level index coefficient of the monitored object are multiplied to obtain the index coefficient of the level index of the monitored object under the cluster scale.

[0136] According to the calculated indicator coefficient, constant coefficient, category indicator value and grouping amplitude value, calculation is performed to obtain the indicator processing values ​​of the triplets of different indicator levels obtained by splitting according to the grouping amplitude under each category, and the maximum value of the indicator processing value is selected.

[0137] Subtract the maximum value of the indicator processing value from the indicator processing peak value; if the result of the subtraction is a negative value or a zero value, split it according to the group amplitude value to obtain a triple; among which, the indicator processing peak value represents the indicator peak value that a single monitoring instance can carry.

[0138] Among them, the indicator coefficient of the level indicator of the monitoring object under the cluster scale can define a vector set of indicator coefficients, as shown below:

[0139] in, Represents the Cluster level index coefficient, Indicates the Namespace level index coefficient, represents the node level index coefficient, Indicates the Container / Pod level indicator coefficient. represents the service level index coefficient, Represents the coefficient ratio of different scales.

[0140] In addition, for indicators of the same category, multi-level aggregation calculations are required at the Container / Pod level, Namespace level, Node level, and Cluster level. High-level indicators rely on low-level indicators for aggregation. The category vector set is defined as follows:

[0141]

[0142] in, Indicates the CPU category indicator value. Indicates the memory category indicator value, Indicates the network category indicator value, Indicates the file system category index value. Indicates the status category indicator value, Indicates other category indicator values.

[0143] After the monitoring object is split, several different triplets are generated (scrape jobs, recording rules, remote write) to ensure that each instance can collect necessary and non-repetitive indicators, aggregate and calculate basic indicators and aggregate indicators, and transfer the required indicators to the time series database as needed. The atomic triplets of the monitoring instance are defined as follows:

[0144]

[0145] in, Indicates the __name__ indicator name regular matching rule set of the scrape jobs to capture the monitoring target. Indicates the record rules custom aggregation indicator set. Indicates the regular matching rule set of the __name__ indicator name of the remote write forwarding address.

[0146] The monitoring object splitting processing function is as follows:

[0147]

[0148] in, It indicates the peak value of the indicator that a single monitoring instance can handle. For example, in a specific version and specific software and hardware environment, the monitoring instance can handle a peak value of 100,000 indicators, which is a fixed constant value. Indicates the grouping range after classification, used to indicate the number of indicators in the group after grouping, and its minimum value is 1; Indicates the result value.

[0149] Among them, according to the calculated indicator coefficient, constant coefficient, category indicator value, and grouping amplitude value, the indicator processing values ​​of the triplets of different indicator levels obtained by splitting according to the grouping amplitude in each category are calculated, and the maximum value of the indicator processing value is selected. For example, the triplets obtained after splitting the cpu according to different levels are: split into cpu1, cpu2 at the Node level, and split into cpu3, ​​cpu4 at the Cluster level; calculate the processing values ​​of the processing indicators of cpu1, cpu2, cpu3, ​​and cpu4 respectively; similarly, calculate the triplets obtained by splitting each category according to different levels; then select the maximum value among the processing values ​​of all processing indicators.

[0150] Subtract the selected maximum value from the peak value that a single monitoring instance can carry. If the result is a negative value or zero, it means that the indicator processing capacity of the monitoring instance created based on the split triples is within the peak upper limit, ensuring that the monitoring instance will not experience memory resource explosion and ensuring the stability of the monitoring function.

[0151] In some examples, the processing module 11 is further configured to reduce the group amplitude value by half if the subtraction result value is a positive value;

[0152] According to the indicator coefficient, constant coefficient, category indicator value, and halved group amplitude value, the maximum value of the indicator processing value is calculated again and selected, and the result value of subtracting the maximum value of the indicator processing value from the indicator processing peak value is calculated until a triplet is obtained.

[0153] Among them, if the result value is a positive value, it means that the indicator processing volume of the monitoring instance created based on the split triplet exceeds the peak upper limit, which will cause the monitoring instance of the monitoring instance to fail; split it according to the grouping amplitude value and split it into more triplets to ensure that the indicator processing volume of the monitoring instance created by each triplet does not exceed the upper limit.

[0154] Therefore, in large-scale scenarios, by splitting the monitoring objects, the indicator processing volume of each monitoring instance will not exceed the peak value, thus ensuring the stability of the monitoring instance. In addition, monitoring instances can be flexibly created for changes in the scale of the business cluster to ensure that the memory resources of the monitoring instance can support the monitoring of the current business cluster scale.

[0155] In some examples, the processing module 11 is further used to create a monitoring instance in the service cluster according to the triple, specifically including:

[0156] Create monitoring instance resources according to the triplet; the monitoring instance resources include the monitoring object of the monitoring instance and the proxy service of the monitoring instance; wherein the proxy service of the monitoring instance provides the monitoring instance access port;

[0157] If the monitoring instance resource is created successfully, a check is performed periodically to see whether the monitoring instance access port can be accessed normally. If the monitoring instance access port can be accessed normally, a monitoring instance is created in the business cluster based on the successfully created monitoring instance resource.

[0158] Among them, the proxy service of the monitoring instance is an intermediate layer service located between the monitoring instance and the monitoring object. In actual applications, the monitoring instance may be in a secure internal network environment, which cannot be directly accessed by the external network; the proxy service can be deployed at the network boundary, providing an externally accessible port to authenticate and authorize external requests, and only verified requests will be forwarded to the monitoring instance. In addition, by using the port of the proxy service, external users and services do not need to know the real IP address and port of the monitoring instance, thereby protecting the monitoring instance from direct attacks. The proxy service can use the internal IP address and port to communicate with the monitoring instance, isolating external requests from the internal monitoring instance, and enhancing the security of the monitoring instance.

[0159] In actual applications, the creation of a monitoring instance in a business cluster according to a triplet is based on a state machine; the state machine can display all state changes of the monitoring instance from creation to operation in real time to the user, so that the user can observe the monitoring process in real time.

[0160] When you create a monitoring instance resource for the first time, it will be displayed in the initialization state; after the monitoring resource is successfully created, it will be displayed in the checking state.

[0161] After the monitoring instance resource is successfully created, a check is made periodically to see whether the monitoring instance access port can be accessed normally. For example, a check is made every 10 seconds to see whether the monitoring instance port can be accessed normally. If it can be accessed normally, a monitoring instance is created based on the monitoring instance resource, and the monitoring instance is run and the running status is displayed.

[0162] In some examples, the processing module 11 is further configured to delete the monitoring instance resource that failed to be created if the monitoring instance resource fails to be created; and recreate the monitoring instance resource.

[0163] If the monitoring resource creation fails, it will be converted to the retry initialization state to remind the user that the monitoring instance resource creation failed.

[0164] In actual applications, the deletion of the monitoring instance resources will be divided into multiple batches. The operation of deleting the monitoring instance resources in the state machine is to delete them in batches every 60 seconds. Among them, it will be determined every 15 seconds whether the monitoring instance resources of the current batch are completely cleared. After the monitoring instance resources are completely deleted, the deletion stop status will be displayed.

[0165] When the creation of monitoring instance resources fails, some resources have been allocated or occupied, but the creation process has not been completed, and they are in an unstable or incomplete state. If these failed resources are not deleted, conflicts may occur with these residual resources when they are recreated. For example, some configuration files, database tables, or network ports may have been created, but the corresponding services failed to start normally. When creating again, the new creation process may fail because the file already exists or the port is already occupied.

[0166] Therefore, before each re-creation, the monitoring instance resources that failed to be created are deleted to prevent useless monitoring instance resources from interfering with the creation of subsequent monitoring instances.

[0167] In some examples, the processing module 11 is further configured to, before recreating the monitoring instance resource, further include:

[0168] It is determined whether the historical reconstruction times of the monitoring instance resource have reached the preset maximum reconstruction times; if the maximum reconstruction times have been reached, it is determined that the monitoring instance creation has failed.

[0169] The preset maximum number of reconstruction times is set according to demand; for example, the maximum number is set to 5 times. When the monitoring instance resource still fails to be created after the fifth reconstruction, it is determined that the monitoring instance creation has failed.

[0170] Creating monitoring instance resources may consume various system resources, such as CPU, memory, storage, network bandwidth, etc. If you try to rebuild them unlimitedly, these resources may be over-consumed, causing other key services or system functions to be affected. For example, in a cloud computing environment, each time you create a monitoring instance, a certain amount of computing resources and storage resources may be allocated. If you rebuild them unlimitedly, you may exhaust the user's quota or exceed the budget.

[0171] In addition, setting a maximum number of rebuild times encourages developers or operation and maintenance personnel to conduct more in-depth analysis and resolution of the problem after the maximum number of times is reached, in order to find the real cause of the rebuild failure.

[0172] In some examples, the processing module 11 is further used to determine whether the reconstruction time of the monitoring instance resource reaches a preset retry time threshold; wherein the reconstruction time is the time from the moment when the monitoring instance resource is first recreated to the current moment;

[0173] If the retry time threshold is reached, it is determined that the monitoring instance creation has failed.

[0174] Among them, the preset reconstruction time threshold can be set according to needs; for example, the reconstruction time threshold is 10 minutes. If the time from the first time the monitoring instance resource is recreated to the current moment exceeds 10 minutes, even if the maximum number of reconstructions is not exceeded, the monitoring instance creation is still deemed to have failed.

[0175] Setting a reconstruction time threshold can avoid resource exhaustion and resource locking; the creation of monitoring instance resources involves multiple system resources, such as CPU, memory, storage, and network bandwidth. If there is no time threshold, long-term creation attempts will continue to occupy these resources and affect the normal operation of the system; in addition, long-term creation attempts will hold system resource locks, preventing other services or processes from using these resources.

[0176] In some examples, the processing module 11 is further configured to periodically check whether the monitoring instance access port is normally accessible, and further includes:

[0177] If the monitoring instance access port cannot be accessed normally, it is determined whether the inspection time reaches the preset inspection time threshold; the inspection time is the time from the first inspection of the monitoring instance access port to the current time;

[0178] If the check time threshold is not reached, the monitoring instance access port is checked again; if the check time threshold is reached, it is determined that the monitoring instance creation has failed.

[0179] The check time threshold is set according to the requirements; for example, if the check time threshold is set to 10 minutes, and the monitoring instance access port is still inaccessible after the check time of 10 minutes, it means that the monitoring instance still failed to be created. In addition, after the cause of the creation failure is found and corrected, it can be rebuilt again.

[0180] If the process of creating a monitoring instance exceeds the inspection time, it may indicate that there are some difficult-to-solve problems, such as network problems, configuration errors, and dependent service failures. If you promptly determine that it has failed, you can quickly locate the problem, avoid long waits, and improve maintenance efficiency.

[0181] In some examples, the processing module 11 is further configured to, before creating a monitoring instance in the service cluster according to the triplet, further include:

[0182] Selecting a node in the service cluster to be set as a monitoring node, adding a node resource flag to the monitoring node, and marking the monitoring node as a node of a first attribute; wherein the node resource flag indicates a monitoring object of the monitoring node;

[0183] A tolerance corresponding to the first attribute is added to the monitoring instance resource to support creation of the monitoring instance resource in the monitoring node marked as the first attribute.

[0184] In the scenario of large-scale business clusters, monitoring instances are resource-consuming components that consume a lot of CPU and memory resources, which in turn affects the management and control services or businesses on the same node and causes resource preemption. Therefore, monitoring nodes are created and marked to achieve node exclusivity.

[0185] In practical applications, marking a monitored node as a node with the first attribute means marking the node with a taint. A taint is a resource marking mechanism that allows certain features or restrictions to be added to a node. Only nodes with the corresponding tolerance can be on the node. Tolerance refers to the ability to tolerate taints on a node.

[0186] For example, the node resource flag is such as label node-roles.kubernetes.io / monitor=; the node taint is such as node-role.kubernetes.io / monitor:NoSchedule; and the tolerance is such as node-role.kubernetes.io / monitor:NoSchedule.

[0187] During the creation of monitoring instance resources, they will be scheduled to the monitoring node. The node will reject scheduling requests from other management and control components and services, so that the monitoring instance exclusively occupies the monitoring node, avoiding resource preemption and mutual impact on management and control components or services.

[0188] In some examples, the processing module 11 is further used to clear the monitoring node, the node resource flag, the monitoring node mark, and the tolerance of the first attribute after the monitoring instance ends.

[0189] When deleting a monitoring node, make sure that no other services depend on the node to avoid monitoring failures caused by incorrect deletion.

[0190] In addition, after the monitoring instance ends, deleting the monitoring node, node resource flag, monitoring node tag, and tolerance of the first attribute is crucial for resource recovery, system maintenance, security, and privacy protection. Through reasonable deletion operations, we can ensure the effective use of system resources, the clarity and security of the system status, and avoid interference and risks to subsequent system operations.

[0191] In some examples, the operation module 12 is further configured to create a monitoring instance in the service cluster according to the triplet, and further include:

[0192] During the operation of the monitoring instance, check whether the monitoring instance is running normally.

[0193] If the monitoring instance runs abnormally, the monitoring instance is repaired according to the triplet, and the abnormal running state of the monitoring instance is displayed on the visual interface.

[0194] If the monitoring instance runs normally, continue to run the monitoring instance and display the successful running status on the visual interface.

[0195] The monitoring instance will continuously check the running status of the monitoring instance at a preset time interval during operation. Specifically, the normal operation of the monitoring instance is checked by checking whether the proxy service of the monitoring instance is normal.

[0196] In actual applications, for example, the time interval is set to 1 minute; then it will check whether the proxy service of the monitoring instance is normal. If it is normal, it will be displayed on the visualization interface. If it is abnormal, the monitoring instance will be repaired according to the triplet, and the abnormal operation status will be displayed on the visualization interface.

[0197] Then, in the next round of inspection, all monitoring instances will still be checked, regardless of whether the previous inspection was normal or abnormal; after checking that the proxy service of the monitoring instance is normal, it will also be determined whether the operating status displayed on the visualization interface in the previous inspection is consistent with this inspection; if the previous inspection showed a normal operating status, it will continue to display the normal operating status; if the previous inspection showed an abnormal operating status, it will be updated to a normal operating status.

[0198] By providing real-time feedback on the running status of the monitoring instance, it is ensured that the monitoring instance can avoid long-term abnormal operation but normal operation, which may lead to the loss of monitoring data. In addition, there is also the situation where the monitoring instance runs normally for a long time but shows abnormal operation, which causes the operation and maintenance personnel to waste time to troubleshoot non-existent faults.

[0199] In actual applications, in the initialization state, the monitoring instance resources are created. If the monitoring instance resources are created successfully, the state is updated to the check state; if the monitoring instance resources fail to be created, the state is updated to the retry initialization state. In the retry initialization state, the monitoring instance resources that failed to be created are deleted, and the reconstruction is started only after determining whether the maximum number of reconstructions or the reconstruction time threshold has been exceeded; if the reconstruction is successful, the state is updated to the check state; if the creation fails, the reconstruction continues within the maximum number of reconstructions or the reconstruction time threshold; until the maximum number of reconstructions or the reconstruction time threshold is exceeded, it is directly determined that the monitoring instance creation has failed. In the check state, the port access of the monitoring instance is checked regularly to see if it is abnormal; if the port access is normal, the state is updated to the running state; if the port access is abnormal, it can be checked again within the check time threshold. If the access is still abnormal after the check time threshold, it is determined that the monitoring instance creation has failed. In the running state, the monitoring instance is checked in a round-robin manner to see if it can run normally, and the running abnormality is repaired according to the triplet.

[0200] A monitoring device for a large-scale scenario provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and this embodiment will not be described in detail here.

[0201] Figure 6 This is a schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the device 60 also includes a communication component 603. The processor 601, the memory 602 and the communication component 603 are connected via a bus 604.

[0202] In a specific implementation process, at least one processor 601 executes the computer execution instructions stored in the memory 602, so that at least one processor 601 executes the above method.

[0203] The specific implementation process of the processor 601 can be found in the above method embodiment, and its implementation principle and technical effect are similar, so this embodiment will not be repeated here.

[0204] In the above embodiments, it should be understood that the processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the invention can be directly implemented as a hardware processor, or can be implemented by a combination of hardware and software modules in the processor.

[0205] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (NVM), such as at least one disk storage.

[0206] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.

[0207] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0208] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the above method is implemented.

[0209] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The readable storage medium can be any available medium that can be accessed by a general or special-purpose computer.

[0210] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0211] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0212] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0213] Finally, it should be noted that those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. The present invention is intended to cover any variations, uses or adaptations of the present invention, which follow the general principles of the present invention and include common knowledge or customary technical means in the art not disclosed by the present invention, are not limited to the precise structure described above and shown in the drawings, and may be modified and changed in various ways without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A large-scale scene monitoring method, characterized in that: The method comprises: According to the scale of the business cluster, the monitoring objects are split based on the hierarchical and category differences to obtain triples; wherein the triples include monitoring targets, custom rules, and remote writing information; Creating a monitoring instance in the business cluster according to the triple; and running the monitoring instance to obtain monitoring indicator data according to the monitoring target, performing aggregation calculation on the monitoring indicator data according to the custom rule, and storing the aggregated monitoring indicator data in a time series database in the management and control cluster according to the remote write information; The monitoring indicator data is obtained from the time series database and displayed on a visualization interface.

2. The method according to claim 1, characterized in that The splitting is performed based on the cluster size and the hierarchical and category differences of the monitored objects to obtain triples, specifically including: Multiplying the business cluster scale coefficient and the level index coefficient of the monitored object to obtain the index coefficient of the level index of the monitored object under the cluster scale; Calculate the indicator coefficient, constant coefficient, category indicator value, and grouping amplitude value obtained by the calculation to obtain the indicator processing values ​​of the triplets of different indicator levels obtained by splitting according to the grouping amplitude under each category, and select the maximum value of the indicator processing value; Subtract the maximum value of the indicator processing value from the indicator processing peak value; if the result of the subtraction is a negative value or a zero value, split it according to the grouping amplitude value to obtain a triple; wherein the indicator processing peak value represents the indicator peak value that a single monitoring instance can carry.

3. The method according to claim 2, characterized in that The method further comprises: If the subtraction result is a positive value, the grouping amplitude value is halved; According to the indicator coefficient, constant coefficient, category indicator value, and halved group amplitude value, the maximum value of the indicator processing value is calculated again and selected, and the result value of subtracting the maximum value of the indicator processing value from the indicator processing peak value is calculated until a triplet is obtained.

4. The method according to claim 1, characterized in that The creating a monitoring instance in the service cluster according to the triplet specifically includes: Creating a monitoring instance resource according to the triple; the monitoring instance resource includes a monitoring object of the monitoring instance and a proxy service of the monitoring instance; wherein the proxy service of the monitoring instance provides a monitoring instance access port; If the monitoring instance resource is successfully created, a check is performed periodically to determine whether the monitoring instance access port can be accessed normally; if the monitoring instance access port can be accessed normally, a monitoring instance is created in the business cluster according to the successfully created monitoring instance resource.

5. The method according to claim 4, characterized in that The method further comprises: If the monitoring instance resource fails to be created, delete the monitoring instance resource that failed to be created; Re-create the monitoring instance resources.

6. The method according to claim 5, characterized in that Before recreating the monitoring instance resources, the following step is also included: It is determined whether the historical reconstruction times of the monitoring instance resource have reached the preset maximum reconstruction times; if the maximum reconstruction times have been reached, it is determined that the monitoring instance creation has failed.

7. The method according to claim 6, characterized in that The method further comprises: Determine whether the reconstruction time of the monitoring instance resource reaches a preset retry time threshold; wherein the reconstruction time is the time from the moment when the monitoring instance resource is first recreated to the current moment; If the retry time threshold is reached, it is determined that the monitoring instance creation has failed.

8. The method according to claim 4, characterized in that After the periodic checking of whether the monitoring instance access port can be accessed normally, the method further includes: If the monitoring instance access port cannot be accessed normally, it is determined whether the inspection time reaches a preset inspection time threshold; the inspection time is the time from the first inspection of the monitoring instance access port to the current time; If the check time threshold is not reached, the monitoring instance access port is checked again; if the check time threshold is reached, it is determined that the monitoring instance creation has failed.

9. The method according to claim 1, characterized in that: Before creating a monitoring instance in the service cluster according to the triplet, the method further includes: Selecting a node in the service cluster to be set as a monitoring node, adding a node resource flag to the monitoring node, and marking the monitoring node as a node of a first attribute; wherein the node resource flag indicates a monitoring object of the monitoring node; A tolerance corresponding to the first attribute is added to the monitoring instance resource to support creation of the monitoring instance resource in the monitoring node marked with the first attribute.

10. The method according to claim 9, characterized in that The method further comprises: After the monitoring instance runs, the monitoring node, the node resource flag, the tag of the monitoring node, and the tolerance of the first attribute are cleared.

11. The method according to any one of claims 1 to 10, characterized in that: After creating a monitoring instance in the service cluster according to the triplet, the method further includes: During the operation of the monitoring instance, checking whether the operation of the monitoring instance is normal; If the monitoring instance runs abnormally, repairing the monitoring instance according to the triplet, and displaying the abnormal running state of the monitoring instance on the visualization interface; If the monitoring instance runs normally, the monitoring instance continues to run and the successful running status is displayed on the visual interface.

12. A monitoring device for large-scale scenes, characterized in that: include: A processing module is used to split the monitored objects according to the scale of the business cluster and the hierarchical and category differences to obtain triples; wherein the triples include the monitored targets, the custom rules, and the remote writing information; An operation module is used to create a monitoring instance in the business cluster according to the triple; and to run the monitoring instance to obtain monitoring indicator data according to the monitoring target, aggregate the monitoring indicator data according to the custom rule, and store the aggregated monitoring indicator data in a time series database in the management and control cluster according to the remote write information; The display module is used to obtain the monitoring indicator data from the time series database and display it on a visual interface.

13. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 11 when executed by a processor.

15. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 11 when being executed by a processor.