A method and system for containerized deployment of a cluster

Through the containerized deployment of the cluster and the consistent hashing algorithm configuration, combined with the data playback mechanism and the Gossip protocol, the stability and data integrity problems of the monitoring and alarm system when the node is restarted are solved, and the reliable and stable operation of the monitoring and alarm system is achieved.

CN119718442BActive Publication Date: 2025-07-22SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510206490.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-07-22
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

The existing monitoring and alarm system has hidden stability risks when the node restarts, resulting in short-term alarm loss or global jitter, and the containerized deployment cannot be achieved.

Method used

The containerized deployment method of the cluster is adopted, and the alarm rules are configured using a consistent hashing algorithm, and the data loss of the restart node is compensated through the data replay mechanism, and the alarm status of the Gossip protocol is synchronized to achieve load balancing.

Benefits of technology

It ensures the stability and data integrity of the monitoring and alarm system when the node is restarted, avoids jitter and data loss, and ensures that the system is reliable and continuous operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119718442B_ABST
    Figure CN119718442B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of monitoring deployment, and particularly relates to a method and system for containerized deployment of a cluster. The method includes: performing cluster initialization for a restarted alarm computing node, obtaining all current alarm states of the cluster for merging after the cluster initialization is completed, and reading and loading all current alarm rules of the cluster; wherein the alarm state is the execution state of an alarm rule within a preset time period; configuring the alarm rules run by the restarted alarm computing node by using the consistent hashing algorithm; when the restarted alarm computing node starts to periodically calculate the configured alarm rules, determining whether data replay is required based on the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm states, so as to make up for the data missing in a short time of the restarted alarm computing node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of monitoring deployment, and particularly relates to a method and system for containerized deployment of a cluster. Background Art

[0002] Currently, existing monitoring and alarm systems already integrate simple alarm capabilities. Users can import alarm rules into the system, and the monitoring and alarm system performs periodic monitoring according to the imported alarm rules and executes an alarm when the trigger conditions are met.

[0003] However, when the monitoring scale increases, due to the single-node deployment method of its alarm system, there are stability hidden dangers and containerized deployment cannot be achieved. If a single node restarts, it will cause the loss of short-term alarms or the jitter of global alarms. Summary of the Invention

[0004] In view of the above-mentioned disadvantages of the prior art, the purpose of the present invention is to propose a containerized deployment solution for a cluster, which can avoid the jitter caused by node downtime or restart to the monitoring and alarm system, and can prevent data loss, thereby ensuring the reliable and stable continuous operation of the monitoring and alarm system.

[0005] To achieve the above object and other related objects, the present invention provides a method for containerized deployment of a cluster, which is applied to a monitoring and alarm system. The monitoring and alarm system includes at least one cluster, and each cluster includes at least one alarm calculation node, which is used to periodically calculate preset alarm rules to send a notification to an alarm management module when the trigger conditions of the alarm rules are met. The method includes: performing cluster initialization for the restarted alarm calculation node, and after the cluster initialization is completed, obtaining all the current alarm states of the cluster for merging, and reading and loading all the current alarm rules of the cluster; wherein, the alarm state is the execution state of an alarm rule within a preset time period; configuring the alarm rules run by the restarted alarm calculation node by using a consistent hashing algorithm; when the restarted alarm calculation node starts to periodically calculate the configured alarm rules, judging whether data replay is required according to the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm state, so as to make up for the data missing in the short time of the restarted alarm calculation node.

[0006] According to a specific embodiment of the present invention, the steps of performing cluster initialization for a restarted alarm calculation node, obtaining all current alarm states of the cluster for merging after the cluster initialization is completed, and reading and loading all current alarm rules of the cluster include: detecting whether there are other alarm calculation nodes; if so, joining the corresponding cluster and sending an online notification to all alarm calculation nodes in the cluster; if not, creating a cluster and waiting for alarm rules to be added and other alarm calculation nodes to join; after the restarted alarm calculation node joins the cluster, obtaining all alarm states from other alarm calculation nodes and performing merging; after the restarted alarm calculation node obtains the alarm states, reading and loading all alarm rules from the database.

[0007] According to a specific embodiment of the present invention, the steps of configuring the alarm rules run by the restarted alarm calculation node using the consistent hashing algorithm include: for each of all alarm rules, judging whether the alarm rule should run on the restarted alarm calculation node according to the consistent hashing algorithm; if not, prohibiting the restarted alarm calculation node from periodically calculating the alarm rule.

[0008] According to a specific embodiment of the present invention, the steps of judging whether data replay is needed based on the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm states to fill in the data missing from the restarted alarm calculation node in a short period of time include: if the difference between the last execution time and the current calculation time is greater than a preset timeout time, no replay is performed; if the difference between the last execution time and the current calculation time is less than the preset operation period of the alarm calculation node, no replay is performed; if the difference between the last execution time and the current calculation time is greater than the preset operation period of the alarm calculation node and less than the preset replayable time of the alarm calculation node, replay is performed; wherein, after the alarm calculation node completes the replay, it periodically calculates the alarm rule to update the local alarm state.

[0009] According to a specific embodiment of the present invention, the steps of the alarm calculation node performing data replay include: starting from the last execution time, for each time period, calling the corresponding data from the cached data for replay and calculating the alarm rule to judge whether the alarm rule is triggered until the current calculation time.

[0010] According to a specific embodiment of the present invention, each alarm calculation node in the cluster is used to manage and maintain all current alarm states of the cluster, and regularly pushes the stored alarm states to other alarm calculation nodes in the cluster in the form of TCP, and / or regularly pulls the alarm states from other alarm calculation nodes in the form of TCP.

[0011] According to a specific embodiment of the present invention, after the alarm status stored in each alarm calculation node in the cluster is updated, it is sent to other alarm calculation nodes in the cluster in the form of UDP.

[0012] According to a specific embodiment of the present invention, any one alarm calculation node in the cluster only stores the alarm status at the latest time of the same alarm rule; wherein, when the alarm calculation node obtains the alarm status from other alarm calculation nodes, it judges whether it stores the alarm status of this alarm rule: if not, it directly merges the alarm status; if so, it retains the one with the latest time among the alarm status and the alarm status stored by itself, and clears the other.

[0013] According to a specific embodiment of the present invention, each alarm calculation node uses the gossip protocol to synchronously share the alarm status.

[0014] A containerized deployment system for a cluster, which is applied to a monitoring and alarm system. The monitoring and alarm system includes at least one cluster, and each cluster includes at least one alarm calculation node, which is used to periodically calculate a preset alarm rule to send a notification to the alarm management module when the trigger condition of the alarm rule is met. The system includes: a cluster initialization module, which is used to perform cluster initialization for the restarted alarm calculation node, and after the cluster initialization is completed, obtain all the current alarm statuses of the cluster for merging, and read and load all the current alarm rules of the cluster; wherein, the alarm status is the execution status of an alarm rule within a preset time period; a load balancing module, which is used to configure the alarm rules run by the restarted alarm calculation node by using the consistent hashing algorithm; a data replay module, which is used to judge whether data replay is needed when the restarted alarm calculation node starts to periodically calculate the configured alarm rule, based on the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm status, so as to make up for the data missing in a short time by the restarted alarm calculation node. Description of the Drawings

[0015] Figure 1 It is a schematic flowchart of a specific embodiment of a containerized deployment method for a cluster provided by the present invention;

[0016] Figure 2 It is a schematic flowchart of another specific embodiment of a containerized deployment method for a cluster provided by the present invention;

[0017] Figure 3 It is a schematic diagram of a specific embodiment of short-time data missing provided by the present invention;

[0018] Figure 4 It is a schematic structural diagram of a specific embodiment of a containerized deployment system for a cluster provided by the present invention. Detailed implementation manners

[0019] To facilitate the understanding of the present application, the present application will be described more comprehensively below with reference to the relevant accompanying drawings. Embodiments of the present application are shown in the drawings. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present application more thorough and comprehensive.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used in the description of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0021] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0022] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0023] Embodiment 1

[0024] Please refer to Figure 1 A containerized deployment method for a cluster shown, which is applied to a monitoring and alarm system, includes:

[0025] Step S100: Perform cluster initialization for the restarted alarm computing nodes, and after the cluster initialization is completed, obtain all the current alarm states of the cluster for merging, and read and load all the current alarm rules of the cluster; wherein, the alarm state is the execution state of an alarm rule within a preset time period.

[0026] It should be noted that to cope with complex scenarios with a large monitoring scale, the monitoring and alarm system can be composed of multiple clusters, and each cluster includes at least one alarm computing node, which is used to periodically calculate alarm rules to detect whether the alarm rules are triggered, and send a notification to the alarm management module when the alarm rules are triggered.

[0027] The alarm status is the execution context of the alarm rule in a timeline, for example, including: last_exec_at (the last execution time), last_updated_at (the last update time of the rule), firing (in the process of being triggered), resolved (recovered), pending (awaiting to be triggered), happened_at (start time), end_at (end time), and so on.

[0028] It can be understood that the last execution time refers to the moment when the alarm rule was last triggered, the last update time of the rule refers to the moment when the alarm rule was last readjusted, firing, resolved, and pending are all the execution results corresponding to the alarm rule, while the start time and end time represent the start and end moments of this timeline.

[0029] It can also be understood that different alarm calculation nodes can be configured to calculate the same alarm rule to detect whether it is triggered. Correspondingly, different alarm calculation nodes may save different time versions of the alarm status. For example, if a certain alarm calculation node crashes abnormally at a certain moment, the timeliness of the alarm status it saves will lag behind that saved by another alarm calculation node.

[0030] It should also be added that each alarm calculation node in the cluster can be used to manage and maintain all the alarm statuses of the current cluster. When any alarm calculation node regularly runs the alarm rule calculation, it will spread the alarm status corresponding to the alarm rule to other alarm calculation nodes, and other alarm calculation nodes will merge it with the data stored in themselves after receiving the alarm status.

[0031] Here, it can be understood that the alarm calculation node will store the alarm statuses of different alarm rules. Regarding the alarm status corresponding to the same alarm rule, the alarm calculation node will only save the one with the best timeliness. For example, among the different alarm statuses of the same alarm rule, if the execution time of the alarm rule in one alarm status is earlier than that in another alarm status, then the alarm calculation node will only save the latter and discard the former. Similarly, if the alarm calculation node saves multiple alarm statuses, then only the alarm status with the latest time will be retained. Therefore, when the alarm calculation node merges the received alarm status, it will first judge whether it has already stored the alarm status of this alarm rule. If it has not stored the alarm status of this alarm rule before, then the alarm status can be directly merged. However, if it has stored the alarm status of this alarm rule and obtains another alarm status of this alarm rule through other means, then it is necessary to identify the timeliness of the two alarm statuses to judge which one to retain, that is, to retain the alarm status with the latest time.

[0032] Specifically, in practical applications, an alarm calculation node periodically runs the calculation of an alarm rule to detect whether it is triggered, and sends the corresponding detection result, that is, the alarm status of the alarm rule, to other alarm calculation nodes to synchronously update the status of the alarm rule. Moreover, since the alarm calculation node needs to synchronously send each time the alarm status is updated, the execution frequency of this sending method is relatively high. Therefore, the alarm calculation node can send data in the form of UDP (User Datagram Protocol), and try to merge the data to the size of the MTU (Maximum Transmission Unit) for sending. In addition, since each alarm calculation node stores the alarm status of all alarm rules, it will be periodically pushed to other alarm calculation nodes in batches, or pulled from other alarm calculation nodes in batches. And, since the execution frequency of this sending method is relatively low, the alarm calculation node can send data in the form of TCP (Transmission Control Protocol).

[0033] In one embodiment, different alarm calculation nodes can use the gossip protocol to synchronously update the alarm status of the alarm rule according to the above method.

[0034] Based on the above, when an alarm calculation node crashes abnormally and causes data loss, it is necessary to redeploy after restart. In this regard, for the restarted alarm calculation node, it is first necessary to perform cluster initialization, that is, to detect whether there are other alarm calculation nodes. If so, try to connect to other alarm calculation nodes in the current cluster to join the cluster, and after joining, send the online notification synchronously to each alarm calculation node in the cluster. Of course, if there are no other alarm calculation nodes, it means that the cluster initialization fails, and a cluster is created alone, the cluster size is set to 1, and wait for the corresponding alarm rules and other alarm calculation nodes to join.

[0035] For the alarm calculation node with successful cluster initialization, it can obtain all the alarm statuses from other alarm calculation nodes for merging to start synchronizing data. Moreover, since the restarted alarm calculation node obtains the alarm status from different alarm calculation nodes, it will obtain the alarm status of the same alarm rule in different time versions. Accordingly, retain the version with a newer version, that is, the alarm status at the latest time, according to the above method, and discard the expired alarm status.

[0036] At the same time, after synchronizing the alarm status, the restarted alarm calculation node will also read and load all the alarm rules from the database or file to import the alarm rules and start running the calculation.

[0037] Step S200, configure the alarm rules run by the restarted alarm calculation node by using the consistent hashing algorithm.

[0038] It should be noted here that the restarted alarm calculation node based on the above steps needs to read and load all the rules. However, in fact, each alarm calculation node is not used to calculate all the alarm rules, but only needs to be responsible for the operation and calculation of some alarm rules, and the query and monitoring of all alarm rules are maintained through the collaborative work of all alarm calculation nodes in the cluster. Therefore, for each alarm rule loaded by the restarted alarm calculation node, the consistent hash algorithm is used to determine whether the alarm rule should run on this alarm calculation node. For example, if a certain alarm rule has been redundantly run and calculated by one or more other alarm calculation nodes, there is no need for other alarm calculation nodes to be responsible. In this regard, when it is judged by the hash algorithm that the alarm rule does not meet the expectation, that is, it should not run and calculate on the restarted alarm calculation node, the corresponding action of this alarm calculation node is prohibited.

[0039] It should also be added that based on the load balancing strategy, when the alarm calculation nodes in the cluster go online and offline, that is, when they crash or restart, a large number of reorientations of alarm rules will occur, causing unnecessary jitters. However, using the consistent hash algorithm on each alarm calculation node can achieve a better balancing effect and avoid causing jitters. When a certain alarm calculation node crashes and is removed from the cluster, or when it restarts and is added to the cluster, using the consistent hash algorithm will only affect the adjacent successor nodes of this node on the hash ring and will not cause jitters to other nodes in the cluster. Therefore, through the consistent hash algorithm, not only can the alarm rules run by the alarm calculation node be reasonably configured when it restarts and goes online, but also the alarm rules it runs can be assigned to other alarm calculation nodes when it crashes and goes offline, thus ensuring the stable operation of the entire monitoring and alarm system.

[0040] Step S300, when the restarted alarm calculation node starts to periodically calculate the configured alarm rules, judge whether data replay is required based on the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm status, so as to make up for the data missing in a short time for the restarted alarm calculation node.

[0041] Based on the above steps S100 and S200, the redeployment of the restarted alarm calculation node has been basically completed. Further, considering that in various scenarios such as system release, pod migration, and storage jitter, the relevant data for the alarm calculation node to run and calculate alarm rules may be missing for a short time, and it is necessary to make up the data for itself. For example, Figure 3As shown in the figure, assume that for a certain alarm rule, the alarm calculation node runs the calculation regularly every 30s. However, after the alarm calculation node crashes and restarts, the migration of the alarm rule is triggered, missing the data at 3 intermediate time points, which can easily lead to data jitter and loss.

[0042] Therefore, after reconfiguring the alarm rule using the hash algorithm, it is necessary to identify the alarm rule currently being calculated by the restarted alarm calculation node. Specifically, it is determined whether data replay is required based on the current calculation time of an alarm rule and the last execution time of the alarm rule in the alarm status merged by the alarm calculation node before. In this regard, if the last execution time is too long, that is, the difference between the last execution time and the current calculation time is greater than the preset timeout, then there is no need for replay accordingly. It can be understood here that since the data has been missing for a long time, the significance of data supplementation is not great, and there is no need for replay accordingly. If the difference between the last execution time and the current calculation time is less than the preset running period of the alarm calculation node, it can be seen that the data loss does not exceed one cycle of the alarm calculation node running the alarm rule calculation, and there is no need for replay either. Just wait for the next time to run the calculation of this alarm rule. If the difference between the last execution time and the current calculation time is greater than the preset running period of the alarm calculation node and less than the preset replayable time of the alarm calculation node, then replay can be performed to supplement the short-term missing data.

[0043] During replay, starting from the last execution time of the alarm rule, for each missing time period, the data corresponding to the missing time period is called from the cached data for replay, and the alarm rule is run for calculation during data replay to detect whether the alarm rule is triggered, so as to complete the data supplementation of the alarm calculation node until the current calculation time is replayed, and the alarm status corresponding to the alarm rule in the missing time can be obtained. It should also be added here that expired alarm statuses can be masked during data replay to avoid generating invalid noise and affecting data replay.

[0044] After completing data replay, the alarm calculation node can run the calculation normally, and after each calculation, the corresponding result is used to update the alarm status stored locally. At the same time, the alarm status will also be stored in the buffer area to be merged to the MTU size and then sent to other alarm calculation nodes in the cluster in the form of UDP.

[0045] Based on the above steps, the containerized deployment of the restarted alarm calculation node can be realized. In addition, the containerized deployment provided in this embodiment still needs to introduce certain external dependencies, such as the MySQL storage medium (the MySQL storage medium does not affect the sharding effect of the alarm rule), and the application load balancer is used to realize alarm service discovery. Otherwise, the restarted alarm calculation node cannot obtain all the nodes in the cluster.

[0046] It should be noted that the step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process are all within the protection scope of this application.

[0047] Embodiment 2

[0048] Based on the technical solution provided in the above Embodiment 1, this embodiment is also used to provide its specific application scenarios.

[0049] In practical applications, the above monitoring and alarm system can be built based on Prometheus. Of course, it is not only applicable to Prometheus. There is no limitation on this. Only Prometheus is taken as an example.

[0050] For any cluster in Prometheus, when the alarm calculation node in it crashes abnormally and restarts, the cluster can be initialized.

[0051] For the new node after the cluster initialization is completed, all Alert States (alarm states) can be obtained from other nodes and merged. At the same time, all alarm rules can be read from the database or file and loaded, and according to the consistent hash algorithm, it is judged whether the rule should run on the current node to complete the configuration of the alarm rules.

[0052] Furthermore, the new node will also judge whether data replay is required for the running alarm rules, that is, judge according to the current calculation time of the alarm rule and the last execution time of the alarm rule merged by the alarm calculation node before. During replay, start running from last_exec_at, increase one running cycle each time, and then perform a time series query until the current time. After the replay is completed, the alarm rules configured by the new node run regularly. After each run, the execution result of the rule is updated to the local Alert State, stored in the buffer, and merged into the MTU and sent to other nodes in the cluster in the form of UDP packets.

[0053] In addition, the new node will also regularly push all its Alert States to other nodes in the cluster in the form of TCP, and regularly pull all the Alert States from other nodes in the cluster in the form of TCP, and after receiving the Alert States returned by other nodes, update them to its own Alert State record.

[0054] Meanwhile, when the alarm rule is triggered, the new node will also send a notification to the Alert Manager.

[0055] During the above process, when other nodes detect the online and offline status of the new node, they will update the cluster member list and trigger a real-time rule load balancing strategy, thereby maintaining the stable operation of the cluster and the entire Prometheus.

[0056] Embodiment 3

[0057] Please refer to Figure 4 As shown, this embodiment also provides a containerized deployment system for a cluster, including:

[0058] A cluster initialization module 10, which is used to perform cluster initialization for the restarted alarm calculation node, obtain all the current alarm states of the cluster for merging after the cluster initialization is completed, and read and load all the current alarm rules of the cluster; wherein, the alarm state is the execution state of an alarm rule within a preset time period.

[0059] A load balancing module 20, which is used to configure the alarm rules run by the restarted alarm calculation node by using the consistent hashing algorithm.

[0060] A data replay module 30, which is used to determine whether data replay is required when the restarted alarm calculation node starts to periodically calculate the configured alarm rules, based on the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm state, so as to make up for the data missing in a short period of time by the restarted alarm calculation node.

[0061] It should be noted that the containerized deployment system for the cluster provided in the above embodiment belongs to the same concept as the containerized deployment method for the cluster provided in the above Embodiment 1. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the containerized deployment method for the cluster provided in the above Embodiment 1 can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here either.

[0062] In summary, the present invention provides a containerized deployment method for a cluster, which can automatically perform load balancing when the alarm calculation node crashes and restarts, so as not to affect the normal operation of the existing nodes, and uses the Gossip protocol to complete the synchronization of the execution status of the alarm rules within the cluster, and avoids data through the replay mechanism, thereby ensuring the reliable and stable continuous operation of the monitoring and alarm system.

[0063] The above embodiments are only illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed by the present invention should still be covered by the claims of the present invention.

Claims

1. A method for containerized deployment of a cluster, characterized in that Applied to a monitoring and alarm system, the monitoring and alarm system includes at least one cluster, and each cluster includes at least one alarm calculation node, which is used to periodically calculate preset alarm rules to send notifications to the alarm management module when the triggering conditions of the alarm rules are met. The method includes: Perform cluster initialization for the restarted alarm calculation node, and after the cluster initialization is completed, obtain all the current alarm states of the cluster for merging, and read and load all the current alarm rules of the cluster; wherein, the alarm state is the execution state of an alarm rule within a preset time period, and when the restarted alarm calculation node obtains the alarm states of different time versions of the same alarm rule, only the alarm state of the latest time version is retained; Configure the alarm rules run by the restarted alarm calculation node using the consistent hashing algorithm; When the restarted alarm calculation node starts to periodically calculate the configured alarm rules, determine whether data replay is required based on the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm state, so as to make up for the data missing in a short time of the restarted alarm calculation node.

2. The containerized deployment method of the cluster according to claim 1, characterized in that The steps of performing cluster initialization for the restarted alarm calculation node, and after the cluster initialization is completed, obtaining all the current alarm states of the cluster for merging, and reading and loading all the current alarm rules of the cluster include: Detect whether there are other alarm calculation nodes; If so, join the corresponding cluster and send an online notification to all the alarm calculation nodes in the cluster; If not, create a cluster and wait for alarm rules to be added and other alarm calculation nodes to join; After the restarted alarm calculation node joins the cluster, obtain all the alarm states from other alarm calculation nodes and merge them; After the restarted alarm calculation node obtains the alarm state, read and load all the alarm rules from the database.

3. The containerized deployment method of the cluster according to claim 1, wherein The steps of configuring the alarm rules run by the restarted alarm calculation node using the consistent hashing algorithm include: For each of all the alarm rules, Judge whether the alarm rule should run on the restarted alarm calculation node according to the consistent hashing algorithm: If not, prohibit the restarted alarm calculation node from periodically calculating the alarm rule.

4. The containerized deployment method of the cluster according to claim 1, characterized in that The steps of determining whether data replay is required based on the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm state, so as to make up for the data missing in a short time of the restarted alarm calculation node include: If the difference between the last execution time and the current calculation time is greater than the preset timeout time, no replay is performed; If the difference between the last execution time and the current calculation time is less than the preset running period of the alarm calculation node, no replay is performed; If the difference between the last execution time and the current calculation time is greater than the preset running period of the alarm calculation node and less than the preset replayable time of the alarm calculation node, replay is performed; Wherein, after the alarm calculation node completes the replay, it periodically calculates the alarm rules to update the local alarm state.

5. The containerized deployment method of the cluster according to claim 1 or 4, characterized in that The steps for the alarm calculation node to perform data replay include: Starting from the last execution time, for each time period, call the corresponding data from the cached data for replay, and calculate the alarm rules to determine whether the alarm rules are triggered until the current calculation time.

6. The containerized deployment method of the cluster according to claim 1, characterized in that Each alarm calculation node in the cluster is used to manage and maintain all the current alarm states of the cluster, and periodically push the stored alarm states to other alarm calculation nodes in the cluster in the form of TCP, and / or periodically pull the alarm states from other alarm calculation nodes in the form of TCP.

7. The containerized deployment method of the cluster according to claim 6, wherein Each alarm calculation node in the cluster sends the updated alarm state to other alarm calculation nodes in the cluster in the form of UDP after the stored alarm state is updated.

8. The containerized deployment method of the cluster according to claim 6 or 7, characterized in that Any alarm calculation node in the cluster only stores the alarm state of the latest time of the same alarm rule; Among them, when the alarm calculation node obtains the alarm state from other alarm calculation nodes, it judges whether it stores the alarm state of this alarm rule itself: If not, directly merge the alarm state; If so, keep the one with the latest time among the alarm state and the alarm state stored by itself, and clear the other.

9. The method for containerized deployment of a cluster according to claim 6 or 7, characterized in that Each alarm calculation node uses the gossip protocol to synchronize and share the alarm state.

10. A containerized deployment system for a cluster, characterized in that, Applied to a monitoring alarm system, the monitoring alarm system includes at least one cluster, and each cluster includes at least one alarm calculation node, which is used to periodically calculate preset alarm rules to send a notice to the alarm management module when the triggering conditions of the alarm rules are met. The system includes: A cluster initialization module, which is used to perform cluster initialization for the restarted alarm calculation nodes, and after the cluster initialization is completed, obtain all the current alarm states of the cluster for merging, and read and load all the current alarm rules of the cluster; among them, the alarm state is the execution state of an alarm rule within a preset time period, and when the restarted alarm calculation node obtains the alarm states of different time versions of the same alarm rule, only the alarm state of the latest time version is retained; A load balancing module, which is used to configure the alarm rules run by the restarted alarm calculation nodes by using the consistent hashing algorithm; A data replay module, which is used to judge whether data replay is required according to the current calculation time of the alarm rule and the last execution time of the alarm rule in the merged alarm state when the restarted alarm calculation node starts to periodically calculate the configured alarm rules, so as to make up for the data missing in a short time by the restarted alarm calculation node.

Citation Information

Patent Citations

  • Metadata group design method based on real-time application group

    CN103795801A

  • Prometheus-based alarm method, device and system

    CN115733732A