Distributed alarm system and method

By designing a distributed alarm system, using a consistent hashing algorithm and rule engine to process equipment data, the problems of data delay and single point of failure in the equipment alarm system in the prior art are solved, and a fast response and high availability alarm system is realized.

CN119996158APending Publication Date: 2025-05-13SHANGHAI WPG WISDOM WATER CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411992105.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing equipment alarm system has problems of data delay and untimely response when processing large amounts of data, and it is difficult to avoid single point of failure, resulting in delayed alarms or missing alarms.

Method used

A distributed alarm system is designed to collect the original data of the device through the acquisition unit, and to establish the mapping relationship between the device and the message queue based on the consistency hashing algorithm using the first allocation unit, and to establish the mapping relationship between the message queue and the rule engine through the second allocation unit, and to process the standard data through the rule engine to identify device exceptions.

Benefits of technology

It realizes rapid data processing and timely response to alarms, avoids single point of failure, improves the real-time and accuracy of alarms, and ensures high availability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996158A_ABST
    Figure CN119996158A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed alarm system and method, and belongs to the technical field of water affair monitoring. The distributed alarm system uses the first distribution unit to establish the mapping relation between the equipment and the message queue based on the first preset rule, so that the data of the equipment is uniformly distributed into the message queue, the speed is high, and the conflict rate is low; establishing a mapping relationship between each message queue in the message array and a preset number of rule engines based on a second preset rule by adopting a second distribution unit; and obtaining a standard data packet in the message queue mapped with the rule engine through the rule engine, and distributing standard data in the standard data packet to a corresponding thread for processing according to a third preset rule so as to identify whether equipment associated with the standard data is abnormal or not. The equipment and the rule engine are decoupled through the first distribution unit and the second distribution unit, the data flow is smoothed, when the number of rule engine instances changes, the second distribution unit can distribute the rule engine again, the real-time performance and the accuracy of alarm are improved, and meanwhile the high availability of the alarm system is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water affairs monitoring, and in particular to a distributed alarm system and method. Background Art

[0002] Most of the existing equipment alarm systems in water supply and water plant operations use a single simple threshold alarm rule; however, the rule-driven alarm for non-threshold alarms is not very timely. Moreover, it cannot effectively reduce noise and convergence control for alarm storms, and cannot configure alarm rules for non-continuous intermittent abnormalities.

[0003] With the development of the Internet of Things and Industry 4.0 technology, devices and sensors are constantly expanding, and the amount of data is growing exponentially. When adding new devices or expanding the system, the centralized device alarm system has the risk of single point failure. When the system is deployed and upgraded or there is a single point failure in the system, the device alarm will stagnate for a period of time, during which major safety hazards may arise, and delayed alarms or missed alarms may occur, which is particularly inconvenient in large-scale industrial application scenarios. It can be seen that traditional alarm systems are difficult to cope with such huge data processing needs, which may lead to data backlogs, processing delays, and untimely responses. Summary of the invention

[0004] In view of the data processing delay and untimely response of the existing equipment alarm system, a distributed alarm system and method are provided which aims to have fast data processing speed, timely response to alarms, distributed high availability, and avoid single point failures.

[0005] The present invention provides a distributed alarm system, characterized in that it comprises:

[0006] A collection unit, used for collecting raw data of the equipment;

[0007] A first allocation unit is used to convert the original data into standard data, and establish a mapping relationship between a device associated with the standard data and a message queue in a message array according to the standard data based on a first preset rule, and the message queue obtains the standard data of the corresponding device according to the mapping relationship;

[0008] A second allocation unit, configured to establish a mapping relationship between each of the message queues in the message array and a preset number of rule engines based on a second preset rule;

[0009] The rule engine is used to obtain the standard data packet in the message queue mapped to the rule engine, and allocate the standard data in the standard data packet to the corresponding thread for processing according to a third preset rule to identify whether the device associated with the standard data is abnormal.

[0010] Preferably, the first preset rule adopts a consistent hashing algorithm.

[0011] Preferably, the first allocation unit comprises:

[0012] A conversion module, used for standardizing the original data, converting the original data into the standard data, wherein the standard data includes a device number of a device associated with the standard data;

[0013] The sharding module is used to obtain the mapping relationship between the device number and the message queue number by adopting a consistent hashing algorithm.

[0014] Preferably, the message queue processes the received standard data in a first-in-first-out manner, writes the standard data sequentially into a log file, and backs up the received standard data.

[0015] Preferably, the second preset rule is:

[0016] Based on the number of cores of each of the rule engines, the message queues in the message array are evenly distributed to a preset number of rule engines, so as to establish a mapping relationship between each of the message queues and the preset number of rule engines.

[0017] Preferably, the third preset rule adopts a hash algorithm.

[0018] Preferably, the rule engine includes a rule management module;

[0019] The thread is used to send the standard data to the rule management module;

[0020] The rule management module is used to create an inverted index data structure according to the alarm rule and the device number of the associated standard data, generate a matching rule, and identify whether the device associated with the standard data is abnormal based on the matching rule.

[0021] Preferably, the rule management module is used to create an inverted index data structure according to the alarm rule and the device number of the associated standard data to generate a matching rule, including:

[0022] Create an empty mapping container of device number and device rule list, obtain all alarm rules by monitoring point number, and each alarm rule corresponds to an alarm rule number;

[0023] The first layer loop processes all the acquired alarm rules and acquires all the associated device numbers according to the alarm rule numbers.

[0024] The second layer loop processes all devices and adds the alarm rules of the current first layer loop to the device rule list;

[0025] When the two-layer loop is completed, the device rule list of the current monitoring point number can be obtained;

[0026] The monitoring points in the device monitoring point data are used to obtain a device rule list, and the device unique number is used to obtain a rule list, that is, the matching rule.

[0027] Preferably, the rule management module identifies whether the device associated with the standard data is abnormal based on the matching rule, including:

[0028] Determine whether the data of the monitoring point is within the preset threshold range. If not, it means that the monitoring point is abnormal. If the matching rule does not have a preset alarm time and alarm abnormality percentage, an alarm signal is generated; if the matching rule sets an abnormal percentage noise reduction rule, a data-driven sliding window algorithm is used to generate an alarm signal;

[0029] If the data of the monitoring point is within the preset threshold range, it means that the monitoring point is normal.

[0030] The present invention also provides a distributed alarm method, comprising the following steps:

[0031] S1. Collect the original data of the equipment;

[0032] S2. converting the original data into standard data, establishing a mapping relationship between the device associated with the standard data and the message queue in the message array according to the standard data based on a first preset rule, and obtaining the standard data of the corresponding device according to the mapping relationship;

[0033] S3. A mapping relationship between each of the message queues in the message array and a preset number of rule engines is established based on a second preset rule;

[0034] S4. Obtain a standard data packet in the message queue mapped to the rule engine, and assign the standard data in the standard data packet to a corresponding thread for processing according to a third preset rule to identify whether a device associated with the standard data is abnormal.

[0035] Beneficial effects of the above technical solution:

[0036] In the technical scheme, the distributed alarm system of the present invention collects the original data of the device through the collection unit; uses the first allocation unit to establish the mapping relationship between the device and the message queue based on the first preset rule, so as to evenly distribute the data of the device to the message queue, with fast speed and low conflict rate; uses the second allocation unit to establish the mapping relationship between each message queue in the message array and a preset number of rule engines based on the second preset rule; obtains the standard data packet in the message queue mapped with the rule engine through the rule engine, and distributes the standard data in the standard data packet to the corresponding thread for processing according to the third preset rule to identify whether the device associated with the standard data is abnormal. The first allocation unit and the second allocation unit decouple the device from the rule engine to smooth the data flow. When the number of rule engine instances changes, the second allocation unit can reallocate the rule engine, which improves the real-time and accuracy of the alarm and ensures the high availability of the alarm system. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of a module of an embodiment of the distributed alarm system of the present invention;

[0038] Figure 2 It is a schematic diagram of the working principle of the distributed alarm system of the present invention;

[0039] Figure 3 This is a schematic diagram of the message queue allocation mechanism of the present invention;

[0040] Figure 4 This is a schematic diagram of internal data processing of the rule engine of the present invention;

[0041] Figure 5 It is a schematic diagram of a sliding window;

[0042] Figure 6 The present invention is a flowchart of an embodiment of the distributed alarm method. DETAILED DESCRIPTION

[0043] The advantages of the present invention are further described below in conjunction with the accompanying drawings and specific embodiments.

[0044] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0045] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. The singular forms of "a", "said" and "the" used in this disclosure and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.

[0046] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0047] In the description of the present invention, it should be understood that the numerical labels before the steps do not identify the order of executing the steps, but are only used to facilitate the description of the present invention and distinguish each step, and therefore cannot be understood as a limitation of the present invention.

[0048] The embodiments of the present application are mainly used in the operation and production of water conservancy and water plants. The distributed alarm system collects the original data of the equipment through the collection unit; uses the first allocation unit to establish a mapping relationship between the equipment and the message queue based on the first preset rule, so as to evenly distribute the data of the equipment to the message queue, with fast speed and low conflict rate; uses the second allocation unit to establish a mapping relationship between each message queue in the message array and a preset number of rule engines based on the second preset rule; obtains the standard data packet in the message queue mapped to the rule engine through the rule engine, and distributes the standard data in the standard data packet to the corresponding thread for processing according to the third preset rule to identify whether the equipment associated with the standard data is abnormal. The first allocation unit and the second allocation unit are used to decouple the equipment from the rule engine, smooth the data flow, and when the rule engine sends a transformation, the second allocation unit can reallocate the rule engine, thereby improving the real-time and accuracy of the alarm.

[0049] The present invention proposes a distributed alarm system to solve the defects of the existing equipment alarm system in processing data delay and untimely response. Figure 1 , which is a module diagram of a distributed alarm system 1 that conforms to a preferred embodiment of the present invention. It can be seen from the figure that a distributed alarm system 1 provided in this embodiment may include: a collection unit 11, a first distribution unit 12, a second distribution unit 13 and a rule engine 14.

[0050] The acquisition unit 11 is used to acquire raw data of the device.

[0051] In this embodiment, the collection unit 11 uses equipment sensors to monitor parameters such as flow, pressure, temperature, etc. during equipment operation and output specific values; the data can be sent to the first distribution unit 12 via MQTT (message publish / subscribe transport protocol) / TCP (transmission control protocol) / UDP (user datagram protocol) / HTTP (application layer protocol) protocol.

[0052] A first allocation unit 12 is used to convert the original data into standard data, and establish a mapping relationship between the device associated with the standard data and the message queue in the message array according to the standard data based on a first preset rule, and the message queue obtains the standard data of the corresponding device according to the mapping relationship;

[0053] It should be noted that the first preset rule in this embodiment adopts a consistent hashing algorithm.

[0054] Furthermore, the first allocating unit 12 may include: a conversion module and a slicing module.

[0055] The conversion module is used to standardize the original data and convert the original data into the standard data, wherein the standard data includes a device number of a device associated with the standard data.

[0056] In this embodiment, the conversion module obtains the data sent by the collection unit 11, and after being processed in a standardized form by the IoT adapter, it is connected to the sharding module. The standardized data structure includes the unique number of the device, the time when the data is generated, and the specific values ​​of all monitoring points of the device. The data structure can be in JSON format, and the standardized data is sent to the sharding module for data sharding.

[0057] The sharding module is used to obtain the mapping relationship between the device number and the message queue number by adopting a consistent hashing algorithm, so as to evenly distribute the device standard data to the message queue.

[0058] In this embodiment, the sharding module performs consistent hashing according to the unique device number, where the consistent hashing algorithm uses Google's Maglev hashing algorithm. The hash function uses the CRC16 Xmodem algorithm, which is adopted by redis and has the advantages of fast speed, low conflict rate, uniform distribution, and hardware friendliness. The message queue number in the specified message array is obtained through consistent hashing to complete the mapping between the device unique number and the message queue ( Figure 2 ). Take the example of a message array including 16 (i.e. 16 slots Solt) message queues, and evenly distribute the devices to the 16 message queues.

[0059] In a preferred embodiment, the message queue processes the received standard data in a first-in-first-out manner, writes the standard data sequentially into a log file, and backs up the received standard data.

[0060] In this embodiment, the message array contains N message queues (where N is a positive integer) when initialized. After the consistent hashing of the sharding module, the standard data is evenly distributed to these N message queues. The message array decouples the collection unit 11 and the rule engine 14 to smooth the data flow. The first-in-first-out of the message queue ensures the order of events. The data of each message queue in the message array is sequentially appended to the log file, and the log file is persisted to a fixed folder on the disk. At the same time, a copy of the data is written to other nodes in the cluster to provide distributed high availability. When a single point of failure occurs, the copy data is read to ensure that the standard data is not lost. When the message queue is processed by the rule engine 14, the rule engine 14 will submit the current processing progress offset offset. After the message array receives the offset offset submitted by the rule engine 14, it records the processing progress of the rule engine 14 in the log.

[0061] A second allocation unit 13, configured to establish a mapping relationship between each of the message queues in the message array and a preset number of rule engines 14 based on a second preset rule;

[0062] Furthermore, the second preset rule is:

[0063] When the total number of the message queues in the message array is an integer multiple of the total number of cores of all rule engines 14, based on the number of cores of each of the rule engines 14, the message queues in the message array are evenly distributed to a preset number of rule engines 14 to establish a mapping relationship between each of the message queues and the preset number of rule engines 14.

[0064] When the total number of message queues in the message array is not an integer multiple of the total number of cores of all rule engines 14, after removing the evenly distributed message queues, the remaining message queues will be sequentially distributed to the rule engines 14 one by one until all message queues are distributed. If the number of rule engines 14 changes, increases or decreases, all data processing will be suspended before reallocation.

[0065] In this embodiment, the second allocation unit 13 allocates the specified message queue in the message array to the corresponding alarm calculation rule engine 14 (see Figure 3), the rule engine 14 (referred to as the consumer) carries its own CPU core number or main frequency information each time it registers, and the second allocation unit 13 controls the partition allocation according to the CPU processing capacity feedback of the consumer. For example, consumer_1_CPU8 means that consumer 1 has an 8-core CPU, consumer_2_CPU16 means that consumer 2 has a 16-core CPU, and consumer_3_CPU8 means that consumer 3 has an 8-core CPU. There are 3 consumers, and the CPU ratio is 1:2:1. If there are 16 message queues, consumer 1 is allocated 4 queues, consumer 2 is allocated 8 queues, and consumer 3 is allocated 4 queues. If it cannot be evenly distributed, the remaining message queues will be allocated to consumers one by one in sequence until all queues are allocated. If the number of consumers changes, increases or decreases, all data processing will be suspended and then reallocated.

[0066] The rule engine 14 is used to obtain the standard data packet in the message queue mapped to the rule engine 14, and assign the standard data in the standard data packet to the corresponding thread for processing according to the third preset rule to identify whether the device associated with the standard data is abnormal.

[0067] In this embodiment, the rule engine 14 (i.e., the consumer) is deployed in multiple instances. Each consumer obtains a batch of standard data packets of devices from the assigned message queue each time. This batch of standard data is arranged in the order of first-in-first-out of the queue and contains data of multiple standardized data structures. In order to process this batch of data with high performance and to ensure the sequence of device events (for example, the door opening and closing events of access control devices need to be sequential), these data are grouped according to the unique device number, and then each group of data is assigned to a thread for processing. The number of threads is the number of threads available for the CPU (see Figure 4 As shown, the number of threads is consistent with the number of CPU threads). This thread is responsible for submitting the standard data of the device to the rule management module for rule matching and execution. The consumer groups standard data packets 1, 2, and 3. Each group corresponds to a thread, and the thread processes the grouped data.

[0068] It should be noted that the third preset rule described in this embodiment adopts a hash algorithm.

[0069] In a preferred embodiment, the rule engine 14 includes a rule management module;

[0070] The thread is used to send the standard data to the rule management module;

[0071] The rule management module adopts the logic of standard data matching alarm rules. The rule management module is used to create an inverted index data structure according to the alarm rules and the device numbers of the associated standard data, generate matching rules, and identify whether the device associated with the standard data is abnormal based on the matching rules.

[0072] In this embodiment, the rule management module is used to manage and maintain the alarm rule entity, rule matching and execution. The alarm rule entity contains the following data fields: unique rule number, alarm rule name, monitoring point unique number metricCode, alarm level, maximum threshold value, minimum threshold value, time required for alarm, number of abnormalities required for alarm, and percentage of abnormalities required for alarm. Each alarm rule can be associated with multiple devices, connected using the unique number of the alarm rule. Provides creation, modification, deletion, and query interfaces to manage alarm rules in the database. In order to achieve data-driven, device data matches the logic of alarm rules, creates an inverted index data structure for alarm rules and their associated devices, and provides high-performance matching of monitoring point metricCode rules. The specific steps are as follows:

[0073] Create an empty mapping container of device number and device rule list, obtain all alarm rules by monitoring point number, and each alarm rule corresponds to an alarm rule number;

[0074] The first layer loop processes all the acquired alarm rules and acquires all the associated device numbers according to the alarm rule numbers.

[0075] The second layer loop processes all devices and adds the alarm rules of the current first layer loop to the device rule list;

[0076] When the two-layer loop is completed, the device rule list of the current monitoring point number can be obtained;

[0077] The monitoring points in the device monitoring point data are used to obtain the device rule list, and the device unique number is used to obtain the rule list, that is, the matching rule. Thus, the rule matching with high performance and time complexity is completed.

[0078] Specifically, after the rule matching is completed, the rules are executed in turn according to the rule definition. Determine whether the data of the monitoring point is within the preset threshold interval (that is, the interval between the minimum and maximum threshold values). If not, it means that the monitoring point is abnormal. When the matching rule does not have a preset alarm time and alarm abnormality percentage, it means that no noise reduction processing is required to directly generate an alarm signal; if the matching rule sets an abnormal percentage noise reduction rule (that is, a preset alarm time and alarm abnormality percentage is configured), since the configured fixed time belongs to a window time, the data-driven sliding sliding window algorithm (see Figure 5 ), generate an alarm signal;

[0079] Here are the steps:

[0080] (1) Each alarm key ([rule number + device number + monitoring point number metricCode] + [message queue number of the current device monitoring point data]) has a sliding window object. Each window is initialized according to the window time set by the rule, and the window time is divided into N grids. The number of grids is determined by the business. The more grids there are, the finer the sliding granularity. Each grid has a counter, for a total of N counters. At the same time, the sliding window object also includes the grid number of the current count and the device measurement point event time that last triggered the sliding.

[0081] (2) Each time the device measurement point data is received, the current device measurement point event time is compared with the last trigger sliding time. If it is less than or equal to the time per grid, the current grid is used directly. If the data is abnormal, the counter of this grid is increased by one. If the comparison time exceeds the time size of each grid, the window slides forward one grid and the new grid is counted again. If the data is late or crosses, the window can slide forward or backward. The number of slides is the difference between the device measurement point time and the last sliding time divided by the time per grid. In this way, the counter abnormality problem caused by data disorder or lateness can be solved.

[0082] (3) Each time the device measurement point data is received, the counter in the window grid performs an atomic accumulation operation, and determines whether the window needs to be slid based on the time. At this point, the abnormality count in the window is completed, that is, the total abnormality count e is equal to the sum of the counters of all grids.

[0083] (4) In order to meet the fixed time exception percentage set by the rule, it is necessary to count the generation time of the device monitoring point data received for the last 10 times, obtain the time interval t1, t2, t3...t10 for each data generation, take the average value avg(t1...t10), and cache it in the memory. The cache key is the device unique number, and the cache value is the average time interval n=avg(t1...t10) seconds for uploading the device monitoring point data.

[0084] (5) Based on the time m minutes required for the noise reduction alarm setting and the average time interval n seconds for uploading the equipment monitoring point data, it is calculated that there are a total of [(m*60) / n] equipment monitoring point data within the m-minute period.

[0085] (6) Then the abnormal percentage p = e / ((m*60) / n). When the abnormal percentage p is greater than or equal to the set percentage, an alarm is generated.

[0086] This embodiment can also provide a cache module to provide high-performance memory reading and writing, which is particularly useful for timely alarms. However, when the server crashes, the cache data will be lost, causing false alarms and missed alarms. The distributed high availability of the alarm system requires the management of the cache and provides a copy mechanism. The cache module provides special high-availability logic for the equipment alarm business scenario. The implementation method includes: the key of the alarm cache is composed of [rule number + equipment number + monitoring point number metricCode] + [message queue number of the current equipment monitoring point data]. The cache provides external query addition and deletion interfaces, where addition and deletion belong to the write interface, and query belongs to the read interface. The write interface provides data dual-write function, writes data to the centralized database, and performs persistent copy backup.

[0087] The specific steps for optimizing cache double writing for device alarm scenarios are as follows:

[0088] (1) First define the alarm cache addition and deletion as two commands, the addition command is ADD and the deletion command is DEL.

[0089] (2) Use a circular queue as a buffer. The circular queue consumer obtains the cached data (key = value) and temporarily places it in a mapping table Map. The key of Map is [rule unique number + device unique number + monitoring point unique number metricCode] + [message queue number of the current device monitoring point data], and the value of Map is the cache value containing ADD and DEL commands.

[0090] (3) Since each atomic increment is performed by first checking and then updating the value before writing it, the key of the alarm cache is transient and does not need to be recorded in history. Therefore, the previous value can be directly overwritten according to the key to compress the data that is double-written to the centralized database and reduce data transmission.

[0091] (4) Finally, every minute or when the number of cached keys reaches the Max value, or when the service goes down, a hook function is created to connect to the centralized database to back up the cached copy.

[0092] Processing logic after the alarm calculation consumption processing node crashes:

[0093] (1) Synchronize the remaining cache to the replica in case of a crash.

[0094] (2) The number of consumers changes, triggering the reallocation of the partition allocation module.

[0095] (3) All other consumers clear their local caches and suspend consumption until the partition allocation module assigns them a suitable message queue.

[0096] (4) After being assigned to the message queue, it uses the message queue number to obtain its own state cache from the cache copy and stores it in the local memory.

[0097] (5) Continue to obtain alarm monitoring data packets, match rules, and execute alarm rules.

[0098] The alarm push module is a public module. Channels such as SMS, WeChat public account, email, and APP can be directly connected to commercial channel suppliers to implement message push logic.

[0099] So far, the core module implementation logic and key steps of the high-performance data-driven distributed equipment alarm software system have been completed ( Figure 2 ). Taking three rule engines, namely rule engine 1, rule engine 2 and rule engine 3 as an example, the message queues 1-5 are allocated to rule engine 1, message queues 6-10 are allocated to rule engine 2, and message queues 11-16 are allocated to rule engine 3 through the second allocation unit. Rule engines 1-3 can read and write local status data respectively. The central replica database includes 16 slots, which respectively store message queue status data and correspond one-to-one with the message queue numbers, thereby realizing functions such as data recovery and incremental synchronization.

[0100] The distributed alarm system 1 of this embodiment collects the original data of the device through the collection unit 11; uses the first allocation unit 12 to establish the mapping relationship between the device and the message queue based on the first preset rule, so as to evenly distribute the data of the device to the message queue, with fast speed and low conflict rate; uses the second allocation unit 13 to establish the mapping relationship between each message queue in the message array and a preset number of rule engines 14 based on the second preset rule; obtains the standard data packet in the message queue mapped with the rule engine 14 through the rule engine 14, and distributes the standard data in the standard data packet to the corresponding thread for processing according to the third preset rule to identify whether the device associated with the standard data is abnormal. The first allocation unit 12 and the second allocation unit 13 decouple the device from the rule engine 14 to smooth the data flow. When the rule engine 14 sends a transformation, the second allocation unit 13 can reallocate the rule engine 14, thereby improving the real-time and accuracy of the alarm.

[0101] Distributed alarm system 1 has the ability to handle complex equipment alarm rules with high performance: while fully utilizing the advantages of multi-core CPUs, it can also solve the problem of alarm sequence of equipment monitoring points. At the same time, when facing fluctuations in equipment data, the alarm is reduced in noise, and an alarm will only be triggered when a certain percentage of abnormalities is reached within a fixed time. In the actual operation and production of a water utility plant, when facing the valve switch alarm, the sequence of switches can be handled well. At the same time, when the pressure and flow data fluctuate abnormally, the alarm rule is used to configure a certain percentage of abnormalities within a fixed time to make a correct alarm and reduce unnecessary alarm push.

[0102] There are two main reasons why the existing equipment alarm system cannot handle complex equipment alarm rules well: (1) When the alarm frequency is too high, the alarm noise reduction convergence process is not perfect, and it cannot be configured to alarm only when the set percentage of abnormalities is reached within a fixed time or the number of non-continuous abnormalities is reached. (2) When using multi-core CPU multi-threaded concurrent computing to improve performance, it cannot handle the problem of the sequence of events at the equipment monitoring point well. In addition, some existing alarm systems are driven by rules or models. The alarm system driven by the rule model drives data collection and processing through a pre-defined rule model, triggers the rule model in a timed cycle, and then obtains data. In some scenarios with high real-time requirements, data lag may occur, resulting in alarm delays, and problems cannot be discovered and solved in time. The distributed alarm system 1 of this embodiment has the advantages of strong real-time performance and high accuracy. By adopting a data-driven method, the system can receive and process the operation data of the equipment in real time, and immediately match it with the preset alarm rules, thereby achieving millisecond-level alarm response. Compared with the traditional rule-driven system, the alarm response speed of the present invention is greatly improved. In actual application tests, when faced with millions of sensor data from a water plant, the average delay time for the system to successfully detect and issue an alarm can be within 1 second, greatly improving the real-time nature of the alarm. At the same time, because real-time data drives reduce the problem of rule lag, the system's alarm accuracy is greatly improved, effectively reducing false alarms and missed alarms.

[0103] The distributed alarm system 1 is also scalable and maintainable: based on the distributed architecture design, the system can be easily expanded when the number and types of equipment increase, avoiding the performance bottleneck problems common in traditional centralized systems. It can also achieve high availability at the software level, and will not miss alarms after the computing node goes down, greatly simplifying the maintenance of the system. In a test of a water group, the system was able to smoothly handle data collection and alarm tasks for more than 5,000 devices and millions of sensor monitoring points, greatly reducing the company's operation and maintenance costs and improving the safety of the company's production.

[0104] The distributed equipment alarm system of this embodiment can efficiently process massive amounts of data, and eliminates the risk of single point failures through special designs such as copy redundancy and load balancing for equipment alarm business scenarios. Even if a node fails, the rest of the system can still operate normally, avoiding untimely equipment alarms, missed alarms, or false alarms. The accuracy and real-time nature of the equipment alarms in the distributed alarm system 1 reduces the detection time of equipment failures, thereby reducing the downtime caused by equipment failures. The distributed high-availability characteristics of the system reduce missed alarms caused by service downtime, improve the safety of industrial production, and bring significant social benefits.

[0105] The present invention also provides a distributed alarm method. Figure 6 As shown, the following steps are included:

[0106] S1. Collect the original data of the equipment;

[0107] In this embodiment, device sensors are used to monitor parameters such as flow, pressure, temperature, etc. during device operation and output specific values; data can be output through MQTT (message publish / subscribe transport protocol) / TCP (transmission control protocol) / UDP (user datagram protocol) / HTTP (application layer protocol) protocols.

[0108] S2. converting the original data into standard data, establishing a mapping relationship between the device associated with the standard data and the message queue in the message array according to the standard data based on a first preset rule, and obtaining the standard data of the corresponding device according to the mapping relationship;

[0109] It should be noted that the first preset rule in this embodiment adopts a consistent hashing algorithm.

[0110] Further, step S2 includes:

[0111] S21. Standardize the original data and convert the original data into the standard data, wherein the standard data includes a device number of a device associated with the standard data.

[0112] In this embodiment, an IoT adapter may be used to standardize the original data. The standardized data structure includes the unique device number, the data generation time, and the specific values ​​of all monitoring points of the device. The data structure may be in JSON format.

[0113] S22. A consistent hashing algorithm is used to obtain a mapping relationship between the device number and the message queue number, so as to evenly distribute the device standard data to the message queue.

[0114] In this embodiment, consistent hashing is performed according to the unique device number, where the consistent hashing algorithm uses Google's Maglev hashing algorithm. The hash function uses the CRC16 Xmodem algorithm, which is adopted by redis and has the advantages of fast speed, low conflict rate, uniform distribution, and hardware friendliness. The message queue number in the specified message array is obtained through consistent hashing to complete the mapping between the device unique number and the message queue ( Figure 2 ).

[0115] In a preferred embodiment, the message queue processes the received standard data in a first-in-first-out manner, writes the standard data sequentially into a log file, and backs up the received standard data.

[0116] In this embodiment, the message array contains N message queues (where N is a positive integer) when initialized. After the consistent hashing of the sharding module, the standard data is evenly distributed to these N message queues. The message array decouples the device sensor and the rule engine to smooth the data flow. The first-in-first-out of the message queue ensures the order of events. The data of each message queue in the message array is sequentially appended to the log file, and the log file is persisted to a fixed folder on the disk. At the same time, a copy of the data is written to other nodes in the cluster to provide distributed high availability. When a single point of failure occurs, the copy data is read to ensure that the standard data is not lost. When the message queue is processed by the rule engine, the rule engine will submit the current processing progress offset. After the message array receives the offset submitted by the rule engine, it records the processing progress of the rule engine in the log.

[0117] S3. A mapping relationship between each of the message queues in the message array and a preset number of rule engines is established based on a second preset rule;

[0118] Furthermore, the second preset rule is:

[0119] When the total number of the message queues in the message array is an integer multiple of the total number of cores of all rule engines, based on the number of cores of each rule engine, the message queues in the message array are evenly distributed to a preset number of rule engines to establish a mapping relationship between each message queue and a preset number of rule engines.

[0120] When the total number of message queues in the message array is not an integer multiple of the total number of kernels of all rule engines, after removing the evenly distributed message queues, the remaining message queues will be distributed to the rule engines one by one in sequence until all message queues are distributed. If the number of rule engines changes, increases or decreases, all data processing will be suspended before reallocation.

[0121] In this embodiment, the specified message queue in the message array is assigned to the corresponding alarm calculation rule engine (see Figure 3 ), the rule engine (referred to as the consumer) carries its own CPU core number or main frequency information each time it registers, and controls the partition allocation based on the consumer's CPU processing power feedback. For example, consumer_1_CPU8 means that consumer 1 has an 8-core CPU, consumer_2_CPU16 means that consumer 2 has a 16-core CPU, and consumer_3_CPU8 means that consumer 3 has an 8-core CPU. There are three consumers, and the CPU ratio is 1:2:1. If there are 16 message queues, consumer 1 is allocated 4 queues, consumer 2 is allocated 8 queues, and consumer 3 is allocated 4 queues. If it cannot be evenly distributed, the remaining message queues will be allocated to consumers one by one in sequence until all queues are allocated. If the number of consumers changes, increases or decreases, all data processing will be suspended and then reallocated.

[0122] S4. Obtain a standard data packet in the message queue mapped to the rule engine, and assign the standard data in the standard data packet to a corresponding thread for processing according to a third preset rule to identify whether a device associated with the standard data is abnormal.

[0123] In this embodiment, the rule engine can be referred to as a consumer, and multiple instances of consumers are deployed. Each consumer obtains a batch of standard data packets of devices from the assigned message queue each time. This batch of standard data is arranged in the order of first-in-first-out of the queue, and contains data of multiple standardized data structures. In order to process this batch of data with high performance and to ensure the sequentiality of device events (for example: the door opening and closing events of access control devices need to be sequential), these data are grouped according to the unique device number, and then each group of data is handed over to a thread for processing. The number of threads is the number of threads available for the CPU. The thread is responsible for submitting the standard data of the device to the rule management module for rule matching and execution.

[0124] It should be noted that the third preset rule described in this embodiment adopts a hash algorithm.

[0125] In a preferred embodiment, the rule engine includes a rule management module;

[0126] The thread is used to send the standard data to the rule management module;

[0127] The rule management module adopts the logic of standard data matching alarm rules. The rule management module is used to create an inverted index data structure according to the alarm rules and the device numbers of the associated standard data, generate matching rules, and identify whether the device associated with the standard data is abnormal based on the matching rules.

[0128] In this embodiment, the rule management module is used to manage and maintain the alarm rule entity, rule matching and execution. The alarm rule entity contains the following data fields: unique rule number, alarm rule name, monitoring point unique number metricCode, alarm level, maximum threshold value, minimum threshold value, time required for alarm, number of abnormalities required for alarm, and percentage of abnormalities required for alarm. Each alarm rule can be associated with multiple devices, connected using the unique number of the alarm rule. Provides creation, modification, deletion, and query interfaces to manage alarm rules in the database. In order to achieve data-driven, device data matches the logic of alarm rules, creates an inverted index data structure for alarm rules and their associated devices, and provides high-performance matching of monitoring point metricCode rules. The specific steps are as follows:

[0129] Create an empty mapping container of device number and device rule list, obtain all alarm rules by monitoring point number, and each alarm rule corresponds to an alarm rule number;

[0130] The first layer loop processes all the acquired alarm rules and acquires all the associated device numbers according to the alarm rule numbers.

[0131] The second layer loop processes all devices and adds the alarm rules of the current first layer loop to the device rule list;

[0132] When the two-layer loop is completed, the device rule list of the current monitoring point number can be obtained;

[0133] The monitoring points in the device monitoring point data are used to obtain the device rule list, and the device unique number is used to obtain the rule list, that is, the matching rule. Thus, the rule matching with high performance and time complexity is completed.

[0134] Specifically, after the rule matching is completed, the rules are executed in turn according to the rule definition. Determine whether the data of the monitoring point is within the preset threshold interval (that is, the interval between the minimum and maximum threshold values). If not, it means that the monitoring point is abnormal. When the matching rule does not have a preset alarm time and alarm abnormality percentage, it means that no noise reduction processing is required to directly generate an alarm signal; if the matching rule sets an abnormal percentage noise reduction rule (that is, a preset alarm time and alarm abnormality percentage is configured), since the configured fixed time belongs to a window time, the data-driven sliding sliding window algorithm (see Figure 5 ), generate an alarm signal;

[0135] Here are the steps:

[0136] (1) Each alarm key ([rule number + device number + monitoring point number metricCode] + [message queue number of the current device monitoring point data]) has a sliding window object. Each window is initialized according to the window time set by the rule, and the window time is divided into N grids. The number of grids is determined by the business. The more grids there are, the finer the sliding granularity. Each grid has a counter, for a total of N counters. At the same time, the sliding window object also includes the grid number of the current count and the device measurement point event time that last triggered the sliding.

[0137] (2) Each time the device measurement point data is received, the current device measurement point event time is compared with the last trigger sliding time. If it is less than or equal to the time per grid, the current grid is used directly. If the data is abnormal, the counter of this grid is increased by one. If the comparison time exceeds the time size of each grid, the window slides forward one grid and the new grid is counted again. If the data is late or crosses, the window can slide forward or backward. The number of slides is the difference between the device measurement point time and the last sliding time divided by the time per grid. In this way, the counter abnormality problem caused by data disorder or lateness can be solved.

[0138] (3) Each time the device measurement point data is received, the counter in the window grid performs an atomic accumulation operation, and determines whether the window needs to be slid based on the time. At this point, the abnormality count in the window is completed, that is, the total abnormality count e is equal to the sum of the counters of all grids.

[0139] (4) In order to meet the fixed time exception percentage set by the rule, it is necessary to count the generation time of the device monitoring point data received for the last 10 times, obtain the time interval t1, t2, t3...t10 for each data generation, take the average value avg(t1...t10), and cache it in the memory. The cache key is the device unique number, and the cache value is the average time interval n=avg(t1...t10) seconds for uploading the device monitoring point data.

[0140] (5) Based on the time m minutes required for the noise reduction alarm setting and the average time interval n seconds for uploading the equipment monitoring point data, it is calculated that there are a total of [(m*60) / n] equipment monitoring point data within the m-minute period.

[0141] (6) Then the abnormal percentage p = e / ((m*60) / n). When the abnormal percentage p is greater than or equal to the set percentage, an alarm is generated.

[0142] This embodiment can also provide a cache module to provide high-performance memory reading and writing, which is particularly useful for timely alarms. However, when the server crashes, the cache data will be lost, causing false alarms and missed alarms. The distributed high availability of the alarm system requires the management of the cache and provides a copy mechanism. The cache module provides special high-availability logic for the equipment alarm business scenario. The implementation method includes: the key of the alarm cache is composed of [rule number + equipment number + monitoring point number metricCode] + [message queue number of the current equipment monitoring point data]. The cache provides external query addition and deletion interfaces, where addition and deletion belong to the write interface, and query belongs to the read interface. The write interface provides data dual-write function, writes data to the centralized database, and performs persistent copy backup.

[0143] The specific steps for optimizing cache double writing for device alarm scenarios are as follows:

[0144] (1) First define the alarm cache addition and deletion as two commands, the addition command is ADD and the deletion command is DEL.

[0145] (2) Use a circular queue as a buffer. The circular queue consumer obtains the cached data (key = value) and temporarily places it in a mapping table Map. The key of Map is [rule unique number + device unique number + monitoring point unique number metricCode] + [message queue number of the current device monitoring point data], and the value of Map is the cache value containing ADD and DEL commands.

[0146] (3) Since each atomic increment is performed by first checking and then updating the value before writing it, the key of the alarm cache is transient and does not need to be recorded in history. Therefore, the previous value can be directly overwritten according to the key to compress the data that is double-written to the centralized database and reduce data transmission.

[0147] (4) Finally, every minute or when the number of cached keys reaches the Max value, or when the service goes down, a hook function is created to connect to the centralized database to back up the cached copy.

[0148] Processing logic after the alarm calculation consumption processing node crashes:

[0149] (1) Synchronize the remaining cache to the replica in case of a crash.

[0150] (2) The number of consumers changes, triggering the reallocation of the partition allocation module.

[0151] (3) All other consumers clear their local caches and suspend consumption until the partition allocation module assigns them a suitable message queue.

[0152] (4) After being assigned to the message queue, it uses the message queue number to obtain its own state cache from the cache copy and stores it in the local memory.

[0153] (5) Continue to obtain alarm monitoring data packets, match rules, and execute alarm rules.

[0154] The alarm push module is a public module. Channels such as SMS, WeChat public account, email, and APP can be directly connected to commercial channel suppliers to implement message push logic.

[0155] So far, the core module implementation logic and key steps of the high-performance data-driven distributed equipment alarm software system have been completed ( Figure 2 ). Taking three rule engines, namely rule engine 1, rule engine 2 and rule engine 3 as an example, the message queues 1-5 are allocated to rule engine 1, message queues 6-10 are allocated to rule engine 2, and message queues 11-16 are allocated to rule engine 3 through the second allocation unit. Rule engines 1-3 can read and write local status data respectively. The central replica database includes 16 slots, which respectively store message queue status data and correspond one-to-one with the message queue numbers, thereby realizing functions such as data recovery and incremental synchronization.

[0156] In this embodiment, the distributed alarm method establishes a mapping relationship between the device and the message queue based on the first preset rule, so that the data of the device is evenly distributed to the message queue, with fast speed and low conflict rate; based on the second preset rule, a mapping relationship between each message queue in the message array and a preset number of rule engines is established; the standard data packet in the message queue mapped to the rule engine is obtained through the rule engine, and the standard data in the standard data packet is allocated to the corresponding thread for processing according to the third preset rule to identify whether the device associated with the standard data is abnormal. The device is decoupled from the rule engine to smooth the data flow. When the rule engine sends a transformation, the rule engine can be reallocated, which improves the real-time and accuracy of the alarm. Taking the configuration of the alarm rule binding device to generate an alarm (the mid-alarm computing node suddenly fails and crashes) as an example, the process of handling faults by the distributed alarm method is as follows:

[0157] Define a pressure anomaly alarm rule with noise reduction (for example, if an abnormal value of less than 800MPa or greater than 1200MPa appears at the pressure monitoring point of a certain device in a water plant, and the abnormal value accounts for 70% within 1 hour, a level 1 alarm is generated). After configuring the alarm rule, associate and bind this rule to all devices with pressure monitoring points in the water plant.

[0158] The device sensor generates device pressure monitoring point data every 2 seconds, which is processed into standard data by the adapter and sent to the message queue through consistent hash routing. The device monitoring point data in the message queue is distributed to the alarm calculation consumer by the partition distribution module (see Figure 3 ), the alarm calculation consumer processes the pressure monitoring point data in sequence according to the equipment number (see Figure 4 ), match the equipment pressure monitoring point data to the defined pressure abnormality alarm rules (see Figure 6 ), execute the rule calculation to determine whether the pressure data is abnormal. When the data of the monitoring point does not reach the amount of 1 hour, save the abnormal count to the local cache and write it to the replica at the same time. At this time, the computing node suddenly crashes. Reallocate the data in the message queue to the healthy rule engine ( Figure 3 ). And cache data recovery, each rule engine pulls and restores the cache data belonging to its own queue according to the queue number ( Figure 2 ).

[0159] Receive the equipment pressure monitoring point data generated by the equipment sensor again, repeat the above operation, if the current equipment monitoring point generation time is greater than or equal to 1 hour from the first recording time, and the abnormal percentage exceeds 70%, a first-level pressure abnormality alarm is generated, the cache is cleared, and the next alarm rule calculation cycle is restarted.

[0160] After an alarm is generated, the operator's SMS platform is connected to push a level 1 alarm that a certain device has experienced abnormal pressure. After the user receives the alarm SMS, the alarm processing process is immediately executed.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A distributed alarm system, characterized in that: include: A collection unit, used for collecting raw data of the equipment; A first allocation unit is used to convert the original data into standard data, and establish a mapping relationship between a device associated with the standard data and a message queue in a message array according to the standard data based on a first preset rule, and the message queue obtains the standard data of the corresponding device according to the mapping relationship; A second allocation unit, configured to establish a mapping relationship between each of the message queues in the message array and a preset number of rule engines based on a second preset rule; The rule engine is used to obtain the standard data packet in the message queue mapped to the rule engine, and allocate the standard data in the standard data packet to the corresponding thread for processing according to a third preset rule to identify whether the device associated with the standard data is abnormal.

2. The distributed alarm system according to claim 1, characterized in that: The first preset rule adopts a consistent hashing algorithm.

3. The distributed alarm system according to claim 2, characterized in that: The first allocation unit comprises: A conversion module, used for standardizing the original data, converting the original data into the standard data, wherein the standard data includes a device number of a device associated with the standard data; The sharding module is used to obtain the mapping relationship between the device number and the message queue number by adopting a consistent hashing algorithm.

4. The distributed alarm system according to claim 1, characterized in that: The message queue processes the received standard data in a first-in-first-out manner, writes the standard data into a log file in sequence, and backs up the received standard data.

5. The distributed alarm system according to claim 1, characterized in that: The second preset rule is: Based on the number of cores of each of the rule engines, the message queues in the message array are evenly distributed to a preset number of rule engines, so as to establish a mapping relationship between each of the message queues and the preset number of rule engines.

6. The distributed alarm system according to claim 1, characterized in that: The third preset rule adopts a hash algorithm.

7. The distributed alarm system according to claim 3, characterized in that: The rule engine includes a rule management module; The thread is used to send the standard data to the rule management module; The rule management module is used to create an inverted index data structure according to the alarm rule and the device number of the associated standard data, generate a matching rule, and identify whether the device associated with the standard data is abnormal based on the matching rule.

8. The distributed alarm system according to claim 7, characterized in that: The rule management module is used to create an inverted index data structure according to the alarm rule and the device number of the associated standard data, and generate a matching rule, including: Create an empty mapping container of device number and device rule list, obtain all alarm rules by monitoring point number, and each alarm rule corresponds to an alarm rule number; The first layer loop processes all the acquired alarm rules and acquires all the associated device numbers according to the alarm rule numbers; The second layer loop processes all devices and adds the alarm rules of the current first layer loop to the device rule list; When the two-layer loop is completed, the device rule list of the current monitoring point number can be obtained; The monitoring points in the device monitoring point data are used to obtain a device rule list, and the device unique number is used to obtain a rule list, that is, the matching rule.

9. The distributed alarm system according to claim 8, characterized in that: The rule management module identifies whether the device associated with the standard data is abnormal based on the matching rule, including: Determine whether the data of the monitoring point is within the preset threshold range. If not, it means that the monitoring point is abnormal. If the matching rule does not have a preset alarm time and alarm abnormality percentage, an alarm signal is generated; if the matching rule sets an abnormal percentage noise reduction rule, a data-driven sliding window algorithm is used to generate an alarm signal; If the data of the monitoring point is within the preset threshold range, it means that the monitoring point is normal.

10. A distributed alarm method, characterized in that: The following steps are involved: S1. Collect the original data of the equipment; S2. Converting the original data into standard data, establishing a mapping relationship between the device associated with the standard data and the message queue in the message array based on the standard data based on the first preset rule, and obtaining the standard data of the corresponding device according to the mapping relationship; S3. A mapping relationship between each of the message queues in the message array and a preset number of rule engines is established based on a second preset rule; S4. Obtain a standard data packet in the message queue mapped to the rule engine, and assign the standard data in the standard data packet to a corresponding thread for processing according to a third preset rule to identify whether a device associated with the standard data is abnormal.

Citation Information

Cited By

  • Internet of Things event center distributed alarm processing method based on Redis Stream

    CN121151183A