Rule updating method based on distributed real-time calculation and storage medium

By adopting a distributed real-time computing framework in real-time data processing, dynamically updating the rule set of the rule engine, the calculation bottleneck and rule static problems in traditional methods are solved, and efficient and flexible real-time data processing and rule matching are achieved.

CN120216547APending Publication Date: 2025-06-27FUJIAN TIANQUAN EDUCATION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510134710.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When traditional real-time data processing methods deal with real-time data flows with high concurrency, high throughput, and low latency, they are prone to computing framework bottlenecks and difficult to scale, and the static nature of the rule engine leads to poor system flexibility and adaptability.

Method used

The rule update method based on distributed real-time computing is adopted to obtain data events through real-time data flow, match rules based on the timestamp of the event and process them, build a view of the rule set, and dynamically update the rule set of the rule engine.

Benefits of technology

It realizes flexible adjustment of rules, ensures the effectiveness and accuracy of rules, improves the system's ability to adapt to dynamic data environments, and avoids the problem of lag in traditional rules.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216547A_ABST
    Figure CN120216547A_ABST
Patent Text Reader

Abstract

The invention discloses a rule updating method based on distributed real-time computing and a storage medium, and the method comprises the steps: obtaining a data event according to a real-time data stream; according to the timestamps of the events, rules corresponding to the events are obtained through matching from a preset rule set, the events are processed according to the rules corresponding to the events, processing results corresponding to the events are obtained, and the events comprise data events; and according to the processing result, constructing a rule set view of the rule set, and updating the rule set of the rule engine. According to the invention, the rule can be flexibly adjusted according to the change of the data stream, and the validity and accuracy of the rule are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of real-time data processing, and particularly to a rule update method and a storage medium based on distributed real-time computing. Background Art

[0002] With the development of big data technology and the continuous expansion of Internet applications, the real-time processing and analysis of massive data have become the core needs of enterprises and organizations in decision-making, monitoring, and optimizing business processes. Especially in the fields of financial monitoring, intelligent manufacturing, Internet of Things, intelligent transportation, etc., the analysis and processing of real-time data streams can help the system capture events in a timely manner and make responses. Therefore, the real-time computing technology based on event streams has become an important research direction. Event stream computing refers to capturing and processing real-time events and data streams, extracting valuable information from them, and making immediate responses. This requires the computing framework to have powerful real-time processing capabilities, flexible rule matching functions, and high scalability to cope with the dynamically changing business environment and the increasing data volume.

[0003] Traditional real-time data processing methods generally adopt a centralized computing architecture and a batch processing mode, and usually use traditional database management systems (DBMS), ETL (Extract, Transform, Load) processes, as well as simple triggers and rule engines for data processing and rule matching. However, this traditional method faces multiple challenges, especially when dealing with real-time data streams with high concurrency, high throughput, and low latency. The centralized architecture is prone to becoming a bottleneck and is difficult to efficiently scale to meet the large-scale data processing requirements. At the same time, traditional rule engines and data storage systems are usually static and cannot dynamically adjust calculations and rules according to the changes in real-time data streams, resulting in poor flexibility and adaptability of the system. In addition, there may be a long delay in the data processing process, affecting the effect of real-time decision-making.

[0004] In traditional methods, computing tasks are generally processed through a single node or a single service, and the scheduling and resource allocation of tasks usually rely on centralized control logic. This limits the load balancing, fault tolerance, and scalability of the system. In addition, due to the lack of flexibility of the rule engine, it cannot be updated and adjusted in real time according to the changes in the event stream, and the rules often lag behind the data stream, resulting in incorrect decisions or missing key information. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a rule update method and a storage medium based on distributed real-time computing, so that the rules can be flexibly adjusted according to the changes in the data stream, and the effectiveness and accuracy of the rules can be guaranteed.

[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows: A rule update method based on distributed real-time computing, comprising: Obtaining data events according to the real-time data stream; According to the timestamps of the events, respectively matching the rules corresponding to each event from a preset rule set, and respectively processing each event according to the rules corresponding to each event to obtain the processing results corresponding to each event, where the events include data events; Constructing a rule set view of the rule set according to the processing results, and updating the rule set of the rule engine.

[0007] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above-mentioned method is implemented.

[0008] The beneficial effects of the present invention are as follows: Through the real-time processing of the data stream, the rules can be flexibly adjusted according to the changes in the data stream, ensuring the effectiveness and accuracy of the rules, significantly improving the adaptability of the system to the dynamic data environment, and avoiding the problem of traditional rule lag. Description of the Drawings

[0009] Figure 1 It is a flowchart of a rule update method based on distributed real-time computing of the present invention; Figure 2 It is a flowchart of the method in the first embodiment of the present invention; Figure 3 It is a schematic structural diagram of the real-time computing framework in the first embodiment of the present invention.

[0010] Label Description: 100, distributed real-time computing framework; 200, distributed file storage system; 101, message queue; 102, management unit; 103, computing unit; 1031, data processing engine; 1032, distributed rule engine; 201, storage node. Detailed Embodiments

[0011] To describe the technical content, the achieved objectives and the effects of the present invention in detail, the following is described in conjunction with the embodiments and the accompanying drawings.

[0012] Please refer to Figure 1 , a rule update method based on distributed real-time computing, comprising: Obtaining data events according to the real-time data stream; According to the timestamps of each event, match the rules corresponding to each event from a preset rule set respectively, and process each event according to the rules corresponding to each event to obtain the processing results corresponding to each event, where the events include data events; Construct a rule set view of the rule set according to the processing results, and update the rule set of the rule engine.

[0013] As can be seen from the above description, the beneficial effect of the present invention is that through the real-time processing of the data stream, the rules can be flexibly adjusted according to the changes in the data stream, ensuring the effectiveness and accuracy of the rules.

[0014] Furthermore, it further includes: Update the rules in the rule set according to the rule update periods of the rules in the rule set respectively.

[0015] As can be seen from the above description, in addition to updating the rules according to the trigger conditions of the events, the rules can also be updated according to the rule update periods, thereby realizing the dynamic update of the rule set to adapt to the changing event data and computing requirements.

[0016] Furthermore, the step of respectively matching the rules corresponding to each event from a preset rule set according to the timestamps of each event, and processing each event according to the rules corresponding to each event to obtain the processing results corresponding to each event includes: Determine the valid time intervals of the rules respectively according to the start time and end time of each rule; If the characteristics of an event meet the matching conditions corresponding to a rule, and the timestamp of the event is within the valid time interval of the rule, then use the rule as the rule corresponding to the event; Process the event according to the execution action corresponding to the rule to obtain the processing result of the event; Wherein, the characteristics of the event include data type, data size, data generation frequency and data source.

[0017] Furthermore, after obtaining the data event according to the real-time data stream, it further includes: Add the data event to an event queue, and determine the data events corresponding to each time window according to the timestamps of each data event and a preset time window length.

[0018] As can be seen from the above description, it is convenient to process the data stream based on the time window later, providing a clear timing basis for event processing.

[0019] Furthermore, the event also includes a rule event; the method further includes: Generate new rule events according to a preset trigger period; If a new real-time data stream arrives in the current time window, a new rule event is generated according to the new real-time data stream and a preset rule template; If no new real-time data stream arrives in the current time window, a new data stream is generated according to the data events and their processing results corresponding to a preset number of previous time windows, and a new rule event is generated according to the new data stream and a preset rule template; Add the new rule event to the event queue.

[0020] Further, the event further includes a rule event; After processing each event according to the rules corresponding to each event to obtain the processing results corresponding to each event, it further includes: Generate a new rule event according to the processing result, and add the new rule event to the event queue.

[0021] As can be seen from the above description, by generating rule events, subsequent processing can be carried out based on these rule events to ensure the efficiency and adaptability of event processing and rule updating in the entire distributed real-time computing process.

[0022] Further, the rules corresponding to the events in the same time window are updated by a parallel processing method.

[0023] As can be seen from the above description, the computing efficiency can be greatly improved by parallel updating.

[0024] Further, before respectively matching the rules corresponding to each event from a preset rule set according to the timestamps of each event, it further includes: If the current time point exceeds the validity period of an event, the event is removed; Or, if the timestamp of an event is not within the current time window, the event is removed.

[0025] As can be seen from the above description, through the aging processing, it can be ensured that historical data does not affect the current calculation result, thereby ensuring the accuracy and effectiveness of the entire system calculation, and avoiding incorrect calculation results and decisions caused by interference from expired data.

[0026] Further, the method is based on a real-time computing framework, the real-time computing framework includes a distributed real-time computing framework and a distributed file storage system, the distributed real-time computing framework includes a message queue, a management unit and at least one computing unit, and the computing unit includes a data processing engine and a distributed rule engine; the distributed file storage system includes at least one storage node; The message queue is used to receive real-time data streams and add the data events in the real-time data streams to the event queue; The management unit is used to allocate computing tasks to computing units through a load balancing algorithm; The computing unit is used to perform rule matching and event processing of events after receiving the computing tasks; The distributed file storage system is used to write event data into corresponding storage nodes according to the timestamps of events and a preset storage policy.

[0027] As can be seen from the above description, the combination of the distributed storage system and the message queue enables large-scale data to be efficiently stored and accessed, and ensures the fault tolerance of the system. In the case of a sharp increase in data traffic or the failure of computing nodes, the system can continue to run and process data, ensuring the high availability of services and the reliability of data. By dynamically scheduling tasks according to the load situation, the bottleneck problem in the centralized architecture is avoided.

[0028] Furthermore, the two engines in the computing unit can be deployed in Docker containers in the form of microservices, so as to achieve service isolation, expansion and high availability. And this deployment method enables the rule engine to run, expand and manage independently, supports high concurrency and horizontal expansion, and improves the flexibility and scalability of the system.

[0029] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.

[0030] Embodiment 1 Please refer to Figures 2 - 3 , Embodiment 1 of the present invention is: a rule update method, which is based on distributed real-time computing and event stream processing technologies and can be widely applied to various scenarios that require efficient processing of large-scale real-time data streams and dynamic rule matching.

[0031] As Figure 2 shown, the method includes the following steps: S1: Generate events and add the events to an event queue.

[0032] The events include data events, which are obtained according to real-time data streams. Specifically, data events are obtained from real-time data streams, and the data events corresponding to each time window are determined according to the timestamps of the data events and a preset time window length.

[0033] Among them, the real-time data stream is a data sequence generated in real time and sorted in chronological order. The time window is a unit for dividing the data stream in the time dimension. Each time window contains a part of the data stream within a certain time period. The system processes the data stream based on the time window, providing a clear time sequence basis for event processing. According to the model of the time stream, a time-window-based calculation and event stream update mechanism is adopted to ensure the efficiency and self-adaptability of event processing and rule update.

[0034] Specifically for the processing flow of the data stream, first, through the initialization of the data stream, the data stream information is loaded into the message queue, and a time node list is created to identify the positions of each time window. The initialization of the time stream arranges the data stream in order and divides it into multiple time windows based on the time dimension, so that events can be processed within each time window. For each time window, the data stream calculation depends on the difference between the previous time window and the current time window to ensure the validity of the event data. That is, from the perspective of dynamically analyzing data changes, valuable information is mined by comparing data in different windows for calculation.

[0035] Events also include rule events, which are generated according to preset triggering methods. That is, rule events are also a special type of event and are also added to the event queue. The purpose is to further process based on these rule events in the future, such as performing rule matching, updating the rule set, etc., to ensure the efficiency and self-adaptability of event processing and rule update in the entire distributed real-time computing process.

[0036] In this embodiment, there are three triggering methods, namely timed triggering, data triggering, and rule triggering, which are specifically as follows: (1) Timed triggering, that is, rule events are generated according to a preset triggering period. The triggering period can be at the second level, minute level, hour level, or day level. According to the triggering period, the triggering time point can be determined. The determination of the timed triggering mechanism can be represented by the following formula: f timer (t)=1(t=T trigger ) Among them, 1( ) represents the indicator function, which returns 1 when the condition is met and 0 when not met. Here, it means that if the current time point t is equal to the triggering time point T trigger , then the generation of rule events is triggered.

[0037] (2) Data triggering, that is, triggering the generation of rule events based on the information of the data source. The data source can include databases, message queues, device gateways, and ETL components, etc. If the data source detects that a new data stream arrives within the current time window, a new rule event is generated. The update period of the data stream and the generation period of the rule event are synchronized. Therefore, when the data stream changes, the corresponding rule event must be updated immediately. In this case, the triggering condition for the arrival of the data stream can be described by the following formula: f data (t)=1(t=T data )·1(T data ∈[T start ,T end ) where T data represents the timestamp of the arrival of the data stream, and [T start ,T end represents the current time window of the data stream. Here, it means that if the data stream arrives within the current time window, the generation of a new rule event is triggered.

[0038] (3) Rule triggering, that is, forming a new data stream according to the calculation result of the data stream in a new time window, and generating a new rule event based on the new data stream. This generation mechanism is particularly important in the process of event stream processing. Especially when there is no new data stream in the data source, the system still needs to calculate a new data stream based on the time window. In this method, if there is no new data stream, a new data stream is generated according to the division of the time window and the data stream. In this process, the processing of the data stream depends on the calculation result of the previous time window and continuously generates a new data stream. It can be represented by the following formula: f dataflow (t)=1(t∈[T start ,T end )·1(newdataflowexisits) where [T start ,T end represents the current time window. Here, it means that if there is a data stream in the current time window, the system generates a rule event and updates the data stream.

[0039] Specifically, when generating a new data stream, operations are first carried out based on the time window as the basic unit according to the calculation results of the data stream. When there is no new data stream in the data source, the system extracts key information from the calculation results of the previous time window, such as the statistical characteristics and change trends of specific data. Based on this information and combined with the rules of the time window, the data is recombined, deduced, or simulated. For example, if the data shows a certain increasing pattern in the previous time window, according to the continuity of the time window, a new data sequence is generated in the new time window according to this pattern, thus forming a new data stream.

[0040] After the new data stream is generated, if there is a data stream within the time window, according to the pre-set rule template, the key data in the data stream, such as specific data values and data change situations, are filled into the corresponding positions in the rule template to generate rule events. For example, if the rule stipulates that a rule event is triggered when a certain type of data value exceeds a specific threshold, then when it is detected in the new data stream that this type of data value exceeds the threshold, the corresponding rule event is generated, clearly recording key information such as the triggering conditions and time.

[0041] Through the above three triggering methods, the generation of the event stream and the update of rule events can be dynamically adjusted according to different triggering conditions. When the time window reaches the triggering condition, rule events are generated through method (1); when there is a new data stream in the data source, event generation is triggered through method (2); when there is no new data stream, a new data stream is calculated through method (3) and events are generated. These generated rule events are finally added to the event queue and form a new event stream for subsequent event processing and rule update.

[0042] S2: Perform de-timeliness processing on event data.

[0043] De-timeliness processing is to ensure that historical data does not affect the current calculation results, that is, to prevent expired data from affecting the current calculation results from the time dimension. Specifically, de-timeliness processing can be carried out according to the timestamps of the data sources, that is, the event data that has exceeded the validity period is removed from the real-time data stream or the event queue.

[0044] Alternatively, the aging process can also be carried out according to the rules of the time window to ensure that only the data within the current time window is valid. Among them, the rules of the time window refer to the rules for determining data validity in distributed real-time computing based on preset time ranges and conditions. For example, it is stipulated that only the data arriving within a specific time window is regarded as valid data for participation in the calculation, and the data outside this time window is considered expired and invalid and not included in the current calculation scope. This ensures that the data used for calculation within each time window meets the requirements of real-time and relevance, thereby ensuring the accuracy and effectiveness of the entire system's calculation and avoiding incorrect calculation results and decisions caused by interference from expired data.

[0045] By aging according to the rules of the data source and time window, the data exceeding the aging period is eliminated, enabling the calculation to be based on valid data. Thus, while leveraging the dynamic changes of the data, the negative impact of historical data is avoided, ensuring the accuracy and reliability of the system's calculation.

[0046] S3: According to the timestamps of each event, respectively match the rules corresponding to each event from the preset rule set, and respectively process each event according to the rules corresponding to each event to obtain the processing results corresponding to each event.

[0047] Specifically, respectively determine the valid time intervals of each rule according to the start time and end time of each rule, that is, each rule is only valid within its valid time interval (i.e., it will be triggered and used for event processing), otherwise it is considered invalid.

[0048] Obtain the events within the current time window, match the valid rules in the rule set according to the characteristics of the events, and process the events according to the valid rules to obtain the processing results. That is, if the characteristics of an event meet the matching conditions corresponding to a rule, and the timestamp of the event is within the valid time interval of the rule, then process the event according to the execution action corresponding to the rule.

[0049] Among them, the characteristics of the event can include data type, data size, data generation frequency, data source, etc. During matching, compare the characteristics of these event data with the corresponding conditions set in the rule. When the characteristics of the event data meet the conditions set in the rule (i.e., belong to the covered range), it is determined that the rule is triggered. After successful matching, process the event according to the action corresponding to the rule and output the processing result.

[0050] Furthermore, after obtaining the processing results, new rule events will also be generated according to the processing results and added to the event queue.

[0051] S4: According to the said processing results, construct a rule set view of the rule set and update the rule set of the rule engine.

[0052] In this distributed real-time computing method, the processing results obtained after event processing are stored in a message queue and used as a condition to trigger subsequent operations. For example, the processing results trigger the update of a rule set. When specific processing result conditions are met, the rule engine updates the rule set based on this trigger condition, thereby achieving the dynamic adaptation of rules to real-time data, ensuring that the system can continuously adjust and optimize the processing logic according to real-time computing results, and maintaining efficient and flexible operation.

[0053] Furthermore, this embodiment further includes: updating the rules in the rule set respectively according to the rule update periods of the rules in the rule set. That is to say, in this embodiment, there are two ways to trigger rule updates. One is to update the rule set according to the rule update period, and the other is when the event processing node receives a new event, the rule set is updated based on the trigger condition of the event.

[0054] To achieve this goal, a rule update function can be defined:

[0055] where f(t i ) represents the output of the rule update function at the timestamp t i of the event, that is, the updated rule set; t i represents the timestamp of the event, T updata represents the rule update period; R k represents the k-th rule in the current rule set, and n is the size of the rule set; δ( ) represents the Dirac delta function, and δ(t i = T updata ) is used to determine whether the time meets the preset trigger condition. When the time meets the condition, its value is 1, and when it does not meet the condition, its value is 0. This rule update function can automatically trigger and update the rule set according to the relationship between the timestamp t i of the event and the rule update period T updata . When t i is equal to T updata , the rule set will be updated and written into the message queue. This function is equivalent to screening out the rules that need to be updated at the moment of t i . Performing such operations on all rules that meet the conditions and summing them up, the updated rule set is obtained, thereby achieving the dynamic update of the rule set to adapt to the continuously changing event data and computing requirements.

[0056] In this way, it can be ensured that during the distributed real-time computing process, the event stream and the rule set can be accurately synchronized and updated according to time conditions, and have the adaptive ability to adapt to the dynamically changing data environment, maintain the efficiency and accuracy of the system's real-time data processing and rule adaptation, and play a key time determination function in the entire distributed real-time computing method of dynamic rule adaptive update. The combination of the event time stream and the rule update mechanism makes the system more flexible and scalable in a dynamic environment.

[0057] In this embodiment, updating the rule set of the rule engine is a core operation. By combining the update of the rule set with the processing of the event stream, the rule engine can efficiently and dynamically adapt to the changing event data in real-time computing. In actual application scenarios, this process can be divided into multiple steps to ensure the accuracy and timeliness of rule updates.

[0058] First, when updating the rule list of the rule set, since events within the same time window may depend on different rules, and these rules can usually be updated independently, a parallel update strategy can be adopted, that is, update the rules on which the events within the same time window depend in parallel, thereby greatly improving the computing efficiency. Specifically, construct the rule set of the time window and update the rule table according to the event characteristics of each time window respectively. These characteristics usually include the trigger time of the event, the time interval between events, and the association conditions of the events, etc. Among them, the association conditions of the events involve the logical relationship and causal relationship between events, such as whether the occurrence of event A is associated with the simultaneous or sequential occurrence of event B, etc.

[0059] When the trigger condition of the time window is met (for example, reaching the preset time window duration, the threshold of the number of specific events, the data feature change conforms to the set pattern, etc.), the system will trigger the rule set in the manner of the time stream to further calculate and update the rules. The update of the rules is ensured by parallel computing so that in a large-scale distributed system, the processing of the event stream is not affected by the update of a single rule. After each rule set is updated, the system will generate a set table of the rule list and write the processing result into the rule set view of the rule engine. The rule set view is a form of presentation and application of the rule set during a specific processing process. It is a specific view based on the rule set generated according to the event processing result for displaying and processing rule-related information. In this embodiment, the rule set view includes the trigger time, the view association conditions, and multiple rule elements. Among them, the view association conditions usually include the identifier of the rule event, the start time, the end time, and the trigger time point, etc., which are used to judge the processing timing of the rule set view.

[0060] According to these conditions, the rule set view will be processed step by step as follows. First, the system determines whether the current time has reached the start time point of the rule set view (this start time point can be determined according to the system's initialization settings, the start time of the time window, or the triggering time of a specific event). If so, the system proceeds to the next step and checks whether the triggering time point has arrived. If the triggering time point has also arrived, the processing result of the rule set view is calculated according to the event processing method in the time flow model, and then the processing result is added to the event flow according to the rule list set table, and the result is written to the start time point of the corresponding time window.

[0061] Specifically, first, based on the event processing logic set in the time flow model, such as according to factors such as the order of events within the time window and the time interval, combined with the triggering conditions of each rule element (including triggering time, start time, end time, etc.), a comprehensive judgment is made. For rules that meet the condition that the current time is within the valid time range of the rule and the triggering conditions are met, calculations are performed according to their defined calculation methods, which involve statistics, analysis of relevant event data, or the application of specific algorithms, and then the calculation results of all rules that meet the conditions are summarized and integrated to obtain the final processing result of the rule view, so as to achieve the dynamic update and adaptation of the rule set. When the processing result is calculated, it will be accurately incorporated into the event flow according to the information in the rule list set table. First, according to the rule list set table, determine the specific position and method where the result should be inserted into the event flow, and then perform the write operation at the start time point of the corresponding time window. This process involves operations on the event flow data structure, such as inserting the processing result data at a specific time series position, ensuring that the event flow can be based on the updated information in subsequent calculations and processing, guaranteeing the integrity and continuity of the event flow data, and enabling the system to accurately use the processing result for subsequent calculations and rule updates according to the predetermined logic and time sequence.

[0062] For example, assume there is a rule set, and its rule set view contains multiple rule elements. The preset triggering times of each rule are t trigger,1 , t trigger,2 , ……, t trigger,n , and the valid time interval of the kth rule is W k , and its start time and end time are T start,k and T end,k The goal of event flow calculation is to update the rule set view according to these rule triggering conditions. The triggering conditions of each rule are calculated through the following formula:

[0063] where t current represents the timestamp in the time flow, and T trigger,k represents the triggering time of the kth rule.

[0064] This formula is based on the current time t current to determine whether the time interval and trigger conditions of each rule are met. If the current time t current is within the valid time range of the rule and the trigger condition has been satisfied, then the rule will be triggered and processed. Through the cumulative calculation of all rules, the final processing result of the rule set view is obtained.

[0065] In this way, the rule set can be dynamically updated according to the time flow model, ensuring that each rule is triggered and executed at the appropriate time. This method not only improves the flexibility of rule processing but also optimizes the processing efficiency of the system through parallel computing.

[0066] This method is based on distributed real-time computing and event stream processing technologies and can be widely applied to various scenarios that require efficient processing of large-scale real-time data streams and dynamic rule matching.

[0067] In practical application scenarios, the rule engine can be first encapsulated into a microservice to implement a dynamic rule engine; then a distributed real-time computing framework based on the event stream can be constructed to form real-time event stream computing; finally, an event time stream processing engine can be constructed, and the real-time computing of the event stream can be realized by adopting a real-time time stream architecture model, that is, the processing of the event stream is triggered through the event processing interface of the real-time event stream architecture model; then the event processing and calculation are carried out according to the time flow of the event stream; finally, the set view of the rule set is constructed from the processing result to update the rule set of the rule engine.

[0068] The microservice architecture and the computing model based on the event stream are combined to construct an efficient, flexible, and scalable real-time computing framework. As Figure 3 shown, this real-time computing framework consists of two parts: a distributed real-time computing framework 100 and a distributed file storage system 200, and can efficiently process massive data in a large-scale distributed environment and perform rule matching, task scheduling, and resource management.

[0069] In the distributed real-time computing framework 100, there are a message queue 101, a management unit 102, and multiple computing units 103. The management unit 102 is used to allocate computing nodes (computing units) for computing tasks and perform task scheduling, termination, and resource allocation according to the usage of system resources. The message queue 101 is mainly used to receive external messages and organize them into an event queue, and at the same time serves as an intermediate link for data transmission and storage during the entire computing process, undertaking the functions of orderly data transfer and temporary storage.

[0070] The computing unit 103 includes a data processing engine 1031 and a distributed rule engine 1032. These two engines are deployed in Docker containers in the form of microservices, that is, the rule engine is encapsulated as a microservice to achieve service isolation, expansion, and high availability. This deployment method enables the rule engine to run, expand, and manage independently. Each rule engine is deployed as an independent microservice, supporting high concurrency and horizontal expansion, enhancing the flexibility and scalability of the system, effectively avoiding the bottleneck problems of the centralized architecture, laying an architectural foundation for implementing a dynamic rule engine, enabling it to better adapt to the changes in real-time data streams and perform dynamic adjustment calculations and rules, and meeting the needs of modern enterprises for real-time data processing and analysis.

[0071] When executing a computing task, the management unit assigns the task to the corresponding computing node. The data processing engine processes the data and performs rule matching, and the calculation result is written into the storage system through a message queue. When executing a data calculation task, the data processing engine will call the distributed rule engine to complete rule matching, ensuring that the rule engine can quickly process and respond to real-time data.

[0072] During the operation of the distributed computing framework, the efficiency of data calculation and the accuracy of rule matching are crucial. The data processing engine and the distributed rule engine jointly complete the processing of the data stream and the execution of the rules. The data processing engine first obtains the data from the message queue, and then executes the computing task according to the requirements of the rule engine. Assume that the data stream of each computing task has a timestamp T data , and the execution period of the rule is T rule (that is, the rule repetition execution interval. In actual application scenarios, it can be determined in combination with specific business scenarios, data characteristics, and system requirements. For example, it can be set according to the urgency of the task, the status of computing resources, etc.). When the data stream arrives, the data processing engine will check whether the timestamp meets the conditions for rule execution. Specifically, the triggering mechanism of the computing task is described by the following formula: f data (t)=1(t≥T data )·1(t≥T rule ) where 1( ) represents the indicator function, which returns 1 when the condition is met and 0 when not met. Here, it means that if the current time t exceeds the data timestamp T data and meets the execution period T of the rule rule , then it returns 1, indicating that the computing task needs to be executed. This triggering mechanism ensures the temporal consistency of data processing and rule matching.

[0073] The distributed file storage system 200 includes multiple storage nodes 201, that is, a unified storage pool is formed by multiple storage nodes together, and the storage capacity is expanded in a cluster manner. The storage nodes interact and share data through a unified access interface, so that the system does not need to re-architect the entire system when expanding storage. This method ensures the high availability and high scalability of data. During the file storage process, data will be dispersed and stored on each node, and the cooperation between the message queue and the storage nodes ensures the persistence and reliability of data.

[0074] When the data traffic is large, the scalability of the distributed file storage system is particularly important, because as the amount of data increases, a single storage node may not be able to meet the requirements. At this time, the system can expand the capacity of the storage pool by adding storage nodes, so as to ensure that the computing framework can continuously process more computing tasks and larger data sets.

[0075] Assume that the storage capacity of each storage node is C node , and the number of storage nodes is N nodes , then the total storage capacity C total of the storage system can be expressed as: C total = N nodes × C node . As the number of storage nodes N nodes increases, the storage capacity of the system is also expanded accordingly, ensuring that it can handle larger-scale data streams.

[0076] For example, assume that the storage capacity of each storage node is 1TB, and the current system has 10 storage nodes, then the total storage capacity of the system is 10TB. If the amount of data continues to increase, the demand can be met by expanding the storage nodes. Assume that 5 new storage nodes are added, and the total storage capacity of the system will be increased to 15TB, thus supporting more computing tasks and data storage requirements.

[0077] This distributed real-time computing framework can not only effectively perform large-scale data processing and rule matching, but also achieve flexible resource scheduling and expansion through the microservices architecture. The collaborative work of the data processing engine and the distributed rule engine ensures the efficient processing of the event stream by the system, while the distributed file storage system provides a powerful storage capacity to support the persistence and fast access of large-scale data. Through this architecture design, the system can cope with dynamic computing requirements and growing data traffic, ensuring efficient and reliable real-time computing and rule updates.

[0078] In the above-mentioned distributed real-time computing framework, the computing process is carried out in a series of ordered steps. The entire process from the reception of the message queue to the final update of the rule set depends on the cooperation of each unit to ensure that the event stream data can be processed in a timely and accurate manner and the corresponding rules can be applied.

[0079] Specifically, it includes the following steps: Step 1: The message queue receives external messages and generates an event queue.

[0080] The message queue is the receiver of the event stream. When triggered by an external data stream or event, the message queue is responsible for organizing this information into an event queue. The event queue contains metadata, timestamps, and the content of various events, providing input for subsequent calculations and rule matching. At this stage, the generation cycle and triggering mechanism of the events are the core. Assuming the timestamp of the event is T event , the message queue receives and sorts these events in chronological order.

[0081] Step 2: The event queue sends messages to the distributed file storage system for storage.

[0082] The distributed file storage system here plays the role of data storage and persistence. Whenever the event queue transfers event data, the storage system writes the event data to the corresponding storage nodes according to the timestamp and storage policy of the event, forming files. The generation of files usually requires a certain amount of storage space and computing resources, so the scalability of the storage system is very important. After the data is stored, it can be read and processed through a unified access interface to support subsequent computing tasks.

[0083] Step 3: The management unit schedules tasks according to the load conditions of each computing unit, that is, the management unit determines a computing unit through a load balancing algorithm and assigns the computing task to the said one computing unit.

[0084] The management unit is responsible for allocating computing tasks to ensure the reasonable use of computing resources. When a computing task needs to be allocated, the management unit will select a computing unit to process the computing task according to the load conditions of each computing unit. For example, the computing task is assigned to the computing unit with the lowest current load. This scheduling process ensures the balanced use of computing resources and improves the overall computing efficiency.

[0085] Step 4: The distributed rule engine completes rule matching, that is, the said one computing unit matches the events in the computing task with a preset rule set and processes the events according to the matching results to obtain a processing result.

[0086] Each rule R in the rule set i has its specific matching conditions and execution actions. The rule engine will match according to the timestamp and characteristics of the event to calculate which rules are triggered. Rule R i 's matching condition can be expressed as: f max (E i ,R i )=1(Ei ∈R i ), where E i represents an event, 1( ) represents an indicator function that returns 1 when the condition holds. Here, it is to determine whether event E i satisfies rule R i 's matching condition. If it is satisfied, 1 is returned, indicating that the rule matching is successful. After the rule matching is successful, the rule engine processes the event according to the execution action corresponding to the rule and outputs the processing result.

[0087] Step 5: Update the event queue according to the processing result.

[0088] During the data processing, the outputs (i.e., processing results) generated through data calculation and rule matching will be added to the event queue as new events. These newly generated events will be further processed to form a new event stream and finally returned to the message queue to wait for the next round of calculation.

[0089] Step 6: Construct a rule set view of the rule set according to the processing result and update the rule set of the distributed rule engine.

[0090] This process involves dynamically updating the rule set stored in the rule engine according to the event processing result. The update of the rule set is adjusted according to different event streams and calculation results, enabling the rule engine to adapt to the changing business requirements and data streams. Each time an update is made, new rules will be added to the rule set or existing rules will be corrected according to the new event stream to maintain the validity and real-time nature of the rules.

[0091] Among them, to determine whether to add new rules or correct existing rules depends on the degree of fit between the new event stream and calculation results and the existing rules. For example, if new data features in a new business scenario appear that cannot be covered by the existing rules, new rules will be added; if the time window limit of a certain rule in the existing rules does not meet the requirements of the new event stream processing, then adjust its time window and other limit conditions to correct the existing rule so that it can adapt to the new situation.

[0092] Through the above calculation process, each step is closely linked, ensuring that the system can flexibly and effectively process event streams from different data sources and perform rule matching. In practical applications, the processing of event streams and the update of the rule engine need to respond quickly to ensure that the system can continuously and efficiently operate in a complex real-time environment.

[0093] For example, in a real-time monitoring system, new data streams of sensor data are sent every second, and these data streams are received and stored by the message queue. The distributed computing framework performs task scheduling according to the load conditions of computing nodes (computing units), and the rule engine determines whether there are abnormalities based on these data streams, such as whether the temperature exceeds a preset threshold. If an abnormal rule is triggered, the system will generate new events and update the rules so that subsequent processing can handle different scenarios. For example, if the temperature exceeds the preset threshold and triggers an abnormal rule, a "high temperature warning event" can be generated, recording in detail information such as the time, location of the high temperature occurrence, and the specific value exceeding the threshold. If the original judgment of temperature abnormality was based only on a single threshold, after this event, the rule can be corrected to comprehensively consider the temperature change trend over a period of time. Suppose the previous rule was that a warning was triggered when the temperature exceeded 80°C, and now it is updated to trigger a warning if the temperature continuously rises and exceeds 75°C within 10 seconds, so that the rule can better handle different scenarios and improve the system's monitoring and handling capabilities for abnormal situations.

[0094] This embodiment can be applied to the following scenarios: 1. Intelligent manufacturing and industrial Internet of Things In intelligent manufacturing and industrial Internet of Things (IIoT), various devices and sensors generate a large amount of real-time data, such as sensor data of temperature, humidity, vibration, pressure, etc. The production line needs to quickly process and analyze this real-time data to achieve goals such as equipment monitoring, fault prediction, and production optimization. Through this solution, it is possible to process the operation data of equipment in real time, dynamically adjust the rule engine, automatically identify potential equipment failures or abnormal states, and trigger alarms or automatically adjust the production process. The distributed computing architecture enables the system to adapt to the real-time processing requirements of large-scale device and sensor data, ensuring efficient and stable operation.

[0095] Application example: In a smart factory, each machine device and sensor upload data in real time through the network, and the system monitors in real time and analyzes the equipment status, automatically triggering maintenance operations. If a certain sensor detects abnormal equipment temperature, through the analysis of the rule engine, a fault alarm is triggered, and even the production is stopped for maintenance in advance through the automation system to avoid production stagnation and equipment damage.

[0096] 2. Financial risk monitoring and anti-fraud In the financial field, real-time transaction monitoring and anti-fraud systems need to quickly analyze a large amount of transaction data to identify potential risks or fraud behaviors. Through this solution, the system can process each transaction data stream in real time, dynamically update risk rules and models, and perform intelligent detection based on multiple triggering conditions (such as transaction amount, transaction frequency, account abnormality, etc.). For example, when the transaction frequency of a certain account exceeds the set threshold, the system can immediately trigger the anti-fraud mechanism, automatically suspend the transaction or conduct further manual review.

[0097] Application Example: Banks or payment platforms use this solution to monitor users' transaction behaviors in real time. When the system detects abnormal transaction patterns (such as large - amount fund transfers, frequent transactions in different regions, etc.), it triggers preset anti - fraud rules, initiates automatic risk reviews, and prevents fund losses.

[0098] 3. Smart City and Traffic Monitoring In the fields of smart cities and traffic management, data such as traffic flow, vehicle positions, and road conditions need to be collected and processed in real time. Through this solution, it is possible to quickly analyze traffic flow data, identify traffic congestion situations, and adjust the traffic signal control system in real time to optimize traffic flow. Event - stream computing can respond immediately to road accidents or traffic incidents, trigger emergency dispatching or release traffic information, and help reduce traffic accidents and congestion.

[0099] Application Example: In an intelligent transportation system, multiple surveillance cameras, traffic sensors, and GPS devices provide real - time data. The system automatically adjusts signal control according to traffic flow information. When an accident or emergency occurs, it issues early warnings in real time and adjusts traffic routes.

[0100] 4. Energy Management and Smart Grid In the fields of smart grids and energy management, the collection and analysis of real - time data are crucial for the stable operation of the power system. This solution can monitor information such as power consumption, power generation, and equipment status in real time. Through a dynamically updated rule engine, it can adjust the operating state of the power grid in real time to avoid overload or power supply interruptions. When power anomaly events (such as overload, equipment failures, etc.) occur, the system can respond immediately and perform automatic dispatching or alarm to ensure the stability and security of power supply.

[0101] Application Example: In a smart grid, sensors and monitoring systems monitor the working status of power equipment in real time, calculate the load distribution in real time, and perform dynamic dispatching based on real - time data. If the power demand in a certain area exceeds the load, the system can identify it in real time and automatically switch the load or start the backup power supply to avoid power supply interruptions.

[0102] 5. E - commerce and Online Marketing In the fields of e - commerce and online advertising, the processing of real - time data streams can help platforms monitor user behaviors, transaction data, and advertising effects in real time. Based on the event - stream computing model, the system can analyze users' browsing, clicking, purchasing, etc. behaviors in real time, automatically adjust advertising push strategies, product recommendations, etc., and enhance the user experience and platform sales. The dynamic rule engine can intelligently recommend personalized advertisements or products based on users' behavior data, improving the advertising conversion rate and sales volume.

[0103] Application Example: An e-commerce platform analyzes users' shopping behaviors through real-time data streams and pushes relevant products in combination with personalized recommendation algorithms. For users who frequently browse a certain category of products, the system automatically triggers personalized advertisement pushing, adjusts marketing strategies in real time, and improves conversion rates.

[0104] 6. Environmental Monitoring and Disaster Warning Environmental monitoring and disaster warning systems need to conduct real-time data analysis on natural phenomena such as air quality, weather changes, and earthquakes. Through this solution, the environmental monitoring system can process sensor data in real time and analyze it according to set rules, and promptly identify environmental pollution or abnormal situations. When the air quality index exceeds the safety standard, the system can issue a warning in real time and activate the emergency response mechanism to notify relevant departments or the public to take protective measures to avoid personal injuries.

[0105] Application Example: In an environmental monitoring system, an air quality monitoring station collects data in real time and analyzes it through a rule engine in real time. When it detects that the air pollution exceeds the standard, the system triggers a warning to notify government departments or residents to take emergency response measures, such as activating air purification facilities or issuing health protection notices.

[0106] This embodiment has the following beneficial effects: 1. Efficient real-time data processing and calculation. Through the event stream-based calculation model, the system can quickly respond after receiving an event and perform real-time calculations according to a predefined time window. Using message queues and event queues, data streams can be processed efficiently in sequence. The rule engine matches according to real-time data streams to ensure that the system responds promptly to changing data. This mechanism can ensure the efficient processing and calculation of data streams in scenarios that require quick decision-making and processing, reduce system latency, and improve real-time response capabilities.

[0107] 2. Flexible scalability and high availability. The system adopts a microservices architecture, and each computing unit and rule engine are independent Docker containers for deployment, supporting horizontal scaling. The distributed architecture ensures that the system can dynamically adjust computing resources according to the load, and when a certain computing unit or service fails, other units can automatically take over tasks to ensure the continuous operation of the system. This flexible resource management and scheduling mechanism enables the system to meet the requirements of high concurrency, large-scale data processing, and dynamic loads.

[0108] 3. Load balancing and optimized resource scheduling. The management unit dynamically assigns tasks according to the load conditions of computing units to ensure the reasonable utilization of computing resources. Through load balancing algorithms, the system can avoid overloading of a single computing node, ensure the smooth distribution of tasks, and improve computing efficiency. For example, adopting a scheduling algorithm based on the current node load situation can ensure that tasks are assigned to the computing node with the lightest load, improving the execution efficiency of computing tasks and the system response capabilities.

[0109] 4. Dynamic rule update and a highly adaptable rule engine. The distributed rule engine can dynamically adjust the rule set according to the real-time calculation results and changes in the data stream. Whether it is a rule triggered at a fixed time or a rule triggered based on a data event, the rule engine can update the rule set in a timely manner according to different scenarios to ensure that the rules are consistent with the real-time data. This dynamic rule update mechanism enables the system to still operate efficiently and flexibly in the face of constantly changing business requirements and data streams.

[0110] 5. High-efficient distributed storage and data processing capabilities. The distributed file storage system consists of multiple storage nodes and a unified access interface to form a storage pool, which can expand the storage capacity to support the storage of massive data. The close combination of the message queue and the storage system enables event data to be persistently stored in a distributed environment and quickly read during subsequent calculation processes. This storage mechanism can ensure that the system always maintains a high storage access efficiency when processing a large amount of data, meeting the requirements of high throughput and low latency.

[0111] 6. High reliability and fault tolerance. This solution adopts a fault tolerance mechanism to ensure that when a computing unit or a storage node fails, other nodes can take over the tasks in a timely manner to ensure the uninterrupted calculation process. The event queue and the message queue provide persistent storage of event data to ensure that even when some computing units are unavailable, the event data can still be stored and waiting to be processed. Through the redundant design and data recovery mechanism of the distributed computing framework, the system can maintain high availability and reliability in the face of hardware failures or network problems.

[0112] Embodiment 2 This embodiment is a computer-readable storage medium corresponding to the above embodiment, on which a computer program is stored. When the program is executed by a processor, it implements each step of a rule update method based on distributed real-time calculation as in the above embodiment and can achieve the same technical effects, which will not be repeated here.

[0113] In summary, the rule update method and storage medium based on distributed real-time computing provided by the present invention can make up for the defects of the prior art in multiple aspects and have significant advantages. First, by adopting a microservices architecture and a distributed computing framework, the system can dynamically schedule tasks according to the load situation, avoiding bottleneck problems in the centralized architecture. Each computing unit and rule engine are deployed as independent microservices, supporting high concurrency and horizontal scalability, thus enhancing the flexibility and scalability of the system. Second, the distributed rule engine and the dynamic rule update mechanism solve the static problem of the traditional rule engine. Through the real-time processing of the event stream, the rules can be flexibly adjusted according to the changes in the data stream, ensuring the effectiveness and accuracy of the rules. This mechanism can significantly improve the system's adaptability to the dynamic data environment and avoid the problem of traditional rule lag. Finally, the combination of the distributed storage system and the message queue enables large-scale data to be efficiently stored and accessed, and ensures the fault tolerance of the system. In the case of a sharp increase in data traffic or the failure of a computing node, the system can continue to run and process data, ensuring the high availability of the service and the reliability of the data. This solution effectively solves the deficiencies of traditional technologies in terms of processing capacity and high availability in the big data environment and can better meet the needs of modern enterprises for real-time data processing and analysis.

[0114] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in the related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A rule updating method based on distributed real-time computing, characterized in that: include: Get data events based on real-time data streams; According to the timestamp of each event, the rules corresponding to each event are matched from the preset rule set, and each event is processed according to the rules corresponding to each event to obtain the processing results corresponding to each event, wherein the events include data events; According to the processing result, a rule set view of the rule set is constructed, and the rule set of the rule engine is updated.

2. The rule updating method according to claim 1, characterized in that: Also includes: The rules in the rule set are updated according to the rule update period of each rule in the rule set.

3. The rule updating method according to claim 1, characterized in that: According to the timestamp of each event, the rules corresponding to each event are matched from the preset rule set respectively, and each event is processed according to the rules corresponding to each event to obtain the processing results corresponding to each event, including: Determine the effective time interval of each rule according to the start time and end time of each rule; If the feature of an event satisfies the matching condition corresponding to a rule, and the timestamp of the event is within the valid time interval of the rule, the rule is used as the rule corresponding to the event; Processing the event according to the execution action corresponding to the rule to obtain a processing result of the event; Among them, the characteristics of events include data type, data size, data generation frequency and data source.

4. The rule updating method according to claim 1, characterized in that: After obtaining data events according to the real-time data stream, the method further includes: The data events are added to an event queue, and the data events corresponding to each time window are determined according to the timestamp of each data event and a preset time window length.

5. The rule updating method according to claim 4, characterized in that: The event also includes a rule event; the method also includes: Generate new rule events according to the preset trigger cycle; If a new real-time data stream arrives in the current time window, a new rule event is generated according to the new real-time data stream and the preset rule template; If no new real-time data stream arrives in the current time window, a new data stream is generated based on the data events corresponding to the previously preset number of time windows and their processing results, and a new rule event is generated based on the new data stream and the preset rule template; Add the new rule event to the event queue.

6. The rule updating method according to claim 4, characterized in that: The events also include rule events; After processing each event according to the rules corresponding to each event and obtaining the processing result corresponding to each event, the method further includes: According to the processing result, a new rule event is generated, and the new rule event is added to the event queue.

7. The rule updating method according to claim 4, characterized in that: Through parallel processing, the rules corresponding to events in the same time window are updated.

8. The rule updating method according to claim 1, characterized in that: Before matching the rules corresponding to the events from the preset rule set according to the timestamps of the events, the method further includes: If the current time point exceeds the validity period of an event, the event is removed; Or, if the timestamp of an event is not within the current time window, the event is removed.

9. The rule updating method according to claim 1, characterized in that: The method is based on a real-time computing framework, which includes a distributed real-time computing framework and a distributed file storage system. The distributed real-time computing framework includes a message queue, a management unit and at least one computing unit, and the computing unit includes a data processing engine and a distributed rule engine; the distributed file storage system includes at least one storage node; The message queue is used to receive the real-time data stream and add the data events in the real-time data stream to the event queue; The management unit is used to allocate computing tasks to computing units through a load balancing algorithm; The computing unit is used to perform event rule matching and event processing after receiving the computing task; The distributed file storage system is used to write event data into corresponding storage nodes according to the timestamp of the event and a preset storage strategy.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.