Method and apparatus for data stream processing based on multi-version rules
Patent Information
- Application Number
- CN202610819956.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-08
- Publication Date
- 2026-09-01
AI Technical Summary
[0003]有鉴于此,本发明的目的在于提供一种基于多版本规则的数据流处理方法及装置,以改善现有技术中跨版本数据处理困难的问题
本发明提供的上述基于多版本规则的数据流处理方法及装置,首先响应于系统规则更新指令,生成目标规则版本,并在规则版本时间轴中构建规则生效时间窗口;然后获取待处理数据流的时间戳区间,并基于时间戳区间和规则生效时间窗口,计算待处理数据流的目标状态参数;接着基于目标规则版本,在隔离的沙箱计算环境中对历史状态数据集执行全量重算,生成差异影响报告;最后基于差异影响报告,对待处理数据流执行路由分发处理。上述方法中,通过构建规则生效时间窗口,实现了对规则生命周期的精细化管理;通过时间戳区间与规则生效时间窗口的匹配计算,改善了跨版本数据流的状态计算难题;引入隔离沙箱环境对历史数据进行全量重算(回归测试),生成差异影响报告,确保了新规则上线的安全性和可预测性;结合差异报告与合规结果动态路由,有效规避了因规则变更带来的业务风险。
Smart Images

Figure CN122673170A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing and rule engine technology, and in particular to a data stream processing method and apparatus based on multi-version rules. Background Technology
[0002] In data-intensive sectors such as fintech, e-commerce, and the Internet of Things, business rules (such as risk control strategies, billing logic, and routing rules) are updated extremely frequently. Existing rule processing systems typically adopt an "instant effect" or "full switchover" model. However, this model has significant drawbacks: when rules change, long-cycle data streams spanning time windows (e.g., days or months) often struggle to determine which version of the rule should apply, easily leading to billing or status calculation errors. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a data stream processing method and apparatus based on multi-version rules, so as to improve the problem of difficult cross-version data processing in the prior art.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a data stream processing method based on multi-version rules, comprising: generating a target rule version in response to a system rule update instruction, and constructing a rule effective time window in the rule version timeline; obtaining the timestamp interval of the data stream to be processed, and calculating the target state parameters of the data stream to be processed based on the timestamp interval and the rule effective time window; performing a full recalculation of the historical state dataset in an isolated sandbox computing environment based on the target rule version, and generating a difference impact report; and performing routing distribution processing on the data stream to be processed based on the difference impact report.
[0005] Optionally, a rule effective time window can be constructed in the rule version timeline, including: obtaining the canary release strategy carried in the system rule update instruction; determining the official effective date of the target rule version based on the canary release strategy on the rule version timeline; and constructing the rule effective time window based on the official effective date.
[0006] Optionally, based on the timestamp interval and the rule effective time window, the target state parameters of the data stream to be processed are calculated, including: determining whether the timestamp interval of the data stream to be processed crosses the boundary of the rule effective time window; if the timestamp interval crosses the boundary of the rule effective time window, the data stream to be processed is divided into a first data sub-stream and a second data sub-stream; wherein, the first data sub-stream corresponds to the old rule version, and the second data sub-stream corresponds to the target rule version; the corresponding rule versions are applied to the first data sub-stream and the second data sub-stream respectively to perform state calculations, and the calculation results are merged to generate the target state parameters of the data stream to be processed.
[0007] Optionally, based on the target rule version, a full recalculation of the historical state dataset is performed in an isolated sandbox computing environment to generate a difference impact report. This includes: constructing a virtual execution container based on the target rule version; loading a snapshot of the historical state dataset from persistent storage and injecting it into the virtual execution container; executing the logical operations of the target rule version in parallel within the virtual execution container to obtain the calculation results of each historical state data, and comparing the calculation results with the original states in the historical state dataset to generate a difference impact report.
[0008] Optionally, the difference impact report includes the compliance verification result of the data stream to be processed; based on the difference impact report, the data stream to be processed is routed and distributed, including: if the compliance verification result of the data stream to be processed is passed, the data stream to be processed is routed to the high-speed processing channel; if the compliance verification result of the data stream to be processed is failed, a manual review process is triggered, and the data stream to be processed is temporarily routed to the dead letter queue or the delayed processing channel until the review approval instruction is received.
[0009] Optionally, after performing routing and distribution processing on the data stream to be processed, the following steps are also included: real-time monitoring of the resource consumption metrics and error logs of the target rule version when processing the data stream to be processed; when it is detected that the error type in the error log belongs to a rule logic conflict, based on the rule version timeline, automatically roll back to the previous version of the target rule version, and add an emergency patch mark to the rule effective time window corresponding to the previous version.
[0010] Optionally, generating a target rule version includes: parsing the system rule update instruction and extracting the rule change difference set; identifying the affected dependent data fields based on the rule change difference set; performing a directed acyclic graph topological sort on the target rule version to determine the rule execution order; and generating an index optimization strategy corresponding to the dependent data fields while generating the target rule version, and applying the index optimization strategy to the reading process of the historical state dataset.
[0011] Secondly, the present invention provides a data stream processing device based on multi-version rules, comprising: a version update module, used to generate a target rule version in response to a system rule update instruction, and construct a rule effective time window in the rule version timeline; a target state parameter calculation module, used to obtain the timestamp interval of the data stream to be processed, and calculate the target state parameters of the data stream to be processed based on the timestamp interval and the rule effective time window; a sandbox simulation module, used to perform a full recalculation of the historical state dataset in an isolated sandbox computing environment based on the target rule version, and generate a difference impact report; and a routing processing module, used to perform routing distribution processing on the data stream to be processed based on the difference impact report.
[0012] Thirdly, the present invention provides an electronic device including a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method provided in any of the first aspects above.
[0013] Fourthly, the present invention provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, performs the steps of the method provided in any of the first aspects above.
[0014] This invention brings the following beneficial effects: The data stream processing method and apparatus based on multi-version rules provided by this invention first responds to system rule update instructions, generates a target rule version, and constructs a rule effective time window in the rule version timeline; then, it obtains the timestamp interval of the data stream to be processed, and calculates the target state parameters of the data stream based on the timestamp interval and the rule effective time window; next, based on the target rule version, it performs a full recalculation of the historical state dataset in an isolated sandbox computing environment to generate a difference impact report; finally, based on the difference impact report, it performs routing and distribution processing on the data stream to be processed. In this method, by constructing a rule effective time window, fine-grained management of the rule lifecycle is achieved; by matching the timestamp interval with the rule effective time window, the state calculation problem of cross-version data streams is improved; by introducing an isolated sandbox environment to perform a full recalculation (regression testing) of historical data and generate a difference impact report, the security and predictability of new rule deployment are ensured; and by combining the difference report with dynamic routing of compliance results, business risks caused by rule changes are effectively avoided.
[0015] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0016] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 A flowchart of a data stream processing method based on multi-version rules provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a data stream processing device based on multi-version rules provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] Currently, existing rule processing systems typically adopt an "instant effect" or "full switch" model. However, this model has significant drawbacks: when rules change, long-cycle data streams spanning time windows (such as days or months) often make it difficult to determine which version of the rule should apply, which can easily lead to billing or status calculation errors.
[0021] Based on this, the present invention provides a data stream processing method and apparatus based on multi-version rules, which can improve the problem of difficult cross-version data processing in the prior art.
[0022] To facilitate understanding of this embodiment, a data stream processing method based on multi-version rules disclosed in this invention will first be described in detail. This method can be executed by electronic devices, such as smartphones, computers, and tablets. See also Figure 1 The flowchart shown illustrates a data stream processing method based on multi-version rules, which mainly includes the following steps S101 to S104: Step S101: In response to the system rule update command, generate the target rule version and construct the rule effective time window in the rule version timeline.
[0023] In one implementation, each time a system rule is modified, the system does not directly overwrite the original database record, but instead generates a new rule version, namely the target rule version V. n+1 And mark its effective time window [T] effective (,infty), and automatically includes the previous version V n The end time closure is T effectiveThe system maintains a continuous, non-overlapping policy version timeline, ensuring that at any historical point in time t, there is one and only one definite rule set P(t) in effect.
[0024] Step S102: Obtain the timestamp interval of the data stream to be processed, and calculate the target state parameters of the data stream to be processed based on the timestamp interval and the rule effective time window.
[0025] In one implementation, for a data stream to be processed (e.g., an in-transit document in the approval flow of a financial system), the system determines the status based on the business occurrence date (i.e., the timestamp interval) of the data stream. For long-cycle documents that span rule change dates, the system executes a time slicing algorithm to split the document into old rule segments and new rule segments, matches the corresponding rule versions for segmented calculations, and then summarizes the results to obtain the target status parameters of the data stream to be processed.
[0026] Step S103: Based on the target rule version, perform a full recalculation of the historical state dataset in an isolated sandbox computing environment to generate a difference impact report.
[0027] In one implementation, before the official release of the target rule version, the administrator can start a simulation program. The system loads a draft of the target rule version into an isolated sandbox computing environment, fully reruns the historical state dataset of the past year (or a specified period), and generates a difference impact report. The difference impact report includes: the increase in total annual cost if the new rule is implemented, changes in budget occupancy rates for each department, and the types of documents with increased violation rates, etc.
[0028] In practice, historical documents are sharded by Dept_ID or Time_Window and distributed to various computing nodes. Each node loads the new rule draft, re-executes the risk control rule engine, and calculates the new control result (approval / rejection) and amount (approved amount).
[0029] Reduce phase: Summarize the calculation results from each node and aggregate them to generate Diff Metrics: Total cost change: ΔCost= Where NewAmount is the total cost under the new rules and OldAmount is the total cost under the old rules.
[0030] Change in compliance rate: ΔComplianceRate=(Count(NewPassed)-Count(OldPassed)) / TotalCount; where Count(NewPassed) represents the number of records that passed verification (compliance) under the new rules, Count(OldPassed) represents the number of records that passed verification (compliance) under the old rules, and TotalCount represents the total number of records participating in the assessment.
[0031] In addition, it supports drill-down analysis of the sources of discrepancies by department, job level, and expense type (e.g., "The sales department's accommodation over-standard rate increased by 15% due to the new policy").
[0032] Step S104: Based on the difference impact report, perform routing distribution processing on the data stream to be processed.
[0033] In one implementation, for pending data streams whose compliance status changes due to rule modifications, the system performs differentiated operations based on the configuration: From lenient to strict (from compliant to non-compliant): Automatically triggers approval blocking, returns the pending data stream to the applicant with the message "Due to policy changes, please readjust", or automatically generates an "exceeding the limit explanation" and marks it as abnormal.
[0034] From strict to lenient (from non-compliant to compliant): Automatically remove the warning mark for exceeding the limit, or re-trigger the automatic approval rule (e.g., directly approve after the amount becomes compliant).
[0035] For scenarios requiring retrospective adjustments (such as retroactively issuing subsidy discrepancies from the past six months), the system automatically generates supplementary payment or deduction documents based on sandbox calculation results, and links them to the original document ID to form a complete traceability evidence chain, eliminating the need for manual entry by employees. Furthermore, regardless of rule changes, the system saves snapshot copies of the rules referenced by the documents at both the generation and posting times, ensuring that the original calculation basis can be restored during audits.
[0036] The data stream processing method based on multi-version rules provided in this invention achieves refined management of the rule lifecycle by constructing a rule effective time window; it improves the problem of state calculation for cross-version data streams by matching and calculating the timestamp interval with the rule effective time window; it introduces an isolated sandbox environment to perform full recalculation (regression testing) on historical data and generate a difference impact report, ensuring the security and predictability of new rule deployment; and it effectively avoids business risks caused by rule changes by combining the difference report with dynamic routing of compliance results.
[0037] In one implementation, for the aforementioned step S101, i.e., when generating the target rule version, the following methods may be used, including but not limited to: First, parse the system rule update instructions and extract the rule change difference set.
[0038] In practical implementation, upon receiving a system rule update instruction, the instruction is first parsed in a structured manner to clarify the type of rule change (e.g., addition, modification, or deletion) and its specific semantics. To accurately extract the rule change difference set, the system can employ an incremental specification mechanism to avoid directly overwriting the original rule text. A difference file containing precise change semantics records the evolution from the old state to the new state. For example, for modification operations, the constraints before and after the change must be recorded; for deletion operations, the removed rule and the reason must be recorded; for addition operations, the content of the new rule and its business context must be recorded. This approach ensures that every change is traceable, providing a deterministic input basis for subsequent impact analysis.
[0039] Secondly, based on the rule change difference set, the affected dependency data fields are identified.
[0040] In practical implementation, after obtaining a precise set of change differences, the impact of these changes on the underlying data is further analyzed. To avoid the high false positive rate of traditional table-level or column-level lineage analysis, this embodiment of the invention can employ operator-level lineage resolution technology. By performing deep parsing of data processing logic (such as SQL scripts or execution plans) at the abstract syntax tree level, the dependencies of internal operators such as filtering, joining, and aggregation are captured. Combined with directed graph traversal algorithms in graph databases, a deep search is performed upstream or downstream from the rule node where the change occurred, thereby accurately locating the data fields and intermediate layer models truly affected by the change, eliminating irrelevant noise data, and significantly narrowing the scope of impact assessment.
[0041] Next, the target rule version is topologically sorted using a directed acyclic graph to determine the rule execution order.
[0042] In implementation, after identifying the affected fields and constructing new rule versions, the execution order of all relevant rules is determined. Since rules often have complex dependencies, the system treats each rule as a node and dependencies as directed edges, constructing a directed acyclic graph (DAG). A topological sorting algorithm is then used to process this graph. During the sorting process, a priority mechanism is employed, comprehensively considering the rule's own declared priority, community or globally shared rules, and user-defined intervention rules. If a circular dependency is detected, a conflict detection and early warning mechanism is triggered. Finally, topological sorting generates a linear rule loading sequence that conforms to all dependency and priority constraints, ensuring the logical correctness of rule execution.
[0043] Finally, while generating the target rule version, an index optimization strategy corresponding to the dependent data fields is generated, and the index optimization strategy is applied to the reading process of the historical state dataset.
[0044] In practical implementation, while generating the target rule version, corresponding index optimization strategies are dynamically generated based on the data query patterns involved in the new rules. For fields with high-frequency access or those used as core filtering conditions, composite indexes or expression indexes are automatically planned to accelerate retrieval under specific conditions. For fields containing complex nested structures or semi-structured data, generated columns are used to extract them into typed scalar values, thereby building more efficient B-tree indexes, or GIN indexes are used to support flexible inclusion queries when the query pattern is uncertain. In addition, for reading historical state datasets, combined with the concept of time interval modeling, dedicated time composite indexes are built for historical records with effective and expiration times. When executing historical backtracking queries, the query engine can directly hit these pre-built optimized indexes, quickly filtering out valid data at the target time point through range scans, thereby significantly improving the reading performance of historical data while ensuring the complete traceability of data.
[0045] In one implementation, for the aforementioned step S101, i.e., when constructing the rule effective time window in the rule version timeline, the following methods may be adopted, including but not limited to: First, obtain the canary release strategy carried in the system rule update instruction; then, on the rule version timeline, determine the official effective period of the target rule version according to the canary release strategy; finally, construct the rule effective time window based on the official effective period.
[0046] In practice, upon receiving a rule update instruction from the system, a canary release strategy is extracted from the instruction's metadata or associated configuration. This strategy typically exists in the form of structured parameters and is used to guide the smooth transition of new version rules.
[0047] After obtaining the gray-scale strategy, the system extrapolates the official full-scale implementation time of the new version on an abstract rule version timeline: First, the system uses the time when the rule enters the test environment or a single-node pilot as the baseline. Then, based on the extracted gray-scale strategy (e.g., a tiered plan of gradually increasing the scale by 10%, 30%, and 50%), and the observation period set for each stage, the system performs an cumulative calculation. Simultaneously, the extrapolation logic employs dynamic evaluation of exit criteria; that is, the next stage's timeline is only included in the plan if the current stage's monitoring metrics (e.g., error rate, response latency) meet the standards. If an exception rollback mechanism is triggered, the implementation period must be recalculated. Finally, the moment when all stages pass verification and traffic reaches 100% is confirmed as the official implementation date of the target rule version.
[0048] Once the official effective date is determined, the system translates it into a rule effective time window at the physical storage and query levels. In the rule model of the database or configuration center, a precise lifecycle timestamp is bound to each rule. Specifically, the system writes the officially effective date derived above into the rule's valid expiration time field, and uses the expiration time of the previous version or the start point of the current time as the valid start time.
[0049] In one implementation, for the aforementioned step S102, i.e., when calculating the target state parameters of the data stream to be processed based on the timestamp interval and the rule effective time window, the following methods may be used, including but not limited to: determining whether the timestamp interval of the data stream to be processed crosses the boundary of the rule effective time window; if the timestamp interval crosses the boundary of the rule effective time window, then dividing the data stream to be processed into a first data sub-stream and a second data sub-stream; wherein, the first data sub-stream corresponds to the old rule version, and the second data sub-stream corresponds to the target rule version; applying the corresponding rule versions to the first data sub-stream and the second data sub-stream respectively to perform state calculations, and merging the calculation results to generate the target state parameters of the data stream to be processed.
[0050] In practice, the minimum and maximum timestamps of the data stream to be processed are extracted to form a dynamic timestamp interval. This interval is then compared to the boundary of the rule's effective time window (i.e., the switching point between the old and new rules). If the start point of this timestamp interval is earlier than the switching point, but the end point is later than the switching point, it can be determined that the data stream to be processed has experienced a time-lapse phenomenon.
[0051] Furthermore, based on the detected rule switching time point, the data stream to be processed is split into time segments: data records with timestamps less than the switching time point are routed to the first data sub-stream, and these records are strictly bound to the old rule version; at the same time, data records with timestamps greater than or equal to the switching time point are routed to the second data sub-stream, and these records are bound to the target new rule version.
[0052] After data segmentation is completed, the system's state calculation engine will perform independent aggregation operations on the two sub-streams in parallel or serially. For the first data sub-stream, the engine loads the old version of the rule configuration and historical dependency states, processes this early data according to the original calculation logic, and generates intermediate state parameters reflecting the old rules. At the same time, for the second data sub-stream, the engine activates the target new version of the rule configuration, processes it according to the new business logic, and generates intermediate state parameters under the new rules.
[0053] After each of the two sub-streams completes its state calculation, a structured reorganization is performed based on the business semantics. For example, if the calculation is an accumulation value, the local aggregation results of the two sub-streams are summed; if the calculation is a deduplication count, the set of elements generated by the two sub-streams is joined.
[0054] In addition, the system needs to update the global water level and window trigger states to ensure that downstream operators can perceive the complete processing progress of this cross-period batch. Through this refined segmentation and merging mechanism, the system not only ensures the absolute accuracy of data processing during rule hot updates, but also avoids business deviations caused by forcibly discarding late data or incorrectly applying new rules.
[0055] For example: For the data stream to be processed, the itinerary T trip (Time interval [D] start D end ]), rule chain P chain Among them, rule A: validity period [...], D change Rule B: Validity period [D] change ,...).
[0056] Detecting T trip Does it cross a policy breakpoint (i.e., a switching time point)? (D) change If the journey crosses a certain distance, the journey will be divided into segments S1[D]. start D change -1] and S2[D change D end ].
[0057] S1 matches rule A, calculates the first target state parameter Limit1=Std A Days (S1).
[0058] S2 matching rule B, calculate the second target state parameter Limit2=Std B Days (S2).
[0059] Target state parameter Limit total =Limit1+Limit2.
[0060] In one implementation, for the aforementioned step S103, that is, when performing a full recalculation of the historical state dataset in an isolated sandbox computing environment based on the target rule version to generate a difference impact report, the following methods may be adopted, including but not limited to: First, construct a virtual execution container based on the target rule version.
[0061] In practice, a completely isolated sandbox runtime environment is created. This virtual execution container fully replicates the current production environment's runtime context during the initialization phase, including loading the latest configuration of the target rule version, business logic scripts, and related dependency libraries. Simultaneously, the system pre-configures an in-memory state storage engine (such as an embedded database or distributed cache) within the container to simulate the real data interaction interface of the production environment. This sandbox mechanism ensures that all subsequent test operations are strictly confined within independent boundaries, eliminating the risk of dirty data contaminating online business operations.
[0062] Then, a snapshot of the historical state dataset is loaded from persistent storage and injected into the virtual execution container.
[0063] In practice, snapshots of historical state datasets at specific points in time are pulled from underlying persistent storage (such as object storage or a data lake). To ensure absolute data consistency, the system employs Multi-Version Concurrency Control (MVCC) when extracting snapshots, ensuring that a complete static view of a transaction is obtained. Subsequently, these historical snapshots are written in batches to the memory-level state storage engine within a virtual container via a high-speed data pipeline.
[0064] Finally, the logical operations of the target rule version are executed in parallel within the virtual execution container to obtain the operation results of each historical state data. The operation results are then compared with the original states in the historical state dataset to generate a difference impact report.
[0065] In practice, the virtual execution container drives the target rule version to replay all historical data. To improve processing efficiency and shorten the evaluation cycle, a parallel computing architecture is adopted, where the historical state dataset is split according to the primary key or sharding strategy and distributed to multiple concurrent threads or worker nodes for simultaneous processing. Each computing node independently executes the complex business logic of the target rule in its own context, and temporarily stores the intermediate results and final output in the container's temporary workspace.
[0066] After all parallel computing tasks are completed, the comparison engine aligns the result set from the new rule calculation with the original injected historical state snapshots line by line. Comparison dimensions include: changes in core business field values (such as changes in amount or status), and changes at the data structure level (such as adding fields or type conversions). For identified discrepancies, a severity level is automatically assigned based on a preset business sensitivity matrix; for example, discrepancies leading to primary key conflicts or changes in funds are marked as high-risk, while minor adjustments to statistical indicators are marked as low-risk. Finally, the system aggregates this structured information and automatically generates a discrepancy impact report, intuitively presenting the potential business impact of the new version of the rules, thus providing solid data support for release decisions.
[0067] In one implementation, the difference impact report includes the compliance verification result of the data stream to be processed. For the aforementioned step S104, that is, when performing routing and distribution processing on the data stream to be processed based on the difference impact report, the following methods may be adopted, including but not limited to: if the compliance verification result of the data stream to be processed is passed, the data stream to be processed is routed to the high-speed processing channel; if the compliance verification result of the data stream to be processed is failed, a manual review process is triggered, and the data stream to be processed is temporarily routed to the dead letter queue or the delayed processing channel until a review approval instruction is received.
[0068] In one implementation, after performing routing and distribution processing on the data stream to be processed, the above method further includes: real-time monitoring of the resource consumption indicators and error logs of the target rule version when processing the data stream to be processed; when it is detected that the error type in the error log belongs to the rule logic conflict, based on the rule version timeline, automatically roll back to the previous version of the target rule version, and add an emergency patch mark to the rule effective time window corresponding to the previous version.
[0069] In practice, probes or exporters mounted on the underlying computing engine collect real-time resource consumption metrics for the current rule instance, such as CPU utilization, memory usage, processing latency, and throughput. Simultaneously, an independent error log collection pipeline is established to perform structured analysis and categorized aggregation of anomalies generated during operation. To ensure timely monitoring, the system sets dynamic sliding time windows (e.g., the past 1 minute or 5 minutes) to continuously assess whether various metrics exceed preset safety thresholds and remains highly sensitive to surges in error logs, thus providing precise triggering criteria for subsequent automated interventions.
[0070] When the error log collection pipeline captures an exception, the system's intelligent diagnostic module immediately intervenes to perform deep semantic analysis of the error information. To accurately identify logical conflicts, the captured exception stack and error code are matched against a predefined fault feature library. For example, if specific keywords or exception classes such as "circular dependency," "deadlock," "field type mismatch," or "conditional branch mutual exclusion" are detected, they can be characterized as serious logical conflicts.
[0071] Upon confirming a rule logic conflict, the system will immediately initiate an emergency recovery process. Based on a pre-established rule version timeline (which records metadata, release times, and snapshot storage paths for all historical versions), it will automatically locate the previous stable version of the current target rule version. Subsequently, a forced switchover command will be issued by calling the hot update interface of the configuration center or control plane. To ensure business continuity during the switchover process, the system can employ double-buffered loading or atomic replacement techniques to seamlessly redirect the data stream being processed to the execution container of the previous version, ensuring that the service recovers to a safe state within seconds.
[0072] While rolling back the version, the system synchronously corrects the rule's lifecycle metadata to prevent subsequent scheduling from triggering new versions with defects. Specifically, the system retrieves the rule's effective time window record for the previous version from the underlying storage and attaches a high-priority "emergency patch flag" to it. This flag not only records the timestamp of the rollback, the cause of the triggered exception, and the associated error log ID, but also locks the validity of the time window. During any future canary rollout or full release of a new version, the scheduling engine will forcibly verify this flag to ensure that versions with serious logical conflicts are completely isolated, thus achieving closed-loop fault defense and version governance.
[0073] The method provided in this invention, by constructing a rule effective time window, performs precise slicing calculations on cross-version data streams, effectively ensuring the consistency and accuracy of long-cycle data processing. Introducing an isolated sandbox environment to fully recalculate historical data and generate a difference report enables security assessment and risk prediction before new rules go live, avoiding blind deployment failures in the production environment. Simultaneously, the system combines the difference report with compliance results to execute intelligent dynamic routing, supporting automatic circuit breaking, degradation, and dead-letter queue interception for anomalies, significantly improving the system's fault tolerance. Furthermore, in conjunction with a canary release strategy and runtime monitoring rollback mechanism, not only is a smooth transition of rules and fine-grained traffic control achieved, but the system's self-healing capabilities and overall operational stability are also significantly enhanced.
[0074] In addition to the data stream processing method based on multi-version rules provided in the foregoing embodiments, this invention also provides a data stream processing apparatus based on multi-version rules, see [link to relevant documentation]. Figure 2 The diagram shown illustrates the structure of a data stream processing device based on multi-version rules, indicating that the device mainly comprises the following parts: Version update module 201 is used to respond to system rule update instructions, generate target rule versions, and construct rule effective time windows in the rule version timeline.
[0075] The target state parameter calculation module 202 is used to obtain the timestamp interval of the data stream to be processed, and calculate the target state parameters of the data stream to be processed based on the timestamp interval and the rule effective time window.
[0076] Sandbox simulation module 203 is used to perform a full recalculation of the historical state dataset in an isolated sandbox computing environment based on the target rule version, and generate a difference impact report.
[0077] The routing processing module 204 is used to perform routing distribution processing on the data stream to be processed based on the difference impact report.
[0078] The data stream processing device based on multi-version rules provided in this embodiment of the invention achieves refined management of the rule lifecycle by constructing a rule effective time window; it improves the problem of state calculation for cross-version data streams by matching and calculating the timestamp interval with the rule effective time window; it introduces an isolated sandbox environment to perform a full recalculation (regression test) of historical data and generate a difference impact report, ensuring the security and predictability of new rule deployment; and it effectively avoids business risks caused by rule changes by combining the difference report with dynamic routing of compliance results.
[0079] In one implementation, the version update module 201 is specifically used to: obtain the canary release strategy carried in the system rule update instruction; determine the official effective period of the target rule version according to the canary release strategy on the rule version timeline; and construct the rule effective time window based on the official effective period.
[0080] In one implementation, the target state parameter calculation module 202 is specifically used to: determine whether the timestamp interval of the data stream to be processed crosses the boundary of the rule effective time window; if the timestamp interval crosses the boundary of the rule effective time window, then divide the data stream to be processed into a first data sub-stream and a second data sub-stream; wherein, the first data sub-stream corresponds to the old rule version, and the second data sub-stream corresponds to the target rule version; apply the corresponding rule version to the first data sub-stream and the second data sub-stream respectively to perform state calculation, and merge the calculation results to generate the target state parameter of the data stream to be processed.
[0081] In one implementation, the sandbox simulation module 203 is specifically used to: construct a virtual execution container based on the target rule version; load a snapshot of the historical state dataset from persistent storage and inject it into the virtual execution container; execute the logical operation of the target rule version in parallel within the virtual execution container to obtain the operation results of each historical state data, and compare the operation results with the original state in the historical state dataset to generate a difference impact report.
[0082] In one implementation, the difference impact report includes the compliance verification result of the data stream to be processed; the sandbox simulation module 203 is specifically used to: if the compliance verification result of the data stream to be processed is passed, then route the data stream to be processed to the high-speed processing channel; if the compliance verification result of the data stream to be processed is failed, then trigger the manual review process, and temporarily route the data stream to be processed to the dead letter queue or the delayed processing channel until the review pass instruction is received.
[0083] In one embodiment, the above-mentioned device further includes a monitoring module, which is used to: monitor the resource consumption indicators and error logs of the target rule version when processing the data stream to be processed in real time; when the error type in the error log is detected to be a rule logic conflict, automatically roll back to the previous version of the target rule version based on the rule version timeline, and add an emergency patch mark to the rule effective time window corresponding to the previous version.
[0084] In one implementation, the version update module 201 is specifically used to: parse the system rule update instruction and extract the rule change difference set; identify the affected dependent data fields based on the rule change difference set; perform a directed acyclic graph topological sort on the target rule version to determine the rule execution order; and generate an index optimization strategy corresponding to the dependent data fields while generating the target rule version, and apply the index optimization strategy to the reading process of the historical state dataset.
[0085] It should be noted that the device provided in this embodiment of the invention has the same implementation principle and technical effects as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment. The specific numerical values provided in this embodiment are merely exemplary and are not intended to limit the scope of the invention.
[0086] This invention also provides an electronic device, specifically, the electronic device includes a processor and a storage device; the storage device stores a computer program, and the computer program, when run by the processor, executes the method described in any of the above embodiments.
[0087] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 30, a memory 31, a bus 32 and a communication interface 33. The processor 30, the communication interface 33 and the memory 31 are connected through the bus 32. The processor 30 is used to execute executable modules, such as computer programs, stored in the memory 31.
[0088] The memory 31 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 33 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc.
[0089] Bus 32 can be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 3 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0090] The memory 31 is used to store programs. After receiving an execution instruction, the processor 30 executes the program. The method executed by the device for defining the flow process disclosed in any of the foregoing embodiments of the present invention can be applied to the processor 30 or implemented by the processor 30.
[0091] Processor 30 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 30 or by instructions in software form. Processor 30 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 31. The processor 30 reads the information in memory 31 and, in conjunction with its hardware, completes the steps of the above method.
[0092] The computer program product of the readable storage medium provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, please refer to the foregoing method embodiments, which will not be repeated here.
[0093] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0094] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A data stream processing method based on multi-version rules, characterized in that, include: In response to system rule update commands, generate the target rule version and construct the rule effective time window in the rule version timeline; Obtain the timestamp range of the data stream to be processed, and calculate the target state parameters of the data stream to be processed based on the timestamp range and the rule effective time window; Based on the target rule version, a full recalculation of the historical state dataset is performed in an isolated sandbox computing environment to generate a difference impact report; Based on the difference impact report, routing and distribution processing is performed on the data stream to be processed.
2. The method according to claim 1, characterized in that, Construct rule effective time windows in the rule version timeline, including: Obtain the canary release strategy carried in the system rule update instruction; On the rule version timeline, the official effective date of the target rule version is determined according to the canary release strategy; The effective time window for rules is constructed based on the aforementioned official effective period.
3. The method according to claim 1, characterized in that, Based on the timestamp interval and the rule effective time window, the target state parameters of the data stream to be processed are calculated, including: Determine whether the timestamp interval of the data stream to be processed crosses the boundary of the rule's effective time window; If the timestamp interval crosses the boundary of the rule's effective time window, the data stream to be processed is divided into a first data sub-stream and a second data sub-stream; wherein, the first data sub-stream corresponds to the old rule version, and the second data sub-stream corresponds to the target rule version; The corresponding rule versions are applied to the first data sub-stream and the second data sub-stream respectively to perform state calculations, and the calculation results are merged to generate the target state parameters of the data stream to be processed.
4. The method according to claim 1, characterized in that, Based on the target rule version, a full recalculation of the historical state dataset is performed in an isolated sandbox computing environment to generate a difference impact report, including: Construct a virtual execution container based on the target rule version; A snapshot of the historical state dataset is loaded from persistent storage and injected into the virtual execution container; The logical operations of the target rule version are executed in parallel within the virtual execution container to obtain the operation results of each historical state data. The operation results are then compared with the original states in the historical state dataset to generate a difference impact report.
5. The method according to claim 1, characterized in that, The difference impact report includes the compliance verification results of the data stream to be processed; Based on the aforementioned difference impact report, routing and distribution processing is performed on the data stream to be processed, including: If the compliance verification result of the data stream to be processed is passed, the data stream to be processed will be routed to the high-speed processing channel. If the compliance verification result of the data stream to be processed is unsuccessful, a manual review process is triggered, and the data stream to be processed is temporarily routed to a dead letter queue or a delayed processing channel until a review approval instruction is received.
6. The method according to claim 1, characterized in that, After performing routing and distribution processing on the data stream to be processed, the process further includes: Real-time monitoring of the resource consumption metrics and error logs of the target rule version when processing the data stream to be processed; When the error type in the error log is detected to be a rule logic conflict, the system automatically rolls back to the previous version of the target rule based on the rule version timeline, and adds an emergency patch marker to the rule effective time window corresponding to the previous version.
7. The method according to claim 1, characterized in that, Generate a version of the target rule, including: Parse the system rule update command and extract the rule change difference set; Based on the rule change difference set, identify the affected dependency data fields; Perform a directed acyclic graph topological sort on the target rule version to determine the rule execution order; While generating the target rule version, an index optimization strategy corresponding to the dependent data field is generated, and the index optimization strategy is applied to the reading process of the historical state dataset.
8. A data stream processing device based on multi-version rules, characterized in that, include: The version update module is used to respond to system rule update commands, generate the target rule version, and construct the rule effective time window in the rule version timeline; The target state parameter calculation module is used to obtain the timestamp interval of the data stream to be processed, and calculate the target state parameters of the data stream to be processed based on the timestamp interval and the rule effective time window. The sandbox simulation module is used to perform a full recalculation of the historical state dataset in an isolated sandbox computing environment based on the target rule version, and generate a difference impact report. The routing processing module is used to perform routing distribution processing on the data stream to be processed based on the difference impact report.
9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions executable by the processor, the processor executing the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method described in any one of claims 1 to 7.