Multi-mode complex event detection and optimization method for semi-structured data

Through the multi-modal regular tree method, it is disassembled into tree mode and regular mode, and combined with TwigList and NFA algorithm, the problem of low efficiency of multi-modal complex event matching of semi-structured data in massive data is solved, and efficient event pattern matching and query planning optimization is achieved.

CN119988681APending Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510067730.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the case of massive data, it is difficult for the prior art to efficiently identify and optimize multiple complex event patterns for semi-structured data, especially in scenarios where the data patterns are not fixed, the structure is complex, and the event triggering behavior is difficult to predict.

Method used

The multi-modal regular tree method is adopted to disassemble the regular tree pattern into two parts: tree pattern and regular pattern, tree pattern matching is achieved through the TwigList algorithm, and regular pattern matching is achieved using NFA. Combined with the dynamically expanded multi-modal matching mechanism, the query plan is optimized to reduce the use of memory resources.

Benefits of technology

It improves the matching efficiency of complex event patterns in semi-structured data, optimizes query plans, reduces system costs, and can effectively handle multiple modes of interest in massive data streams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988681A_ABST
    Figure CN119988681A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode complex event detection and optimization method for semi-structured data, and relates to the field of data stream processing. According to the method, mode matching of the semi-structured data is realized by using a multi-mode normal tree. In order to enable a regular tree mode to be matched with a semi-structured data stream, the patent adopts a method of disassembling the regular tree mode into two parts, namely a tree mode and a regular mode, and realizing matching step by step. The tree mode designs a node relationship in which a user is interested into a logic tree, and a TwigList algorithm is used for matching data meeting conditions in a data stream and storing the data. The regular pattern defines a behavior sequence in which a user is interested, and state transition is realized by using NFA. And carrying out NFA processing on the data stream which is successfully matched with the tree pattern to obtain a final result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data stream processing, and in particular to a multi-mode complex event matching and optimization method for semi-structured data (such as Json, XML), which is used for real-time recognition of complex event patterns. Background Art

[0002] Query optimization is an important part of CEP system optimization. In scenarios where semi-structured or nested data (such as JSON) need to be processed, query plan design and optimization of CEP become particularly important. Especially when the data pattern is not fixed, the structure is complex, and the event triggering behavior is difficult to predict, how to efficiently process and optimize the detection of complex JSON events becomes a core research issue. In order to cope with massive data streams, users may be interested in multiple behavior patterns at the same time, and corresponding optimizations need to be made for the complexity of multi-mode detection. Complex event processing (CEP) is mainly used to extract information or behavior patterns of interest to users from large amounts of event data streams generated in real time. It is usually combined with an event-driven architecture to enable the system to give timely feedback when user-defined patterns occur. Due to their powerful and expressive query languages ​​and performance potential, they are becoming increasingly popular in many fields, including financial services, electronic health record systems, sensor networks, and the Internet of Things.

[0003] The present invention proposes a multi-mode complex event detection for semi-structured data and optimizes the query plan, which fills the gap in existing research on detecting multiple modes in semi-structured data. Summary of the invention

[0004] The purpose of the present invention is to propose a method for multi-mode complex event detection and optimization for semi-structured data, while achieving matching of specific nested relationships of semi-structured data and matching of specific behavior sequences by traditional complex event processing. The problem that there are multiple patterns of interest to users of semi-structured data that need to be identified simultaneously in the current massive data situation is solved, and the efficiency of event matching is improved. The present invention uses a multi-mode regular tree to achieve pattern matching of semi-structured data. In order to make the regular tree pattern match the semi-structured data stream, this method adopts a method of decomposing the regular tree pattern into two parts, a tree pattern + a regular pattern, and implementing the matching step by step. The tree pattern will design the node relationship of interest to the user into a logical tree, and use the TwigList algorithm to match and store the qualified data in the data stream. The regular pattern defines the behavior sequence of interest to the user, and uses NFA to achieve state transfer. The data stream that successfully matches the tree pattern is processed by NFA to obtain the final result.

[0005] In order to achieve the above object, the technical solution adopted by the present invention consists of the following steps:

[0006] (1) Multi-pattern regular tree matching semi-structured events

[0007] The regular tree model is a hybrid model of regular expression and tree structure. It supports the structural constraints and predicate constraints of semi-structured data, and uses regular expressions to describe the constraint relationship between consecutive events in the event stream. This method proposes a formal expression of multi-mode regular trees, and expands the expressive power of regular tree models to adapt to the current scenario of responding to massive events. The matching principle of regular tree models is defined to ensure the correctness of the final result.

[0008] (2) Propose a system-wide cost model to support optimization

[0009] According to the proposed regular tree pattern matching mechanism, the consumption of the entire system is considered from multiple dimensions of comparison, construction, and expiration and removal costs. The cost models of the tree structure matching part and the regular matching part are given respectively. This provides theoretical support for the subsequent query plan optimization. A multi-weight matching mechanism is designed to optimize the query plan to achieve the purpose of reducing the use of memory resources. This method comprehensively considers the impact of operator selectivity, predicate selectivity, and event arrival rate on event processing and proposes a cost model for query plan optimization.

[0010] (3) Dynamically Expanded Multi-Pattern Matching Mechanism

[0011] The present invention adopts the method of dynamic disassembly + reconstruction, presents a new perspective, and breaks through the limitations of static structure. It will use the cost model proposed above to evaluate the shared state and path, and re-evaluate whether these shares are reasonable according to the newly added mode. When it is found that the current shared state uses a large total cost, this method allows the shared path to be dynamically split into independent paths. This ensures that the reconstructed NFA achieves the global optimal performance.

[0012] Step 1: Initialize the multi-modal normal tree.

[0013] The workload of the multi-mode regular tree is expressed as A single normal tree pattern can be represented as a four-tuple P Ti =(E i ,R i ,P i ,W i ), where E i =E1,···,E n} is the regular tree pattern All event types included in R i =

[0014] R1,···,R n}yes All relations contained in R i ={(E i / E j ),(E i / / E j ),(E i / P i ),(E i / / P i ),(P i / E i ),(P i / / E i )}, the relational operators / and / / respectively represent the parent-child relationship and grandparent-grandchild relationship in the regular tree structure. They are binary operators and have sequentiality. i =

[0015] {P1,···,P n} is a common pattern contained in the regular tree pattern, each pattern can be represented as a four-tuple P i ={E i ,S i ,C i ,W i}. C i A set of clauses that constrain the attribute values ​​of an event. and the normal pattern P contained in the regular tree pattern i ShareE i ,W i .W i is the time window defined for this pattern.

[0016] S i Specifies how to combine the events requested by the pattern to form a match. It is defined by a combination of event type and operator. In this invention, the most common operators such as AND, SEQ, and OR will be considered.

[0017] Step 2: Design a regular tree pattern cost model

[0018] The meanings of the symbols used in the query cost of the regular tree mode are shown in Table 1:

[0019] Table 1 Symbols and meanings of cost model

[0020]

[0021]

[0022] Step 2.1: Tree pattern matching cost

[0023] If the traditional method is used to query the tree structure, it may be necessary to completely traverse all parent-child nodes and nested relationships of the event, resulting in a large number of redundant operations and resource waste. When performing layer-by-layer parsing, each nested layer needs to be indexed separately or multiple query operations need to be performed. The present invention adopts the Twiglist algorithm, so that the algorithm complexity increases linearly with the increase of the number of pattern occurrences, and the space complexity increases linearly with the increase of the number of pattern occurrences.

[0024] Step 2.1.1: The main task of TwigList-Construct (explanation) is to build a node list for subsequent query operations. It is known that the input Json data contains |X| nodes, and the algorithm needs to traverse these nodes to build different types of node lists.

[0025] L V1 ,L V2 ,…,L Vn During the traversal, each node needs to be visited once. The depth d of the tree affects the processing time of each node, especially when dealing with nested structures. For each node, some operations may need to be performed on the path from the root node to the node (for example, updating the pointer of the parent node or child node), so its time complexity is:

[0026] O(d·|X|).

[0027] Step 2.1.2: The main task of TwigList-Enumerate is to extract all matching n-tuples from the constructed node list. Its time complexity is calculated as follows: Each time the moreMatch function is called, the algorithm finds a new match and adds it to the result set T. Assume that there are a total of |T| matching results.

[0028] Each time a match is found, the operation inside the moreMatch function needs to traverse the structure associated with each node in the query tree, which usually involves checking whether the current node meets the query conditions, updating pointers, etc. Therefore, when each match is found, the operation time complexity of the moreMatch function is O(n) because it needs to process n nodes in the query tree. Based on the above analysis, it can be concluded that the time complexity of TwigList-Enumerate is: O(n·|T|).

[0029] Step 2.1.3: Calculate the tree pattern matching cost. |X| is the total number of nodes in the semi-structured data input stream (such as json stream), and d is the maximum degree of the nodes in the query tree Q. Where |R| is the total number of query pattern matches and n is the number of nodes in the query tree.

[0030] So for tree query, assuming the cost of constructing a node using TwigList-Construct is C T-C , we can infer that the cost of constructing a node list for the input Json stream is as shown in formula (1).

[0031] C T-C ·(d·|X|) (1)

[0032] Assume that the cost of querying a node list is C T-E , then the query cost for the query tree Q is shown in formula (2).

[0033] C T-E ·(n·|T|) (2)

[0034] In summary, the partial matching cost of matching the query tree Q in the json stream S is shown in formula (3).

[0035] Cost TwigList (S,Q)=C T-C ·(d·|X|)+C T-E ·(n·|T|) (3)

[0036] Step 2.2: Regular expression partial matching cost

[0037] Since the logic of NFA is very intuitive, complex regular expressions can be directly expressed through simple state descriptions and edge connections. Therefore, the regular expression partial matching of the present invention is implemented using NFA. This part of the cost is mainly formed by three parts of the cost of creating partial matches in NFA, clearing expired partial matches, and comparing events with partial matches in NFA, namely:

[0038] (Cost-EST)+(Cost-CLN)+(Cost-CMP)

[0039] Step 2.2.1: Create the cost of the partial match (Cost-EST).

[0040] For a state k in NFA, the number of partial matches created and maintained is shown in formula (4).

[0041]

[0042] The maintenance cost of a regular pattern depends on the NFA state that defines the pattern and the number of partial matches that need to be created in each state. Therefore, the maintenance cost calculation steps for a single regular pattern are shown in formulas (5) to (6).

[0043] Cost-EST=CountP*C e (5)

[0045]

[0046] Step 2.2.2: The comparison cost between partial matching and arrival events is shown in formula (7).

[0047]

[0048] Let the single comparison cost of a single event and a partial match be C i , the number of comparisons between two adjacent states is the product of the number of partial matches maintained by the previous state and the number of events selected by the predicate in the desired event type.

[0049] It can be seen that the comparison cost between states M and N is shown in formula (8).

[0050]

[0051] Step 2.2.3: Clear the expired portion of the matching cost

[0052] After the pattern window expires, the expired partial matches are cleared. The cost of this part is shown in formula (9).

[0053] Cost-CLN=Count(inst) timeout ·C d (9)

[0054] Assume that the time window slides by 1 second, and all partial matches in the window composed of the previous second are cleared. From the formula for creating partial matches, it can be seen that the calculation process of the number of partial matches cleared in each state is shown in formulas (10) to (11).

[0055]

[0056] Step 2.2.4: The total cost of the NFA matching pattern can be calculated as shown in formula (12).

[0057]

[0058] Step 3: Optimize multi-mode query plan. For multi-mode optimization of semi-structured data, this method starts from two parts: tree pattern sharing and regular optimization.

[0059] Step 3.1: Multiple tree patterns share the maximum prefix.

[0060] Step 3.1.1: There are certain common parts among multiple tree patterns, that is, these logical trees may share some parent-child relationships and grandparent-grandchild relationships. According to the TwigList-Enumerate algorithm, the query cost of each single relationship is consistent. This method designs multiple tree patterns to share common relationships and form a new tree pattern. By merging repeated relationships, the number of queries is greatly reduced to achieve the purpose of reducing system costs.

[0061] Step 3.2: Optimization of the regularization part.

[0062] Step 3.2.1: For each event type E, use the arrival rate and selection rate to calculate its importance. According to existing research, the earlier low-probability events are filtered by the system, the fewer intermediate results are generated and the less burden on the system. Therefore, this method designs formula (13) to calculate the importance of each event type.

[0063] IMP(E)=r E ×sel E (13)

[0064] Step 3.2.2: According to the event importance formula, the smaller the probability of an event, the smaller the corresponding IMP(E) value. Therefore, this method designs all event types to be sorted in ascending order according to the event importance value to generate an event importance list ImpList as shown in formula (14), so as to generate the initial query plan later and ensure that low-probability events are filtered as early as possible.

[0065] ImpList={Ea:IMPa,Eb:IMPb,Ec:IMPc,…} (14)

[0066] Step 3.2.3: Generate an initial query plan QPlan1 through the event importance list. The initial query plan QPlan1 is a query plan generated from low to high according to the importance of the events, which can give priority to matching low-probability events, thereby reducing the generation of intermediate results.

[0067] Step 3.2.4: Generate a sub-pattern set SubPi for all regular patterns Pi contained in the multi-pattern, and re-sort the events in the sub-pattern according to the event importance list ImpList as shown in formula (15), then find the common sub-pattern as shown in formula (16), and sort the common sub-patterns. Sort the sub-patterns by length as shown in formula (17), that is, longer sub-patterns are shared first and the cost is calculated (the longer the shared sub-pattern, the smaller the total system cost).

[0068] For example, for the patterns P1 = AND (A, B, C, D), P2 = AND (B, D, F, G), and P3 = AND (A, B, D, E), the maximum common sub-pattern is P = AND (B, D). Assume that the generated event importance list ImpList is

[0069] ImpList = {A: 0.08, D: 0.21, C: 0.3, B: 0.56, G: 0.71, E: 0.8, F: 0.88}, it can be seen that the probability of D type event in the common sub-pattern P = AND (B, D) is lower than that of B type event. Therefore, the common sub-pattern P = AND (B, D) is adjusted to P new =AND(D, B), ensuring that low-probability events are filtered out as early as possible.

[0070] SubP(Pi)'=Sort(E∈SubP(Pi),key=IMP(E)) (15)

[0071] CSubP=SubP(P1)∩SubP(P2)∩...∩SubP(Pn) (16)

[0072] CSubP'=Sort(SubP(Pi)∈CSubP,key=Length(SubP)) (17)

[0074] Step 3.2.5: Based on the sorted common sub-pattern set, take the first order CsubP[0] to generate the query plan QPlan2 and calculate its execution cost.

[0075] Step 3.2.6: Compare the execution costs of QPlan1 and QPlan2, and select the plan with lower execution cost as the optimal plan as shown in formula (18);

[0076]

[0077] Step 3.2.7: Traverse the remaining common sub-patterns, generate a new query plan and calculate the system cost. If the system cost is less than the cost of the current optimal plan, update the optimal query plan;

[0078] Step 3.2.8: Output the optimal query plan finally generated. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 System architecture diagram of the present invention.

[0080] Figure 2 Example diagram of a regular tree pattern.

[0081] Figure 3 Regular tree pattern matching flow chart.

[0082] Figure 4 Comparison results of optimization algorithms. DETAILED DESCRIPTION

[0083] According to the above description, the following is a specific implementation process.

[0084] Step 1: Initialize two normal trees according to the defined normal tree pattern as shown in Table 2 and Table 3 respectively.

[0085] Table 2 Formal expression of regular tree pattern 1

[0086]

[0087]

[0088] Table 3 Formal expression of regular tree pattern 2

[0089]

[0090] Step 2: Design a regular tree pattern cost model

[0091] There are 1000 json objects input by users within the initialization time window of 1s, the average number of nodes in each object is 100, and the average depth of the json tree is 8. The final output is 18 objects that conform to regular tree mode 1 (hereinafter referred to as mode 1) and 30 objects that conform to regular tree mode 2 (hereinafter referred to as mode 2). For the regular part in mode 1, the arrival rate of D type events is 30 / s, the arrival rate of E type events is 50 / s, the arrival rate of G type events is 20 / s, and the arrival rate of H type events is 60 / s. The selection rate of operator And for D type events is 90%, for E type events is 40%, for G type events is 50%, and for H type events is 10%. Since there is no predicate constraint between the defined events, the predicate selection rate is 1.

[0092] Step 2.1: Tree pattern matching cost

[0093] Calculate the tree pattern matching cost. The total number of nodes in the semi-structured data input stream (json stream) |X|=100*100, and the maximum degree of the nodes in the query tree Q is d=8. The total number of query pattern 1 matches |T1|=18, and the total number of query pattern 2 matches |T2|=30. From Table 2 and Table 3, we can see that the number of nodes in the query tree corresponding to pattern 1 and pattern 2 is 6 and 7 respectively.

[0094] Therefore, for tree queries, it can be deduced that the cost of constructing a node list for the input Json stream is:

[0095] C T-C*(8*100*100)=80000*C T-C

[0096] The query cost for mode 1 and mode 2 query trees is

[0097] C T-E* (18*6)+C T-E *(30*7)=310*C T-E

[0098] In summary, the cost of matching the tree pattern portion of pattern 1 and pattern 2 in the initialized json stream is:

[0099] Cost TwigList (S,Q)=80000*C T-C +310*C T-E

[0100] Step 2.2: Regular expression partial matching cost

[0101] Since the logic of NFA is very intuitive, complex regular expressions can be directly expressed through simple state descriptions and edge connections. The regular pattern part contained in Mode 1 and Mode 2 is implemented using NFA. The cost of this part is mainly composed of three parts: creating partial matches in NFA, clearing expired partial matches, and comparing events with partial matches in NFA, namely:

[0102] (Cost-EST)+(Cost-CLN)+(Cost-CMP)

[0103] For the NFA constructed in mode 1, it consists of four states: the initial state S1 receives a D-type event that meets the conditions and then transfers to the S2 state; the state S2 receives an E-type event that meets the conditions and then transfers to the S3 state; the state S3 receives a G-type event that meets the conditions and then transfers to the S4 state. S4 is the final state, indicating that the pattern matching is completed. The NFA constructed in mode 2 can be deduced by analogy.

[0104] Step 2.2.1: Create the cost of the partial match (Cost-EST).

[0105] For mode 1, the cost of creating partial matches for the four states of the constructed NFA (the initial state does not require the creation of partial matches, so the cost is not calculated) can be calculated based on the initialized data:

[0106]

[0107] Substituting into formula (4), we can get the cost of creating a partial match for mode 1 as:

[0108] (30*0.9+30*0.9*50*0.4+30*0.9*50*0.4*20*0.5)*C e =5967*C e

[0109] Similarly, the cost of creating a partial match for mode 2 is calculated as:

[0110] (30*0.9+30*0.9*20*0.5+30*0.9*20*0.5*60*0.1)*C e =1917*C e

[0111] Step 2.2.2: Comparison cost of partial matches and arrival events.

[0112] The comparison cost of mode 1 is divided into three parts: the comparison cost of S1 and S2, S2 and S3, and S3 and S4. According to formula (8), we can get:

[0113] (30+30*0.9*50+30*0.9*50*0.4*20)*C i =12180*C i

[0114] Similarly, the comparison cost of mode 2 is calculated as:

[0115] (30+30*0.9*20+30*0.9*20*0.5*60)*C i =16770*C i

[0116] Step 2.2.3: Clear the expired portion of the matching cost.

[0117] After the pattern window expires, the expired partial matches will be cleared. If the time window slides by 1 second, all partial matches in the window consisting of the previous second will be cleared.

[0118] The clearing partial matching cost of mode 1 is calculated by formulas (9) to (11):

[0119] 0.1*(30*0.9+27*50*0.4+540*20*0.5)*C d =596*C d

[0120] Similarly, the cost of clearing partial matches in mode 2 is:

[0121] 0.1*(30*0.9+27*20*0.5+270*60*0.1)*C e =191*C e

[0122] Step 2.2.4: Total cost of NFA matching pattern

[0123] According to formula (12), the sum of the NFA partial costs of mode 1 and mode 2 is calculated as:

[0124] 7884*C e +28950*C i +788*C d

[0125] 7290*C e +8500*C i +729*C d

[0126] Step 3: Optimize multi-mode query plan. For multi-mode optimization of semi-structured data, this method starts from two parts: tree pattern sharing and regular optimization.

[0127] Step 3.1: Multiple tree patterns share the maximum prefix.

[0128] Step 3.1.1: There are certain common parts among multiple tree patterns, that is, these logical trees may share some parent-child relationships and grandparent-grandchild relationships. According to the TwigList-Enumerate algorithm, the query cost of each single relationship is consistent. This method designs multiple tree patterns to share common relationships and form a new tree pattern. By merging repeated relationships, the number of queries is greatly reduced to achieve the purpose of reducing system costs.

[0129] For the tree mode part of mode 1 and mode 2, the relationship that can be shared by this method is: share ={A / B,A / / C}. After sharing, the query cost of the tree mode part is:

[0130] C T-E *(18*6)+C T-E *(12*7)=192*C T-E

[0131] It can be seen that compared with the query cost before sharing 310*C T-E , the system cost of the tree mode part is reduced by about 40%.

[0132] Step 3.2: Optimization of the regularization part.

[0133] Step 3.2.1: For each event type E, use the arrival rate and selection rate to calculate its importance. According to existing research, the earlier low-probability events are filtered by the system, the fewer intermediate results are generated and the less burden on the system. According to formulas (13) to (14), the event importance list can be obtained: ImpList =

[0134] Eh:6,Eg:10,Ee:20,Ed:27.

[0135] Step 3.2.3: Generate the initial query plan QPlan1 through the event importance list and calculate the system cost corresponding to the query plan. The initial query plan QPlan1 is a sequential plan generated from low to high according to the importance of the events, which can give priority to matching low-probability events, thereby achieving the purpose of reducing intermediate results.

[0136] For modes 1 and 2, QPlan1 is obtained as And(G,E,D) and And(H,G,D). The total NFA cost corresponding to And(G,E,D) is calculated as: 5610*C e +6520*C i +561*C d The total cost of the NFA corresponding to And(H,G,D) is: 1686*C e +1980*C i +168*C d .

[0137] Therefore, the total cost of QPlan1 is: 7290*C e +8500*C i +729*C d .

[0138] Step 3.2.4: Generate a sub-pattern set SubPi for all regular patterns Pi contained in the multi-pattern, and re-sort the events in the sub-pattern according to the event importance list ImpList as shown in formula (15), then find the common sub-pattern according to formula (16), and sort the common sub-patterns, and sort the sub-patterns according to the length as shown in formula (17), that is, the longer sub-patterns are shared first and the cost is calculated (the longer the shared sub-pattern, the smaller the total cost of the system). Calculation can obtain CSubP = {(G, D), (G), (D)}.

[0139] Step 3.2.5: Based on the sorted common sub-pattern set, take the first order CSubP[0] to generate the query plan QPlan2, and calculate the cost: 7300*C e +30020*C i +730*C d

[0140] Step 3.2.6: Compare the execution costs of QPlan1 and QPlan2, and select the plan with the lower execution cost as the optimal plan. The above calculation shows that QPlan1 is the optimal query plan in this example.

[0141] Step 3.2.7: Traverse the remaining common sub-patterns, generate a new query plan and calculate the system cost. If the system cost is less than the cost of the current optimal plan, update the optimal plan. The calculated QPlan1 is the optimal query plan, and its cost is: 7290*C e +8500*C i +729*C d Compared with the query plan before optimization, the cost of creating partial matches and clearing partial matches was reduced by 7.5 percentage points, and the cost of comparing partial matches with arrival events was reduced by 70 percent.

[0142] Step 3.2.8: Output QPlan1 as the final optimal query plan.

[0143] In summary, the present invention fills the gap in existing research on the matching of multiple patterns for semi-structured data, and optimizes the query plan of the regular pattern by using the proposed cost model.

Claims

1. A multi-modal complex event detection and optimization method for semi-structured data, characterized in that: include: Step 1: Initialize the multi-mode regular tree; The workload of the multi-mode regular tree is expressed as A single normal tree pattern is represented as a four-tuple Where E i ={E1,···,E n } is the regular tree pattern All event types included in R i ={R1,···,R n }yes All relations contained in R i ={(E i / E j ),(E i / / E j ),(E i / P i ),(E i / / P i ),(P i / E i ),(P i / / E i )}, the relational operators / and / / respectively represent the parent-child relationship and grandparent-grandchild relationship in the regular tree structure, which are binary operators and have sequentiality; i ={P1,···,P n } is a common pattern contained in the regular tree pattern, each pattern is represented by a four-tuple P i ={E i ,S i ,C i ,W i }; C i Refers to the set of constraints on the attribute values ​​of an event; regular tree mode and the normal pattern P contained in the regular tree pattern i ShareE i ,W i ; W i is the time window defined for this mode; S i Specifies how the events requested by the pattern are combined to form a match, as defined by the combination of event type and operator; Step 2: Design a regular tree pattern cost model; The meanings of the symbols used in the query cost of the regular tree mode are shown in Table 1: Table 1 Symbols and meanings of cost model Step 3: Optimize multi-mode query plans. For multi-mode optimization of semi-structured data, start with tree mode sharing and regular optimization.

2. The method for multi-modal complex event detection and optimization of semi-structured data according to claim 1, characterized in that: The implementation process of step 2 is as follows: Step 2.1: Tree pattern matching cost The Twiglist algorithm is used, so that the algorithm complexity increases linearly with the number of times the pattern appears, and the space complexity increases linearly with the number of times the pattern appears; Step 2.1.1: The main task of TwigList-Construct is to build a node list for subsequent query operations; it is known that the input Json data contains |X| nodes, and the algorithm needs to traverse these nodes to build different types of node lists L V1 ,L V2 ,…,L Vn ; During the traversal process, each node needs to be visited once; the depth d of the tree affects the processing time of each node, especially when dealing with nested structures; for each node, operations need to be performed on the path from the root node to the node; Step 2.1.2: The main task of TwigList-Enumerate is to extract all matching n-tuples from the constructed node list; the time complexity is calculated as follows. Each time the moreMatch function is called, the algorithm finds a new match and adds it to the result set T; assuming that there are a total of |T| matching results; When finding each match, the operation time complexity of moreMatch function is O(n), and the time complexity of TwigList-Enumerate is: O(n·|T|); Step 2.1.3: Calculate the tree pattern matching cost; |X| is the total number of nodes in the semistructured data input stream, d is the maximum degree of the nodes in the query tree Q; where |R| is the total number of query pattern matches, and n is the number of nodes in the query tree; For tree queries, the cost of constructing a node using TwigList-Construct is C T-C , the cost of constructing a node list for the input Json stream is shown in formula (1); C T-C ·(d·|X|) (1) Assume that the cost of querying a node list is C T-E , then the query cost for the query tree Q is shown in formula (2); C T-E ·(n·|T|) (2) In summary, the partial matching cost of matching the query tree Q in the json stream S is shown in formula (3); Cost TwigList (S,Q)=C T-C ·(d·|X|)+C T-E ·(n·|T|) (3) Step 2.2: Regularized partial matching cost; Regular expression partial matching is implemented using NFA. This part of the cost is formed by three parts of the cost: creating partial matching in NFA, clearing expired partial matching, and comparing events with partial matching in NFA, namely: (Cost-EST)+(Cost-CLN)+(Cost-CMP) Step 2.2.1: Create the cost of the partial match Cost-EST; For a state k in NFA, the number of partial matches created and maintained is shown in formula (4); The calculation steps of the maintenance cost of a single regular pattern are shown in formulas (5) to (6); Cost-EST=Count P *C e (5) Step 2.2.2: The comparison cost between partial matching and arrival events is shown in formula (7); Let the single comparison cost of a single event and a partial match be C i ,The number of comparisons between two adjacent states is the product of the number of partial matches maintained by the previous state and the number of events selected by the predicate in the required event type; It can be seen that the comparison cost between states M and N is shown in formula (8); Step 2.2.3: Clear the expired portion of the matching cost; After the pattern window expires, the expired part of the match is cleared. The cost of this part is shown in formula (9); Cost-CLN=Count(inst) timeout ·C d (9) Assume that the time window slides by 1 second, and all partial matches in the window composed of the previous second are cleared. From the formula for creating partial matches, it can be seen that the calculation process of the number of partial matches cleared in each state is shown in formulas (10) to (11). Count(inst) timeout =∑Count(K,P) timeout (11) Step 2.2.4: Calculate the total cost of the NFA matching pattern as shown in formula (12); 3. The multi-modal complex event detection and optimization method for semi-structured data according to claim 1, characterized in that: The implementation process of step 3 is as follows: Step 3.1: Multiple tree patterns share the maximum prefix; Step 3.1.1: There are certain common parts among multiple tree patterns, that is, these logical trees may share some parent-child relationships and grandparent-grandchild relationships. According to the TwigList-Enumerate algorithm, the query cost of each single relationship is consistent. Design multiple tree patterns to share common relationships to form a new tree pattern. By merging repeated relationships, the system cost can be reduced by reducing the number of queries. Step 3.2: Optimization of the regular part; Step 3.2.1: For each event type E, calculate its importance using the arrival rate and selection rate; Formula (13) is designed to calculate the importance of each event type; IMP(E)=r E ×sel E (13) Step 3.2.2: According to the event importance formula, the smaller the probability of an event, the smaller the corresponding IMP(E) value; all event types are sorted in ascending order according to the event importance values ​​to generate the event importance list ImpList as shown in formula (14), so as to generate the initial query plan later to ensure that low-probability events are filtered out as early as possible; ImpList={Ea: IMPa, Eb: IMPb, Ec: IMPC,...} (14) Step 3.2.3: Generate an initial query plan QPlan1 through the event importance list. The initial query plan QPlan1 is a query plan generated from low to high importance of events, so that low-probability events are matched first, thereby reducing the generation of intermediate results; Step 3.2.4: Generate a sub-pattern set SubPi for all regular patterns Pi contained in the multi-pattern, and re-sort the events in the sub-pattern according to the event importance list ImpList as shown in formula (15), then find the common sub-patterns as shown in formula (16), and sort the common sub-patterns, and sort the sub-patterns according to the length as shown in formula (17), that is, longer sub-patterns are shared first and the cost is calculated; For the patterns P1 = AND (A, B, C, D), P2 = AND (B, D, F, G), and P3 = AND (A, B, D, E), the maximum common subpattern is P = AND (B, D); The generated event importance list ImpList is: ImpList = {A: 0.08, D: 0.21, C: 0.3, B: 0.56, G: 0.71, E: 0.8, F: 0.88}, it can be seen that the probability of D type event in the common sub-pattern P = AND (B, D) is lower than that of B type event; Therefore, the common sub-pattern P=AND(B, D) is adjusted to P new =AND(D, B), ensuring that low-probability events are filtered out as early as possible; SubP(Pi)′=Sort(E∈SubP(Pi), key=IMP(E)) (15) CSubP=SubP(P1)∩SubP(P2)∩…∩SubP(Pn) (16) CSubP′=Sort(SubP(Pi)∈CSubP, key=Length(SubP)) (17) Step 3.2.5: Based on the sorted common sub-pattern set, take the first order CsubP[0] to generate the query plan QPlan2 and calculate its execution cost; Step 3.2.6: Compare the execution costs of QPlan1 and QPlan2, and select the plan with lower execution cost as the optimal plan as shown in formula (18); Step 3.2.7: Traverse the remaining common sub-patterns, generate a new query plan and calculate the system cost. If the system cost is less than the cost of the current optimal plan, update the optimal query plan; Step 3.2.8: Output the optimal query plan finally generated.

Citation Information

Cited By

  • Complex event processing method and system under intelligent scene of Internet of Things terminal

    CN120825459A