Graph flow tense query method based on FlaatMap line segment tree

By using FlatMap line segment tree structure and greedy query algorithm in graph stream temporal query, the problems of low query accuracy and large delay in the existing technology are solved, and efficient tense query of large-scale graph stream data is realized.

CN120045636AActive Publication Date: 2025-05-27NORTHEASTERN UNIV CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510115057.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The prior art has problems such as low query accuracy and large insertion and query delay in graph stream temporal query, especially when processing large-scale graph stream data, it is difficult to meet the needs of efficient query.

Method used

The graph stream tensile query method based on FlatMap segment tree is adopted. By converting the timestamp of the flow edge into the corresponding time period and performing one-time positioning and inserting in the storage stage, the target query range is efficiently decomposed to the specific interval that can be queried in the query stage, and querying any interval is realized.

Benefits of technology

It effectively reduces hash collisions, improves query accuracy, reduces insertion delay, and supports efficient tense query of large-scale graph stream data, suitable for graph stream scenarios with large data volume.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045636A_ABST
    Figure CN120045636A_ABST
Patent Text Reader

Abstract

The invention provides a FlatMap line segment tree-based graph stream tense query method, which belongs to the technical field of data processing and comprises the following steps of: finding a corresponding storage position in a graph stream abstract storage structure according to a target query edge; the graph flow abstract storage structure comprises a Hash compression matrix and a buffer area, and each cell in the Hash compression matrix comprises a fingerprint pair of flow edges and a group of FlatMap line segment trees; decomposing the target query range to corresponding intervals in the pre-updated FlaatMap line segment tree to obtain a queriable sub-interval set; according to a corresponding storage position and a queried sub-interval set found in the graph flow abstract storage structure, nodes corresponding to all queried sub-intervals are found in the pre-updated FlaatMap line segment tree, and the sum of weights stored in the found nodes serves as a graph flow tense query result. According to the method, Hash collision is reduced, query accuracy is improved, and query delay is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a graph stream temporal query method based on a FlatMap segment tree. Background Art

[0002] With the advent of the big data era, the data scale has been continuously expanding, and the form of data is no longer limited to traditional relational data, showing a diversified trend. As a new form of big data, graph streams have been widely applied in many fields such as network traffic, transaction systems, and social media. A graph stream is an infinite sequence of edge streams containing time information, and each item is represented as (s, d, w, t), representing an edge with weight w from node s to node d arriving at time t. These stream edges together form a continuously and dynamically changing stream graph, which can represent the connections or interactions between entities and play a role in discovering malicious attacks in network traffic data, detecting financial fraud in transaction systems, analyzing user behavior in social networks, etc.

[0003] In actual graph stream analysis, temporal query is a common and key issue. For example, we may encounter the following types of queries: "Between January 1, 2024 and January 31, 2024, has a certain IP accessed this host and how many bytes of data have been downloaded?" "What is the change in the connectivity of a certain website between this Monday and this Sunday?" "Between 22:00 and 23:00 today, has there been a direct or indirect transaction between two suspicious bank accounts?" All these types of questions are graph data queries within a specific time range (such as temporal edge weight queries, temporal boolean queries, etc.), which have important value in practical applications and require effective methods to store the topological structure and time-related information in graph streams and support efficient temporal queries.

[0004] Due to the extremely large amount of data in real-world graph streams, the cost of accurately and completely storing such a huge amount of graph stream data is unacceptable. Therefore, graph stream summaries based on approximate storage have been widely applied, aiming to achieve lower storage costs and support various queries related to graph topologies at the cost of sacrificing a small amount of accuracy. Currently, temporal queries are mainly implemented based on graph stream summary structures, and existing methods can be roughly divided into the following three categories:

[0005] (1) Directly utilize the timestamps on the stream edge for storage. This type of method usually directly saves the timestamps on the stream edge or directly uses them as calculation parameters for storage. For example, the TCM method (DOI: 10.1145 / 2882903.2915223) uses multiple m×m compressed adjacency matrices M to represent the graph stream. For each element (s, d, w, t) in the graph stream, TCM adds the weight value w to the bucket (referred to as "bucket") in the h(s|t)-th row and h(d|t)-th column of each matrix, where s|t and d|t respectively represent the concatenation of the source node s and destination node d with the arrival timestamp t of the edge. In this way, TCM retains the temporal information of the edge. When performing a temporal query, it only needs to find the corresponding position in the same way as when inserting, and extract the weight value therein. Further, the GSS method (Patent No.: CN112800288A) introduces fingerprint matching technology to mitigate the impact of hash collisions. However, in terms of implementing temporal queries, the GSS method and TCM adopt the same idea, that is, concatenate the node with the timestamp and then perform a hash mapping to determine the position to complete storage and query.

[0006] The ideas and implementation methods of this type of method are very simple, but the defects are also relatively obvious. On the one hand, a large number of hash collisions result in low query accuracy; on the other hand, during query, it is necessary to traverse all timestamps within the target query range to complete the retrieval, and this workload is extremely large, so it also causes a huge query delay.

[0007] (2) Aggregate timestamps according to different granularities and store them hierarchically. This type of method does not directly utilize the timestamps on the stream edge, but aggregates the timestamps according to different granularities. Each granularity corresponds to a level, and when the stream edge is inserted, it will be updated on each level in turn. For example, the PGSS method (DOI: 10.1007 / s11280-023-01165-z) stores a hierarchical hashmap at each position in the matrix, where the key is the index of the interval in this level and the value is the aggregated weight of the edges falling within the corresponding interval. When an edge arrives, first locate the specific bucket in each compressed adjacency matrix, and then update the hashmap at this position layer by layer. When performing a temporal query, PGSS will map the target query range to the intervals of different hashmap levels and accumulate the entry values of these intervals layer by layer to complete the temporal query.

[0008] Although this kind of method completes the aggregation of time and can locate the target query range to the aggregated interval to reduce the workload, since the aggregation method is still at the timestamp level, for the temporal queries in the actual scenario, it still has to face a large number of traversals of the aggregated intervals, and the query latency is not effectively alleviated. At the same time, it will also bring a huge memory overhead due to the aggregated timestamps, increasing the storage burden.

[0009] (3) First, convert the timestamp into the corresponding time period, and then aggregate and store the time periods according to different granularities. The main difference between this kind of method and the second method is that when integrating time, the aggregated object is the time period rather than the timestamp. When the user wants to complete a query with the granularity of "day", the time granularity (i.e., the query granularity) is set to one day, so that the aggregation is completed according to different granularities in units of "day". When inserting and performing temporal queries on the stream edge, it is also based on the time period corresponding to the timestamp of the stream edge. For example, the Horae method (DOI: 10.1109 / ICDE53745.2022.00254) stores the stream edge in a multi-layer summary structure, and each layer corresponds to a kind of aggregation granularity of the time period. When inserting, the stream edge will be inserted into each layer of the summary structure at the same time. When querying, the target query range is also mapped to different summary levels, and the storage location is found in each respective level, and then the weight value is obtained.

[0010] Since this kind of method aggregates time periods instead of single timestamps, it better meets the temporal query requirements in practical applications and alleviates the large number of traversal problems faced by the previous two methods. However, the current Horae method still has certain defects: on the one hand, the multi-layer summary structure will lead to an increase in query errors because the number of hash collisions accumulates with the increase in the number of layers; on the other hand, whether it is insertion or query, multiple position calculations need to be performed on the same edge, thus increasing the insertion and query latency. Summary of the Invention

[0011] Aiming at the deficiencies of the prior art, the present application proposes a graph stream temporal query method based on FlatMap segment tree. In the storage stage, the present application converts the timestamp of the stream edge into the corresponding time period and completes the insertion storage through one-time positioning. In the query stage, the present application efficiently decomposes the target query range into specific queryable intervals to achieve queries for any interval. In this way, the present application overcomes the problems of low query accuracy, large insertion and query latency in the existing methods, is more suitable for graph stream scenarios with large amounts of data, and provides practical and efficient temporal graph stream query services.

[0012] The present application proposes a graph stream temporal query method based on FlatMap segment tree, including:

[0013] Obtain a temporal query task, where the temporal query task includes: a source node, a destination node, and a target query range at a specified time granularity, and use the edge from the source node to the destination node as the target query edge;

[0014] According to the target query edge, find the corresponding storage location in the graph stream summary storage structure; the graph stream summary storage structure includes: a hash compression matrix and a buffer. Each cell in the hash compression matrix includes: a fingerprint pair of a stream edge and a group of FlatMap segment trees; the FlatMap segment tree is a τ-layer tree structure, and the number of nodes on the g-th layer is 2 g-1 ones, and the time interval length corresponding to the node is 2 τ-g , and the node is used to save the weights of the stream edges arriving within the corresponding time interval; the buffer is used to save the graph stream data that cannot be saved into the hash compression matrix;

[0015] Decompose the target query range onto the corresponding intervals in the pre-updated FlatMap segment tree to obtain a set of queryable sub-intervals, where the set of queryable sub-intervals contains at least one queryable sub-interval;

[0016] According to the corresponding storage location found in the graph stream summary storage structure and the set of queryable sub-intervals, find the nodes corresponding to all queryable sub-intervals in the pre-updated FlatMap segment tree, and take the sum of the weights saved in the found nodes as the graph stream temporal query result.

[0017] For the updated FlatMap segment tree, the update process includes:

[0018] Step S100: Obtain a data set, where the data set includes multiple stream edges, read the start time of the data set and specify the time granularity of the data set;

[0019] Step S101: Traverse each stream edge in the data set and calculate the time period to which the stream edge belongs;

[0020] Step S102: According to the source node and destination node of each stream edge, determine the storage location in the graph stream summary storage structure, and insert each stream edge into the FlatMap segment tree according to the storage location in the graph stream summary storage structure;

[0021] Step S103: Use the weight of each stream edge and the time period to which the stream edge belongs to update the nodes in the corresponding FlatMap segment tree;

[0022] Step S104: Repeat Step S101 to Step S103 until all stream edges have completed the update of the nodes in the FlatMap segment tree, and obtain the pre-updated FlatMap segment tree.

[0023] Finding the corresponding storage location in the graph stream summary storage structure according to the target query edge includes:

[0024] Mapping the source node and the destination node in the temporal query task through a hash function to calculate the hash address of the source node in the temporal query task, the hash address of the destination node in the temporal query task, the fingerprint of the source node in the temporal query task, and the fingerprint of the destination node in the temporal query task;

[0025] Calculating the hash address sequence of the source node and the hash address sequence of the destination node according to the random sequence, the hash address of the source node in the temporal query task, and the hash address of the destination node in the temporal query task;

[0026] Obtaining the storage locations in r 2 hash compression matrices according to the hash address sequence of the source node and the hash address sequence of the destination node, where r is the length of the random sequence;

[0027] Among the storage locations in r 2 hash compression matrices, finding the fingerprint pair that is the same as the fingerprint of the source node in the temporal query task and the fingerprint of the destination node in the temporal query task. The corresponding set of FlatMap segment trees for the same fingerprint pair is the corresponding storage location found in the graph stream summary storage structure.

[0028] Inserting each flow edge into the FlatMap segment tree according to the storage location in the graph stream summary storage structure includes:

[0029] If an empty location is found among the r 2 storage locations, or the fingerprint of the source node of the flow edge and the fingerprint of the destination node of the flow edge are the same as the fingerprint pair stored in any one of the r 2 storage locations, then save the fingerprint of the source node of the flow edge and the fingerprint of the destination node of the flow edge at the empty location or the any one location;

[0030] If an empty location is not found among the r 2 storage locations and the fingerprint of the source node of the flow edge and the fingerprint of the destination node of the flow edge are not the same as the fingerprint pairs stored in all of the r 2 storage locations, then save the fingerprint of the source node of the flow edge and the fingerprint of the destination node of the flow edge in the buffer;

[0031] Calculating the index number k of the FlatMap segment tree to be updated by the current flow edge in the storage location and the element subscript i in the FlatMap segment tree;

[0032] The calculation formula for the index number k of the FlatMap segment tree to be updated is as follows:

[0033]

[0034] Among them, k is the index number of the FlatMap segment tree, and t i is the time period to which the flow edge belongs, and τ is the total number of layers of the FlatMap segment tree;

[0035] For the element subscript i in the FlatMap segment tree, the calculation formula is as follows:

[0036] i = t i +(1 - k)2 τ-1 -2

[0037] Among them, i is the element subscript in the FlatMap segment tree;

[0038] If the index number k of the FlatMap segment tree to be updated is less than or equal to the maximum index value of the current FlatMap segment tree, then use the weight of the current flow edge to update the element with the index number k of the FlatMap segment tree and the element subscript i in the FlatMap segment tree;

[0039] If the index number k of the FlatMap segment tree to be updated is greater than the maximum index value of the current FlatMap segment tree, then create a new FlatMap segment tree with the current index number k of the FlatMap segment tree, and insert the weight of the current flow edge into the new FlatMap segment tree.

[0040] The decomposition of the target query range into the corresponding intervals on the pre-updated FlatMap segment tree to obtain a set of queryable sub-intervals includes:

[0041] If the target query range is on a single FlatMap segment tree, then add the target query range to the intermediate set S';

[0042] If the target query range spans different FlatMap segment trees, then decompose the target query range to a single FlatMap segment tree, and add the sub-ranges obtained after splitting to the intermediate set S', and the sub-ranges in S' all fall on a single FlatMap segment tree;

[0043] According to the greedy query algorithm, decompose the sub-ranges in the intermediate set S' to obtain queryable sub-intervals, and add them to the result set S.

[0044] The decomposition of the sub-ranges in the intermediate set S' according to the greedy query algorithm to obtain queryable sub-intervals and adding them to the result set S includes:

[0045] Step a: If in the sub-range [t b , t e t b > t e, directly end the decomposition; if the sub - range [t b , t e , if t b ≤t e , then the length l of the current sub - range is l = t e - t b +1. If l is 1, add the sub - range [t b , t e to the result set S and end the decomposition;

[0046] Step b: When the length l of the current sub - range is not 1, determine whether t b %2 is 0. If t b %2 = 0, then add the sub - range [t b , t b to the result set S, let t b = t b +1, l = t e - t b +1, and go to step a;

[0047] Step c: If t b %2≠0, then determine whether the current sub - range length l is a power of 2 and satisfies t b %l = 1. If the current sub - range length l is a power of 2 and satisfies t b %l = 1, then no decomposition is required. Add the sub - range [t b , t e to the result set S and end the decomposition;

[0048] Step d: If the current sub - range length l is not a power of 2 or does not satisfy t b %l = 1, then decompose the sub - range [t b , t e into [t b , t b +maxLength - 1] and [t b +maxLength, t e , where maxLength takes the value of the largest power of 2. Add the decomposition result [t b , t b +maxLength - 1] to the result set S, and the remaining [t b +maxLength, t e is used as the target query range for the next iteration. Go to step a to continue the iterative decomposition until the decomposition ends;

[0049] Based on finding the corresponding storage location and the set of queryable sub - intervals in the graph stream summary storage structure, find all the corresponding nodes in the pre - updated FlatMap segment tree for the set of queryable sub - intervals, and take the sum of the weights saved in the found nodes as the graph stream temporal query result, including:

[0050] Determine the FlatMap segment tree index number k corresponding to any queryable sub - interval in the set of queryable sub - intervals;

[0051] Determine the element subscript i in the FlatMap segment tree corresponding to any queryable sub - interval in the set of queryable sub - intervals:

[0052] Based on finding the corresponding storage location, FlatMap segment tree index number k, and element subscript i in the FlatMap segment tree in the graph stream summary storage structure, obtain the weight corresponding to each queryable sub - interval, and take the sum of all weights as the graph stream temporal query result.

[0053] For the FlatMap segment tree index number k corresponding to any queryable sub - interval, the calculation formula is as follows:

[0054]

[0055] where k is the FlatMap segment tree index number, t b is the start position of the queryable sub - interval, and τ is the total number of layers of the FlatMap segment tree.

[0056] For the element subscript i in the FlatMap segment tree corresponding to any queryable sub - interval, the calculation formula is as follows:

[0057]

[0058] where k is the FlatMap segment tree index number, i is the element subscript in the FlatMap segment tree, t b is the start position of the queryable sub - interval, and l is the length of the queryable sub - interval.

[0059] Beneficial effects:

[0060] This application proposes a graph stream temporal query method based on the FlatMap segment tree. Among them, the FlatMap segment tree structure realizes the efficient integration of time logic, supports one - time calculation and insertion of a stream edge, effectively reduces hash collisions, improves query accuracy, and reduces insertion latency at the same time. In addition, the greedy query algorithm can support the efficient decomposition of any target range and shows optimal query performance. Brief description of the drawings

[0061] Figure 1Flowchart of a graph stream temporal query method based on FlatMap segment tree according to an embodiment of the present invention;

[0062] Figure 2 Schematic diagram of a graph stream summary structure supporting temporal queries according to an embodiment of the present invention;

[0063] Figure 3 Schematic diagram of the flow edge insertion process according to an embodiment of the present invention;

[0064] Figure 4 In the embodiment of the present invention, for the target query range [T b , T e decomposition flowchart;

[0065] Figure 5 In the embodiment of the present invention, for the sub - range [t b , t e decomposition flowchart;

[0066] Figure 6 Schematic diagram of the change trend of the average relative error (ARE) of each method according to an embodiment of the present invention;

[0067] Figure 7 Schematic diagram of the change trend of the average query latency of each method according to an embodiment of the present invention. Detailed implementation manners

[0068] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present application.

[0069] Embodiment:

[0070] This embodiment proposes a graph stream temporal query method based on FlatMap segment tree, as Figure 1 shown, including:

[0071] Step S1: Obtain a temporal query task, where the temporal query task includes: a source node, a destination node, and a target query range at a specified time granularity, and use the edge from the source node to the destination node as the target query edge;

[0072] This embodiment uses the lkml-reply dataset, which is derived from the real Linux kernel email network and is a collection of communication records, containing 63,399 email addresses (nodes) and 1,096,440 communication records (edges). The query set consists of 1,000 temporal queries for edge aggregation weights, and the edges and query ranges in each query are randomly generated. In the storage stage of the embodiment, specifically, a stream edge (11, 91, 1, 1139121093) is taken as an example for processing. This stream edge indicates that the user with email address ID 11 sent a communication to the user with email address ID 91 at the timestamp 1139121093. In the embodiment, the time granularity is set to one day, granularityLength = 86400, m = 1430, the fingerprint length is 6 bits, the number of row addresses and column addresses are 4 respectively, and the coverage length of a single FlatMap is 2 τ-1 = 4096, where τ = 13. In the query stage of the embodiment, specifically, a query (11, 91, 8, 15) is taken as an example for processing. This query represents " 8 , T 15 What is the aggregation weight of the edge (11, 91)?". Since the time granularity of this embodiment is "day", actually this query reflects "How many communications did the user with email address ID 11 send to the user with email address ID 91 between the 8th day and the 15th day?".

[0073] Step S2: According to the target query edge, find the corresponding storage location in the graph stream summary storage structure; the graph stream summary storage structure includes: a hash compression matrix and a buffer. Each cell in the hash compression matrix includes: a fingerprint pair of the stream edge and a set of FlatMap segment trees; the FlatMap segment tree is a τ-layer tree structure, and the number of nodes on the g-th layer is 2 g-1 nodes, and the time interval length corresponding to the node is 2 τ-g , and the node is used to save the weights of the stream edges arriving within the corresponding time interval; the buffer is used to save the graph stream data that cannot be saved into the hash compression matrix.

[0074] In this embodiment, the complete graph stream summary storage structure, as Figure 2 shown, includes the following two parts: a hash-based compression matrix (referred to as sketch) and a buffer. Among them, each bucket (i.e., each cell) in the compression matrix stores:

[0075] (1) A fingerprint pair of the edge <f s , f d >, as the identity id of the edge.

[0076] (2) A group of FlatMap segment trees, used to store the weights of edges, and can read the edge weight values that arrive within the corresponding time range in the nodes. Each FlatMap segment tree stores the weights of edges within a time period, and the FlatMap segment tree is the key to solving the temporal query problem.

[0077] First, each node in the segment tree corresponds to an interval (i.e., a time range), and there are several time granularities within the interval. In this way, each node of the segment tree stores the edge weights that arrive within this time range, where the interval length corresponding to the leaf node is 1, and it contains only one time granularity. The number of nodes on the g-th layer of the segment tree is 2 g-1 ^g, and the interval length corresponding to each layer of nodes is 2 τ-g ^τ, where τ is the total number of layers of the segment tree. Such a design structure can ensure that any target query range can find several nodes on the segment tree that just cover it, thus completing the temporal query. The FlatMap segment tree is formally an array, which is proposed to adapt to the characteristics of graph stream data and reduce the insertion delay. It is equivalent to the result obtained by breadth-first traversal after the number of layers of the segment tree reaches a certain threshold. The time organization logic in the FlatMap segment tree remains, and each element in the array corresponds to a node in the segment tree, that is, an interval. Assuming that the FlatMap segment tree corresponds to a segment tree with τ layers, then the FlatMap segment tree is formally an array with a length of 2 0 ^τ + 2 1 ^(τ - 1) + … + 2 τ-2 ^1 + 2 τ-1 ^0 = 2 τ-1 ^(τ + 1) - 1, and the covered time range length is 2 τ-1 ^τ. The first element in the array (i.e., the root node) corresponds to an interval length of 2 τ-1 ^τ, the next 2 1 ^(τ - 1) elements correspond to an interval length of 2 τ-2 ^(τ - 1), the next 2 2 ^(τ - 2) elements correspond to an interval length of 2 τ-3 ^(τ - 2), ……, the next 2 τ-2 ^1 elements correspond to an interval length of 2 1 ^1, and the last 2 τ-1 ^0 elements (i.e., the leaf nodes) correspond to an interval length of 2 0 ^0. When the time exceeds the range covered by the current FlatMap segment tree, a new FlatMap segment tree will be dynamically opened for storage. Each FlatMap segment tree is represented by the FlatMap segment tree index k, and the nodes in the FlatMap segment tree are represented by the element subscript i in the FlatMap segment tree.

[0078] Step S3: Decompose the target query range onto the corresponding intervals in the pre-updated FlatMap segment tree to obtain a set of queryable sub-intervals, where the set of queryable sub-intervals contains at least one queryable sub-interval;

[0079] Step S4: Based on finding the corresponding storage location in the graph stream summary storage structure and the set of queryable sub-intervals, find the corresponding nodes in all the queryable sub-intervals in the pre-updated FlatMap segment tree, and use the sum of the weights saved in the found nodes as the graph stream temporal query result.

[0080] In this embodiment, before the query process, the FlatMap segment tree needs to be updated first. Therefore, the update process of the FlatMap segment tree is described as follows:

[0081] The updated FlatMap segment tree, as Figure 3 shown, the update process includes:

[0082] Step S100: Obtain a data set, where the data set includes multiple flow edges, read the start time of the data set and specify the time granularity of the data set;

[0083] In this embodiment, read the first flow edge of the lkml-reply data set and use its arrival timestamp as the start time, startTime = 1136080607; the time granularity is specified as one day (the time granularity is the finest granularity in the query process), and the length of the time granularity granularityLength = 86400.

[0084] Step S101: Traverse each flow edge in the data set and calculate the time period to which the flow edge belongs;

[0085] The calculation method of the time period t i to which the flow edge belongs is as follows:

[0086]

[0087] For the flow edge (11, 91, 1, 1139121093), its arrival timestamp t = 1139121093, so the time period to which it belongs is

[0088] Step S102: Determine the storage location in the graph stream summary storage structure according to the source node and destination node of each flow edge, and insert each flow edge into the FlatMap segment tree according to the storage location in the graph stream summary storage structure.

[0089] For a flow edge (11, 91, 1, 1139121093), first determine the position in the compression matrix where it can be stored, that is, the corresponding coordinates in the compression matrix sketch; if the insertion cannot be completed in the sketch, store the flow edge in the buffer. The specific steps are as follows:

[0090] Step S102.1: Map the source node and destination node of the flow edge through a hash function to calculate the hash address of the source node of the flow edge, the hash address of the destination node of the flow edge, the fingerprint of the source node of the flow edge, and the fingerprint of the destination node of the flow edge;

[0091] In this embodiment, when calculating the hash address and fingerprint, the calculation methods of the hash address and fingerprint of the node are as follows:

[0092] f(v) = H(v) % F

[0093] where F is the maximum size of the fingerprint. Map the node id through the hash function H(·) to obtain the node hash values H(s) = H(11) = 173226315 and H(d) = H(91) = 708853215; then calculate the hash address of the source node and fingerprint f(s) = 173226315 % 63 = 3, and the hash address of the destination node and fingerprint f(d) = 708853215 % 63 = 21.

[0094] Step S102.2: Calculate the hash address sequence of the source node and the hash address sequence of the destination node according to the random sequence, the hash address of the source node of the flow edge, and the hash address of the destination node of the flow edge;

[0095] Further, to calculate the hash address sequence, first calculate the random sequence {q i (v) | 1 ≤ i ≤ r}, and then calculate the hash address sequence {h i (v) | 1 ≤ i ≤ r} according to the random sequence and the fingerprint and hash address obtained in step 3-1. The calculation method is as follows:

[0096]

[0097] h i (v) = (h(v) + q i (v)) % m, 1 ≤ i ≤ r

[0098] where a, b, and p are all parameters used to generate the random sequence, m is the matrix width, r is the number of random sequences, each node will generate r random address sequences, so each edge will have r 2A bucket that can be used for storage. In this embodiment, a = 5, b = 739, p = 1048576, and the calculation shows that: q 1 (s) = 754, q 2 (s) = 4509, q 3 (s) = 23284, q 4 (s) = 117159. h 1 (s) = 488, h 2 (s) = 1383, h 3 (s) = 138, h 4 (s) = 1063. Therefore, the hash address sequence corresponding to the source node is {h i (s)} = {488, 1383, 138, 1063}. Similarly, the hash address corresponding to the destination node is {h i (d)} = {1242, 1017, 1372, 287}.

[0099] Step S102.3: According to the hash address sequence of the source node and the hash address sequence of the destination node, obtain r 2 storage positions in the hash compression matrix, where r is the length of the random sequence, and use the r 2 storage positions as the storage positions in the graph stream summary storage structure;

[0100] Step S102.4: If an empty position is found among the r 2 storage positions or the fingerprints of the source node of the flow edge and the fingerprints of the destination node of the flow edge are the same as the fingerprint pairs stored in any position among the r 2 storage positions, then save the fingerprints of the source node of the flow edge and the fingerprints of the destination node of the flow edge in the empty position or that any position;

[0101] Step S102.5: If no empty position is found among the r 2 storage positions and the fingerprints of the source node of the flow edge and the fingerprints of the destination node of the flow edge are not the same as all the fingerprint pairs stored in the r 2 storage positions, then save the fingerprints of the source node of the flow edge and the fingerprints of the destination node of the flow edge in the buffer;

[0102] In this embodiment, further traverse the addresses, perform fingerprint comparison, and determine the final storage position. According to the obtained hash address sequence, determine the r of the flow edge 2A storable bucket address (where the elements in the compressed matrix, the row number and column number are the address), traverse these addresses in row-major order. If the first bucket is empty, directly store the fingerprint pair <f(s), f(d)> into this bucket and proceed to step S102.6; if the first bucket is not empty, compare the fingerprint pairs already stored in this bucket with the fingerprint pair of the current flow edge. If the fingerprint pairs are the same, it is considered that the edge stored in this bucket and the current flow edge are the same edge, and proceed to step S102.6; if the fingerprint pairs are different, continue to traverse the next bucket, and so on, until an empty bucket or a bucket with the same stored fingerprint pair is found. If r 2 This edge cannot be inserted into any of the buckets, then insert this edge into the buffer and proceed to step S102.6.

[0103] Specifically: Given that the fingerprint of the source node is f(s) = 3 and the hash address sequence is {h i (s)} = {488, 1383, 138, 1063}, and the fingerprint of the destination node is f(d) = 21 and the hash address sequence is {h i (d)} = {1242, 1017, 1372, 287}, then the 16 bucket addresses where the flow edge can be stored are: <488, 1242>, <488, 1017>, <488, 1372>, <488, 287>,

[0104] <1383, 1242>, <1383, 1017>, <1383, 1372>, <1383, 287>,

[0105] <138, 1242>, <138, 1017>, <138, 1372>, <138, 287>,

[0106] <1063, 1242>, <1063, 1017>, <1063, 1372>, <1063, 287>. Next, traverse these 16 buckets in order. First, check <488, 1242>. Assume that the fingerprint pair <11, 31> has been stored at this position, which means this position has been occupied by other edges. Continue to traverse the next address <488, 1017>. Assume that the fingerprint pair <3, 21> is stored at this position, which means the same flow edge has reached this position before. It can be determined that the flow edge insertion address is <488, 1017>, and the traversal ends.

[0107] Step S102.6: Calculate the index k of the FlatMap segment tree to be updated by the current flow edge in the storage location and the element subscript i in the FlatMap segment tree;

[0108] Step S102.7: If the index number k of the FlatMap segment tree to be updated is less than or equal to the maximum index of the current FlatMap segment tree, update the element with FlatMap segment tree index k and element subscript i in the FlatMap segment tree with the weight of the current flow edge;

[0109] Step S102.8: If the index number k of the FlatMap segment tree to be updated is greater than the maximum index of the current FlatMap segment tree, create a new FlatMap segment tree with the current FlatMap segment tree index number k, and insert the weight of the current flow edge into the new FlatMap segment tree.

[0110] In this embodiment, according to the time period t to which the flow edge belongs i and the coverage range 2 of the FlatMap segment tree τ-1 , the index number k (starting from 0) of the FlatMap segment tree corresponding to the flow edge can be determined, and the calculation method is as follows:

[0111]

[0112] If k is greater than the current FlatMap segment tree index number, start the next FlatMap segment tree.

[0113] The length of the covered time period of a single FlatMap segment tree in this embodiment is 2 τ-1 = 4096. The time period to which the current flow edge (11, 91, 1, 1139121093) belongs is 36, then the weight value of this flow edge should be stored in the FlatMap segment tree with k = 0.

[0114] To reduce the processing delay of the flow edge, this embodiment adopts lazy update, that is, only update the weights at the corresponding positions in the last 2 τ-1 elements of the FlatMap segment tree. According to the time period t to which the flow edge belongs i , the coverage range 2 of the FlatMap segment tree τ-1 and the confirmed index number k, the array subscript i (starting from 0) to be updated can be determined, and the calculation method is as follows:

[0115] i = t i +(1 - k)2 τ-1 - 2

[0116] After confirming the array subscript i, only need to update the element fm k [i]+ = w. In this embodiment, for the flow edge (11, 91, 1, 1139121093), t i= 36, k = 0, the coverage time period length of a single FlatMap segment tree is 4096, then the array subscript i to be updated is i = 36 + 4096 - 2 = 4130, and add 1 to the weight value in fm 0

[4130] .

[0117] Thus, the insertion process for the flow edge (11, 91, 1, 1139121093) ends.

[0118] Step S103: Update the nodes in the corresponding FlatMap segment tree by using the weight of each flow edge and the time period to which the flow edge belongs;

[0119] Step S104: Repeat Step S101 to Step S103 until all flow edges have completed the update of the nodes in the FlatMap segment tree, and obtain the pre-updated FlatMap segment tree.

[0120] In this embodiment, continue the Flash update of the FlatMap segment tree. This step can be selected to be performed when no more flow edges arrive, when a query task arrives, or when performing FlatMap expansion in Step S102.6. In this example, when a query task arrives, a one-time Flash update is performed on all FlatMap segment trees. The update process is to update level by level from the back to the front until all updates are completed (updated to i = 0). First, update the (τ - 1)-th layer: Start updating from the element with subscript i = 2 τ-2 - 1, and update a total of 2 τ-2 elements backward. The update method is fm k [2 τ-2 - 1] = fm k [2 τ-1 - 1] + fm k [2 τ-1 , and so on. The update of each element is the sum of the corresponding two elements in the next layer; then update the (τ - 2)-th layer: Start updating from the element with subscript i = 2 τ-3 - 1, and update a total of 2 τ-3 elements backward;...; update the first layer: Start updating from the element with subscript i = 0, and update a total of 2 0 elements backward. Thus, the Flash update operation of the FlatMap segment tree is completed.

[0121] In this embodiment, τ = 13, and the coverage time period length of a single FlatMap segment tree is 2 τ-1 = 4096. Then during the Flash update: First, update the elements with i = 2047, 2048,..., 4094. The update method is

[0122] fm k

[2047] = fmk

[4095] + fm k

[4096] , fm k

[2048] = fm k

[4097] + fm k

[4098] , …,

[0123] fm k

[4094] = fm k

[8189] + fm k

[8190] , thus completing the Flash update for the τ - 1 = 12th layer; then update the elements with i = 1023, 1024, …, 2046, and the update method is fm k

[1023] = fm k

[2047] + fm k

[2048] , fm k

[1024] = fm k

[2049] + fm k

[2050] , …, fm k

[2046] = fm k

[4093] + fm k

[4094] , thus completing the Flash update for the τ - 2 = 11th layer; …; and so on until the elements with i = 0 are updated, and the update method is

[0124] fm k [0] = fm k [1] + fm k [2], after completing the update of layer 1, the Flash ends.

[0125] After updating the FlatMap segment tree, the following combines Figure 4 and Figure 5 shown in the query range decomposition flowchart to illustrate the steps of a temporal query task (11, 91, 8, 15).

[0126] In step S2, finding the corresponding storage location in the graph stream summary storage structure according to the target query edge includes:

[0127] Step S2.1: Map the source node and destination node in the temporal query task through a hash function to calculate the hash address of the source node in the temporal query task, the hash address of the destination node in the temporal query task, the fingerprint of the source node in the temporal query task, and the fingerprint of the destination node in the temporal query task;

[0128] Step S2.2: Query the hash address of the source node in the temporal query task and the hash address of the destination node in the temporal query task according to the random sequence, and calculate the hash address sequence of the source node and the hash address sequence of the destination node;

[0129] Step S2.3: Obtain the storage positions in r 2 hash compression matrices according to the hash address sequence of the source node and the hash address sequence of the destination node, where r is the length of the random sequence;

[0130] Step S2.4: Among the storage positions in r 2 hash compression matrices, find the fingerprint pairs that are the same as the fingerprint of the source node in the temporal query task and the fingerprint of the destination node in the temporal query task. The set of FlatMap segment trees corresponding to the same fingerprint pairs is the corresponding storage position found in the graph stream summary storage structure.

[0131] In this embodiment, finding the corresponding position of the target query edge is similar to the step of determining the storage position of each flow edge in the graph stream summary storage structure in step S102. When performing a temporal query, it is necessary to find its corresponding position according to the target query edge. First, find the hash address sequences of the source node and the destination node of the target query edge, which is equivalent to finding the r 2 bucket addresses that can be stored. Next, traverse these r 2 addresses in sequence until the fingerprint pair stored in the bucket is the same as the fingerprint pair of the destination query edge. If no result is found in the r 2 buckets, search for the target query edge in the buffer. If it still cannot be found, it means that the edge has not been reached, and 0 is directly returned to end the query. This step is similar to step S102, and the specific process will not be elaborated. Finally, the corresponding address of the target query edge (11, 91) can be determined as <488, 1017>.

[0132] In step S3, the decomposing the target query range into the corresponding intervals on the pre-updated FlatMap segment tree to obtain a set of queryable sub-intervals includes:

[0133] Step S3.1: If the target query range is on a single FlatMap segment tree, add the target query range to the intermediate set S';

[0134] Step S3.2: If the target query range spans different FlatMap segment trees, decompose the target query range onto a single FlatMap segment tree, and add the split sub-ranges to the intermediate set S'. The sub-ranges in S' all fall on a single FlatMap segment tree;

[0135] Step S3.3: Decompose the sub-ranges in the intermediate set S' according to the greedy query algorithm to obtain queryable sub-intervals, and add them to the result set S.

[0136] In this embodiment, when performing a temporal query, the target query range [T b , T e needs to be decomposed into corresponding intervals in the FlatMap segment tree to obtain a set S of queryable sub-intervals. This process mainly includes two parts: The first part is to split the target query range that spans different FlatMap structures, and add the obtained sub-ranges to the intermediate set S'. The sub-ranges in S' all fall on a single FlatMap structure; The second part is to decompose the sub-ranges in S', and add the obtained sub-intervals to the result set S. The sub-intervals in S are all queryable sub-intervals on a single FlatMap structure.

[0137] If the length L of the current range > 2 τ-1 or L ≤ 2 τ-1 and T e > dp, where dp = 2 τ-1+n is the first FlatMap demarcation point greater than or equal to currentT b (currentT b is the start time of the current range, and initially currentT b = T b ), then the current range involves multiple FlatMaps. Split the interval into [currentT b , dp] and [dp + 1, T e . Add the split [currentT b , dp] to S', and use [dp + 1, T e as the current range to continue the decomposition. Repeat the above steps until the current range does not involve multiple FlatMap structures, and directly add it to S' to end.

[0138] In this implementation, the target query range is [T 8 , T 15 , with a length of L = 8 = 2 3 and T e = 15 > 8, so the current query range involves multiple FlatMap structures. Split the current range into [t 8 , t 8 and [t 9 , t 15 , where [t 8 , t 8 is added to S', and [t 9 , t 15Continue the decomposition as the new current range; the length of the current query range is L = 7 < 2 3 and T e = 15 < 16, so there is no need to decompose further. Directly add [t 9 , t 15 to S', and end. At this time, the sub - ranges in S' are: [t 8 , t 8 , [t 9 , t 15 .

[0139] The sub - ranges in S' are already ranges that fall on a single FlatMap structure. Next, decompose each sub - range [t b , t e in S' to obtain the queryable sub - intervals of the corresponding segment tree interval and add them to the result set S. The decomposition steps for each sub - range [t b , t e are as follows:

[0140] According to the greedy query algorithm, decompose the sub - ranges in the intermediate set S' to obtain queryable sub - intervals and add them to the result set S, including:

[0141] Step a: If in the sub - range [t b , t e , t b > t e , directly end the decomposition; if in the sub - range [t b , t e , t b ≤t e , then the length l of the current sub - range is l = t e - t b + 1. If the length l of the current sub - range is 1, add the sub - range [t b , t e to the result set S and end the decomposition;

[0142] Step b: When the length l of the current sub - range is not 1, judge whether t b % 2 is 0. If t b % 2 = 0, then add the sub - range [t b , t b to the result set S, let t b = t b + 1, l = t e - t b + 1, and go to step a;

[0143] Step c: If t b % 2 ≠ 0, then judge whether the current sub - range length l is a power of 2 and satisfies tb %l = 1. If the length l of the current sub - range is a power of 2 and satisfies t b %l = 1, then there is no need to decompose. Add the sub - range [t b , t e to the result set S and end the decomposition;

[0144] Step d: If the length l of the current sub - range is not a power of 2 or does not satisfy t b %l = 1, then decompose the sub - range [t b , t e into [t b , t b + maxLength - 1] and [t b + maxLength, t e , where the value of maxLength is the largest power of 2. Add the decomposition result [t b , t b + maxLength - 1] to the result set S, and the remaining [t b + maxLength, t e is used as the target query range for the next iteration. Go to step a and continue the iterative decomposition until the decomposition ends.

[0145] In this implementation, the target query range is [T 8 , T 15 . The sub - ranges in S' obtained through the above are [t 8 , t 8 , [t 9 , t 15 .

[0146] Perform decomposition on [t 8 , t 8 : First, execute step a. Since l = 8 - 8 + 1 = 1, directly add [t 8 , t 8 to the result set S and end.

[0147] For [t 9 , t 15 perform decomposition: First, execute step a. Since l = 7, perform step b. Since 9 % 2 = 1, perform step c. Since l = 7, perform step d. The current sub - range is decomposed into [t 9 , t 12 and [t 13 , t 15 , where maxLength = 4. Add [t 9 , t 12 to the result set S, and [t 13 , t 15Continue with the decomposition. First, execute step a. Since l = 3, proceed to step b. Since 13 % 2 = 1, proceed to step c. Since l = 3, proceed to step d. The current sub-range is decomposed into [t 13 , t 14 and [t 15 , t 15 , where maxLength = 2. Add [t 13 , t 14 to the result set S. [t 15 , t 15 continues with the decomposition. First, execute step a. Since l = 1, directly add [t 15 , t 15 to the result set S, and end.

[0148] So far, all the sub-ranges in S' have been traversed, and S is obtained. The sub-intervals in S are: [t 8 , t 8 , [t 9 , t 12 , [t 13 , t 14 and [t 15 , t 15 .

[0149] In this embodiment, each interval in S is traversed and the query operation is executed. After obtaining the storage location of the target query edge, find all the elements corresponding to the queryable sub-intervals [t b , t e in S on its corresponding FlatMap, and then add up the weight values obtained by the traversal and return them as the temporal query result.

[0150] In step S4, the method of finding all the corresponding nodes in the pre-updated FlatMap segment tree according to the corresponding storage location and the set of queryable sub-intervals in the graph flow summary storage structure, and taking the sum of the weights saved in the found nodes as the graph flow temporal query result includes:

[0151] Step S4.1: Determine the FlatMap segment tree index number k corresponding to any queryable sub-interval in the set of queryable sub-intervals. The calculation formula is as follows:

[0152]

[0153] Where k is the FlatMap segment tree index number, t b is the start position of the queryable sub-interval, and τ is the total number of layers of the FlatMap segment tree.

[0154] Step S4.2: Determine the element subscript i in the FlatMap segment tree corresponding to any queryable sub-interval in the set of queryable sub-intervals. The calculation formula is as follows:

[0155]

[0156] where i is the element subscript in the FlatMap segment tree, and l is the length of the queryable sub-interval.

[0157] Step S4.3: Obtain the weight corresponding to each queryable sub-interval according to the corresponding storage location found in the graph stream summary storage structure, the FlatMap segment tree index k, and the element subscript i in the FlatMap segment tree. Take the sum of all weights as the graph stream temporal query result.

[0158] In this embodiment, for the queryable sub-interval [t 8 ,t 8 in S, its corresponding FlatMap index The element subscript corresponding to it in the 0th FlatMap segment tree The weight value corresponding to this interval is fm 0

[5102] ;

[0159] For the queryable sub-interval [t 9 ,t 12 in S, its corresponding FlatMap index The element subscript corresponding to it in the 0th FlatMap segment tree The weight value corresponding to this interval is fm 0

[1025] ;

[0160] For the queryable sub-interval [t 13 ,t 14 in S, its corresponding FlatMap index The element subscript corresponding to it in the 0th FlatMap segment tree The weight value corresponding to this interval is fm 0

[2053] ;

[0161] For the queryable sub-interval [t 15 ,t 15 in S, its corresponding FlatMap index The element subscript corresponding to it in the 0th FlatMap segment tree The weight value corresponding to this interval is fm 0

[4109] .

[0162] The fm in the bucket with the address <488,1017> 0

[4102] + fm 0

[1025] + fm 0

[2053] + fm 0

[4109] Return as a result.

[0163] Thus, a (11, 91, 8, 15) temporal query task ends.

[0164] A graph stream temporal query method based on FlatMap segment tree proposed in this embodiment, in terms of storage, proposes a FlatMap segment tree structure to retain the time information of edges and support temporal queries. The FlatMap segment tree is formally an array of length 2 0 + 2 1 + … + 2 τ-2 + 2 τ-1 = 2 τ - 1, which is equivalent to being obtained by breadth-first traversal after the number of layers of the segment tree reaches a certain threshold. The FlatMap segment tree contains time organization logic. Each element in the array corresponds to a node in the segment tree, that is, an interval. Each interval contains several time periods and stores the edge weight information arriving in these time periods. The interval corresponding to the first element (i.e., the root node) in the array has a length of 2 τ-1 , and the intervals corresponding to the subsequent 2 1 elements have a length of 2 τ-2 , the intervals corresponding to the subsequent 2 2 elements have a length of 2 τ-3 , ……, the intervals corresponding to the subsequent 2 τ-2 elements have a length of 2 1 , and the intervals corresponding to the last 2 τ-1 elements (i.e., the leaf nodes) have a length of 2 0 and contain only one time period (i.e., one time granularity). When a flow edge is inserted, only the leaf nodes of the FlatMap segment tree are updated (lazy update). When the time exceeds the range covered by the current FlatMap segment tree, a new FlatMap segment tree is dynamically opened for storage.

[0165] In terms of query, this embodiment proposes a greedy query algorithm, which decomposes the time range given by the user into the node intervals of the corresponding segment tree to complete the temporal query. Specifically, for any target query range, a set of sub-intervals needs to be found that can exactly cover the target query range without overlap, and at the same time, the sub-intervals are required to match the node intervals in the above FlatMap segment tree structure to support queries for any time range.

[0166] By applying the method of this embodiment, edge weight query experiments were conducted on the real lkml-reply dataset within different query ranges of lengths (L = 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 1536, 2560). The experimental results show that the DSTGS method proposed in this embodiment exhibits better performance in terms of query accuracy compared to other methods. Figure 6 It shows that DSTGS maintained a lower average relative error (ARE) across all query ranges, and in some cases, the ARE was less than 10 -2 even 10 -3 resulting in an improvement of two orders of magnitude compared to the existing optimal method. Additionally, while ensuring better query accuracy, DSTGS achieved lower query latency. Figure 7 It shows that DSTGS maintained a lower average query latency (unit: ms) across all query ranges, representing an improvement of one order of magnitude compared to the existing optimal method.

[0167] From the experimental results, the method of this implementation can provide more accurate and efficient temporal query services in real-world graph stream applications, fully meeting the actual application requirements.

[0168] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.

[0169] The protection scope of this application is not limited to the above embodiments. Obviously, those skilled in the art can make various changes and deformations to this disclosure without departing from the scope and spirit of this disclosure. If these changes and deformations fall within the scope of the claims of this disclosure and their equivalent technologies, the intention of this disclosure also includes these changes and deformations.

Claims

1. A graph stream temporal query method based on FlatMap segment tree, characterized in that: include: Acquire a temporal query task, wherein the temporal query task includes: a source node, a destination node, and a target query range at a specified time granularity, and uses the source node to the destination node as a target query edge; According to the target query edge, find the corresponding storage location in the graph stream summary storage structure; the graph stream summary storage structure includes: a hash compression matrix and a buffer, each cell in the hash compression matrix includes: a fingerprint pair of the flow edge and a set of FlatMap segment trees; the FlatMap segment tree is a τ-layer tree structure, and the number of nodes on the g-th layer is 2 g-1 The length of the time interval corresponding to the node is 2 τ-g , the node is used to save the weight of the flow edge arriving in the corresponding time interval; the buffer is used to save the graph flow data that cannot be saved in the hash compression matrix; Decomposing the target query range into corresponding intervals in the pre-updated FlatMap segment tree to obtain a queryable sub-interval set, wherein the queryable sub-interval set includes at least one queryable sub-interval; According to the corresponding storage location and the set of queryable subintervals found in the graph stream summary storage structure, the nodes corresponding to all queryable subintervals are found in the pre-updated FlatMap segment tree, and the sum of the weights stored in the found nodes is used as the graph stream temporal query result.

2. According to claim 1, a graph stream temporal query method based on FlatMap segment tree is characterized in that: The updated FlatMap segment tree includes the following steps: Step S100: obtaining a data set, the data set including a plurality of flow edges, reading the start time of the data set and specifying the time granularity of the data set; Step S101: traverse each flow edge in the data set and calculate the time period to which the flow edge belongs; Step S102: Determine the storage location in the graph flow summary storage structure according to the source node and the destination node of each flow edge, and insert each flow edge into the FlatMap segment tree according to the storage location in the graph flow summary storage structure; Step S103: using the weight of each flow edge and the time period to which the flow edge belongs, updating the corresponding node in the FlatMap segment tree; Step S104: Repeat steps S101 to S103 until all flow edges have completed the update of the nodes of the FlatMap segment tree, and obtain a pre-updated FlatMap segment tree.

3. According to the FlatMap segment tree-based graph stream temporal query method of claim 1, it is characterized in that: The step of finding a corresponding storage location in the graph stream summary storage structure according to the target query edge includes: Mapping the source node and the destination node in the temporal query task through a hash function, calculating the hash address of the source node in the temporal query task, the hash address of the destination node in the temporal query task, the fingerprint of the source node in the temporal query task, and the fingerprint of the destination node in the temporal query task; Calculate a hash address sequence of the source node and a hash address sequence of the destination node according to the random sequence, the hash address of the source node in the temporal query task, and the hash address of the destination node in the temporal query task; According to the hash address sequence of the source node and the hash address sequence of the destination node, we can get r 2 The storage locations in the hash compression matrix, where r is the length of the random sequence; In r 2 In the storage locations in the hash compression matrix, a fingerprint pair that is the same as the fingerprint of the source node in the temporal query task and the fingerprint of the destination node in the temporal query task is found. A set of FlatMap segment trees corresponding to the same fingerprint pair is the corresponding storage location found in the graph stream summary storage structure.

4. According to claim 2, a graph stream temporal query method based on FlatMap segment tree is characterized in that: The method of inserting each flow edge into the FlatMap segment tree according to the storage position in the graph flow summary storage structure includes: If in r 2 The fingerprint of the source node of the empty position or the flow edge and the fingerprint of the destination node of the flow edge are combined with r 2 If any of the storage locations are the same, the fingerprint of the source node of the flow edge and the fingerprint of the destination node of the flow edge are saved in the empty location or the any location; If in r 2 No empty location is found in the storage locations and the fingerprint of the source node of the flow edge and the fingerprint of the destination node of the flow edge are the same as r 2 If all the locations in the storage locations are different, the fingerprint of the source node of the flow edge and the fingerprint of the destination node of the flow edge are stored in the buffer; Calculate the FlatMap segment tree index k to be updated by the current flow edge in the storage location and the element subscript i in the FlatMap segment tree; The FlatMap line segment index k to be updated is calculated as follows: Where k is the index number of the FlatMap segment tree, t i is the time period to which the flow edge belongs, τ is the total number of layers of the FlatMap segment tree; The element subscript i in the FlatMap segment tree is calculated as follows: i=t i +(1-k)2 τ-1 -2 Among them, i is the element subscript in the FlatMap segment tree; If the index number k of the FlatMap segment tree to be updated is less than or equal to the maximum index number of the previous FlatMap segment tree, then the element with the index k of the FlatMap segment tree and the element subscript i in the FlatMap segment tree is updated with the weight of the current flow edge; If the index number k of the FlatMap segment tree to be updated is greater than the maximum index value of the previous FlatMap segment tree, a new FlatMap segment tree is created using the current FlatMap segment tree index number k, and the weight of the current flow edge is inserted into the new FlatMap segment tree.

5. According to the FlatMap segment tree-based graph stream temporal query method of claim 1, it is characterized in that: The target query range is decomposed into corresponding intervals in the pre-updated FlatMap segment tree to obtain a set of queryable sub-intervals, including: If the target query range is on a single FlatMap segment tree, add the target query range to the intermediate set S'; If the target query range spans across different FlatMap segment trees, the target query range is decomposed into a single FlatMap segment tree, and the sub-ranges obtained after the decomposition are added to the intermediate set S', and the sub-ranges in S' all fall on a single FlatMap segment tree; According to the greedy query algorithm, the sub-ranges in the intermediate set S' are decomposed to obtain the queryable sub-intervals and added to the result set S.

6. According to claim 5, a graph stream temporal query method based on FlatMap segment tree is characterized in that: According to the greedy query algorithm, the sub-ranges in the intermediate set S' are decomposed to obtain queryable sub-intervals, which are added to the result set S, including: Step a: If the sub-range [t b ,t e ] b >t e , directly end the decomposition; if the sub-range [t b ,t e ] b ≦t e , then the length of the current sub-range l = t e -t b +1, if l is 1, the subrange [t b ,t e ] is added to the result set S to end the decomposition; Step b: If the length l of the current sub-range is not 1, determine t b %2 is 0, if t b %2=0, then the subrange [t b ,t b ]Add the result set S, let t b =t b +1, l = t e -t b +1, go to step a; Step c: If t b % 2≠0, then determine whether the current sub-range length l is 2 to the power of n and satisfies t b % l = 1, if the current sub-range length l is a power of 2 and satisfies t b %l=1, no decomposition is required, and the subrange [t b ,t e ]Add result set S to end decomposition; Step d: If the current sub-range length l is not a power of 2 or does not satisfy t b %l=1, then the subrange [t b ,t e ] is decomposed into [t b ,t b +maxLength-1] and [t b +maxLength,t e ], where maxLength is the largest power of 2, and the decomposition result [t b ,t b +maxLength-1] is added to the result set S, and the remaining [t b +maxLength,t e ] as the target query range for the next iteration, and go to step a to continue iterative decomposition until the decomposition is completed.

7. According to claim 1, a graph stream temporal query method based on FlatMap segment tree is characterized in that: The method finds the corresponding storage location and the queryable sub-interval set in the graph stream summary storage structure, finds the corresponding nodes in all queryable sub-interval sets in the pre-updated FlatMap segment tree, and uses the sum of the weights stored in the found nodes as the graph stream temporal query result, including: Determine the FlatMap segment tree index number k corresponding to any queryable subinterval in the queryable subinterval set; Determine the element index i in the FlatMap segment tree corresponding to any queryable subinterval in the queryable subinterval set: According to finding the corresponding storage location in the graph stream summary storage structure, the FlatMap segment tree index number k and the element subscript i in the FlatMap segment tree, the weight corresponding to each queryable subinterval is obtained, and the sum of all weights is used as the graph stream temporal query result.

8. According to claim 7, a graph stream temporal query method based on FlatMap segment tree is characterized in that: The FlatMap segment tree index number k corresponding to any queryable subinterval is calculated as follows: Where k is the index number of the FlatMap segment tree, t b is the starting position of the queryable subinterval, and τ is the total number of layers of the FlatMap segment tree.

9. According to claim 7, a graph stream temporal query method based on FlatMap segment tree is characterized in that: The element subscript i in the FlatMap segment tree is calculated as follows: Among them, k is the index number of the FlatMap segment tree, i is the element subscript in the FlatMap segment tree, and t b is the starting position of the queryable subinterval, and l is the length of the queryable subinterval.

Citation Information

Patent Citations

  • Graph stream data processing method

    CN112800288A

  • Double-hash table association method for inquiring interval durability top-k

    CN102663030A

  • Safety nearest neighbor query method and system based on maximum division and random data block

    CN102999594A

  • Stream data cluster search method based on query points

    CN114510506A

  • Efficient graph flow measurement method based on adjacent matrix

    CN116628025A