Distributed storage and query optimization system and method for dynamic graph data

By introducing graph data compression coding mechanism, searchable coding tree path index, reinforcement learning query optimization and multi-dimensional similarity edge cache technology in the dynamic graph data processing system, the problems of low storage efficiency and slow query response of large-scale dynamic graph data are solved, efficient storage and fast query are achieved, and system performance is significantly improved.

CN120067402APending Publication Date: 2025-05-30SHANGHAI GUOJIE INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510253707.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When processing large-scale dynamic graph data, the prior art has problems such as low storage efficiency, slow query response, and unbalanced resource utilization.

Method used

Through innovative graph data compression and coding mechanism, path index structure of searchable coding trees, query strategy optimization based on reinforcement learning, edge cache mechanism for multi-dimensional similarity calculation and other technologies, efficient storage and rapid query of dynamic graph data can be achieved.

Benefits of technology

It achieves a data compression rate of up to 40-60%, improves path query efficiency, improves query efficiency by 3-5 times, enhances cache hit rate, reduces load imbalance, and improves the overall system throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067402A_ABST
    Figure CN120067402A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of database systems, distributed systems and graph data processing, in particular to a distributed storage and query optimization system and method for dynamic graph data, and the distributed storage and query optimization system comprises a dynamic graph data management module, a graph data compression module, an edge cache module, a path index module, a reinforcement learning module and a graph coding module. The dynamic graph data management module updates node data in real time to ensure data consistency; the graph data compression module realizes the compression rate of 40-60% through an innovative coding mechanism, and the storage demand is reduced; the edge cache module stores frequent query sub-graphs, so that the access speed is increased; the path index module adopts a searchable coding tree structure, so that the query complexity is reduced to be close to O (logn), and the efficiency is remarkably improved; the reinforcement learning module dynamically optimizes a strategy according to historical query, so that the query efficiency is improved by 3-5 times; the graph coding module provides a structure coding scheme, efficient data management, compression, caching and query are integrally achieved, and the graph data processing performance is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of database systems, distributed systems, and graph data processing, and particularly to a distributed storage and query optimization system and method for dynamic graph data. Background Art

[0002] With the rapid development of applications such as social networks, knowledge graphs, and the Internet of Things, the scale and complexity of dynamic graph data have been continuously increasing. Dynamic graph data refers to graph-structured data that changes over time, where nodes and edges may be added, deleted, or modified at any time. Existing graph data processing systems face problems such as low storage efficiency, slow query response, and unbalanced resource utilization when dealing with large-scale dynamic graph data.

[0003] Traditional graph data storage methods usually adopt relational databases or dedicated graph databases. Although relational databases are mature and stable, their design concepts are not suitable for processing graph-structured data, especially when facing complex path queries, the efficiency is relatively low; although dedicated graph databases are optimized for graph structures, when dealing with large-scale dynamically changing graph data, they still face problems such as poor distributed scalability, fixed caching strategies, and insufficient query optimization.

[0004] The dynamic graph data processing systems in the prior art mainly have the following problems: First, there is a lack of an efficient compression algorithm for the characteristics of dynamic graph data, resulting in waste of storage space; second, the path index structure is not optimized enough, affecting query efficiency; third, the caching strategy is fixed and cannot be dynamically adjusted according to the access pattern; fourth, the query optimization strategy is simple and cannot adapt to diverse query requirements; fifth, the partitioning algorithm is not flexible enough, resulting in load imbalance.

[0005] Therefore, there is an urgent need for a distributed storage and query optimization system that can efficiently process large-scale dynamic graph data to solve the above problems. Summary of the Invention

[0006] The purpose of the present invention is to provide a distributed storage and query optimization system and method for dynamic graph data, aiming to solve the problems of low storage efficiency, slow query response, and unbalanced resource utilization of dynamic graph data in the prior art. Through innovative graph data compression and encoding mechanisms, path index structures of searchable encoding trees, query strategy optimization based on reinforcement learning, edge caching mechanisms for multi-dimensional similarity calculation, and other technologies, efficient storage and fast query of dynamic graph data are realized.

[0007] The present invention proposes a distributed storage and query optimization system for dynamic graph data, including: A dynamic graph data management module, which is used to manage the latest graph data stored on each node, update the data block numbers of the storage nodes according to the updated edge data, and update the path index module and the cache module; A graph data compression module, which is used to compress graph data and store it in a data block, where one data block corresponds to one storage node; An edge cache module, which is used to store frequently queried subgraphs; A path index module, which is used to maintain index information of graph paths and quickly query graph data; A reinforcement learning module, which is used to record historical query information and dynamically select query strategies; A graph encoding module, which is used to provide an encoding scheme for the graph structure.

[0008] Preferably, the compression encoding of the graph data compression module includes: Initialization encoding: Traverse and store the entire graph data, and use the encoding of each node as the bucket value of the hash bucket corresponding to the node; Update encoding: Calculate the number of neighbor nodes of the current node, and search for the optimal hash bucket in the hash bucket according to the number of neighbor nodes of the node. When searching for an empty bucket, judge whether the bucket value of the hash bucket is greater than the encoding of the current node, and replace it in the encoding order to achieve compression.

[0009] Preferably, the method for traversing and storing the entire graph data includes: Visit each node of the graph data in the form of depth-first traversal or breadth-first traversal, where each node corresponds to a hash bucket and an encoding; Use the encoding of the current node to represent the attributes of the current node, and use the hash bucket corresponding to the current node to represent the neighbor nodes of the current node; When updating the neighbor nodes of the current node, while ensuring that the bucket value of the hash bucket corresponding to the current node is the smallest encoding, store the bucket value as the bucket value corresponding to the number of the node.

[0010] Preferably, the path index module includes: Establish a hash bucket for each node using the structure of a searchable encoding tree, and store the path corresponding to the current node in the key-value storage of each hash bucket; The hash bucket corresponding to each non-leaf node in the searchable encoding tree stores the path encoding of its child nodes. The distance between this node and the encoding of the current node is recorded as 1, and the path encoding of each leaf node is used as the path of the current node; A query path part, which is used to retrieve the path encoding layer by layer and obtain the target path when querying the path.

[0011] Preferably, the query path part includes: Retrieve the last node of the path in the path index table; Use the last node as the parent node and obtain all child nodes corresponding to the parent node; Calculate the similarity between the codes of each child node and the last node, and obtain the leaf node with the largest similarity as the current node; When the code of the current node corresponds to the code of the parent node, update the current node and use the current node as the parent node; Continuously update the parent node of the current node to the current node until the parent node number matches the next node number in the path, then obtain the remaining nodes in the path and generate the query result.

[0012] Preferably, the edge cache module includes: A similarity calculation unit for calculating the similarity threshold of each code using the graph data encoding and finding the K nearest neighbor nodes of each code according to the similarity threshold; A frequency statistics unit for obtaining the number of times the node is cached according to the access frequency of each node; A data storage unit for storing the data of each node in the edge node to implement subgraph caching.

[0013] Preferably, the reinforcement learning module includes: A query pattern recognition unit for analyzing historical queries and identifying common query patterns and access patterns; A policy selection unit for selecting the optimal query policy based on the query pattern and the system state; A learning optimization unit for continuously updating and optimizing the policy selection model to improve the query efficiency.

[0014] Preferably, the dynamic graph data management module includes: A data update unit for storing the numbers at the source nodes by reading the data of each edge and storing all nodes distributively according to the path; A record management unit for numbering the newly inserted edge records and updating the data stored for all nodes; A synchronization processing unit for simultaneously updating the data of the edge cache module and the path index module.

[0015] Preferably, the graph encoding module includes: A node encoding unit for assigning a unique code to each node in the graph to represent the node attributes; A topological relationship encoding unit for storing the topological relationship between nodes through a hash bucket; An encoding optimization unit for dynamically adjusting the encoding scheme to improve the compression efficiency and query performance.

[0016] A distributed storage and query optimization method for dynamic graph data, including: Compress the graph data and store the graph data encoding as the graph partition key value in the nodes; Calculate the hash buckets of each node according to the graph partition, calculate the similarity threshold of each encoding using the graph data encoding, and find the K nearest neighbors of each encoding according to the similarity threshold; Obtain the number of times each node is cached according to the access frequency of each node, and store the data of each node in the edge node; According to the graph path query times, record the frequent query paths of each node, and generate the subgraph data cached by the edge node; According to the amount of compressed data when storing the graph data, taking the compression ratio of the dynamic graph data as the objective function and the edge cache size as the limit, select the data compression ratio; Establish a path index for the dynamic graph data, and calculate the target path according to the query path to achieve efficient query.

[0017] The present invention has the following beneficial effects: 1. Through an innovative graph data compression and encoding mechanism, a data compression rate of up to 40 - 60% is achieved, significantly reducing the storage space requirements; 2. Adopt a path index structure of a searchable coding tree, reduce the path query complexity from the traditional O(n²) to nearly O(logn), and greatly improve the query efficiency; 3. Based on the optimization of the query strategy by reinforcement learning, the system can dynamically select the optimal query strategy according to historical query information, and the query efficiency is increased by 3 - 5 times; 4. The edge caching mechanism for multi - dimensional similarity calculation improves the cache hit rate by 40 - 60%, reduces the network transmission overhead and query latency; 5. The dynamic adaptive multi - dimensional graph partition algorithm realizes the balanced distribution of data, the overall throughput of the system is increased by more than 30%, and the load imbalance degree is reduced by 60%. Brief Description of the Drawings

[0018] Figure 1 It is a schematic diagram of the overall architecture of the distributed storage and query optimization system for the dynamic graph data of the present invention; Figure 2 It is a schematic diagram of the working process of the graph data compression module of the present invention; Figure 3 It is a schematic diagram of the structure of the searchable coding tree of the present invention; Figure 4 It is a schematic diagram of the query process of the path index module of the present invention; Figure 5 It is a schematic diagram of the working principle of the edge cache module of the present invention; Figure 6 It is a schematic diagram of the query strategy selection process of the reinforcement learning module of the present invention; Figure 7Flowchart of the method for optimizing distributed storage and query of dynamic graph data of the present invention; Figure 8 Simulation effect diagram of the distributed storage and query optimization system for dynamic graph data of the present invention. Detailed implementation manners

[0019] Please refer to the appendix Figure 1-8 , the distributed storage and query optimization system for dynamic graph data of the present invention includes: a dynamic graph data management module 1, a graph data compression module 2, an edge cache module 3, a path index module 4, a reinforcement learning module 5, and a graph encoding module 6.

[0020] The dynamic graph data management module 1 is used to manage the latest graph data stored on each node, update the data block numbers of the storage nodes according to the updated edge data, and update the path index module 4 and the cache module 3. The graph data compression module 2 is used to compress the graph data and store it in a data block, and one data block corresponds to one storage node. The edge cache module 3 is used to store frequently queried subgraphs. The path index module 4 is used to maintain the index information of the graph paths and quickly query the graph data. The reinforcement learning module 5 is used to record historical query information to dynamically select query strategies. The graph encoding module 6 is used to provide an encoding scheme for the graph structure.

[0021] The compression encoding of the graph data compression module 2 mainly includes two parts: initialization encoding 21 and update encoding 22.

[0022] The initialization encoding 21 traverses and stores the entire graph data, and uses the encoding of each node as the bucket value of the corresponding hash bucket of the node. In the preferred embodiment of the present invention, the initialization encoding adopts an integer encoding method, starting from 0 and incrementing, and assigns a unique encoding to each node. For example, for graph data with 1000 nodes, the node encoding range is 0-999.

[0023] The update encoding 22 calculates the number of neighbor nodes of the current node, and searches for the optimal hash bucket from the hash bucket according to the number of neighbor nodes of the node. When searching for an empty bucket, it judges whether the bucket value of the hash bucket is greater than the current node encoding, and replaces it in the encoding order to achieve compression. Preferably, when the bucket value of the hash bucket is greater than the current node encoding, the bucket value of the hash bucket is replaced with the current node encoding, so as to minimize the encoding.

[0024] The method for traversing and storing the entire graph data includes: accessing each node of the graph data in the form of depth - first traversal or breadth - first traversal, where each node corresponds to a hash bucket and a code; using the code of the current node to represent the attributes of the current node, and at the same time using the hash bucket corresponding to the current node to represent the neighbor nodes of the current node; when updating the neighbor nodes of the current node, while ensuring that the bucket value of the hash bucket corresponding to the current node is the minimum code, storing the bucket value as the bucket value corresponding to the node number.

[0025] In an embodiment of the present invention, the hash bucket is implemented using open addressing, and the hash function uses quadratic probing to handle conflicts. The hash function can be expressed as: QUOTE , where, QUOTE is the basic hash function, QUOTE is the key value, QUOTE is the probing sequence number, QUOTE and QUOTE are constants (usually QUOTE is the size of the hash table.

[0026] Searching for the optimal hash bucket from the hash bucket includes: traversing all neighbor nodes of the current node, obtaining the sizes of the bucket values stored in all hash buckets of the current node; when there is a situation where there is no neighbor node belonging to the same hash bucket as the current node, stop traversing.

[0027] The graph data compression algorithm of the present invention can effectively reduce the storage space requirement. Experiments show that for graph data containing millions of nodes and edges, the compression ratio can reach 40 - 60%, significantly better than the 20 - 30% compression ratio of traditional compression methods.

[0028] The path index module 4 uses the structure of a searchable coding tree to establish the hash bucket of each node, and the key - value storage of each hash bucket stores the path corresponding to the current node. As Figure 3 shown, the hash bucket corresponding to each non - leaf node in the searchable coding tree stores the path code of its child nodes, the distance between this node and the code of the current node is recorded as 1, and the path code of each leaf node is used as the path of the current node.

[0029] In a preferred embodiment of the present invention, the searchable coding tree adopts a B + tree structure, and each internal node contains at most QUOTE Child nodes (QUOTE usually take values of 3 or 5), and leaf nodes contain actual path data. The B+ tree structure can ensure that the time complexity of query and update operations is O(logn), where n is the number of nodes.

[0030] The query path part of the path index module 4 includes: retrieving the last node of the path in the path index table; taking the last node as the parent node and obtaining all child nodes corresponding to the parent node; calculating the similarity between the encoding of each child node and the last node, and obtaining the leaf node with the largest similarity as the current node; when the encoding of the current node corresponds to the encoding of the parent node, updating the current node and taking the current node as the parent node; continuously updating the parent node of the current node to the current node until the parent node number matches the next node number of the path, obtaining the remaining nodes in the path, and generating a query result.

[0031] The calculation of node encoding similarity uses the cosine similarity formula: QUOTE , where, QUOTE and QUOTE are node encoding vectors, QUOTE represents the dot product of vectors, QUOTE and QUOTE respectively represent the modulus lengths of vectors QUOTE and QUOTE The similarity value ranges between [-1, 1], and the larger the value, the more similar. In practical applications, when the similarity is greater than 0.8, the nodes can be considered highly similar; when the similarity is between 0.5 - 0.8, the nodes are considered moderately similar; when the similarity is less than 0.5, the nodes are considered significantly different.

[0032] Through this hierarchical path index structure, the present invention significantly improves the path query efficiency. For graph data containing millions of nodes, the average response time of path queries using traditional methods is several hundred milliseconds, while the method of the present invention reduces it to dozens of milliseconds, with a performance improvement of about 10 times.

[0033] The edge cache module 3 includes a similarity calculation unit 31, a frequency statistics unit 32, and a data storage unit 33.

[0034] The similarity calculation unit 31 calculates the similarity threshold for each encoding using the graph data encoding, and finds the K nearest neighbor nodes for each encoding according to the similarity threshold. In the embodiments of the present invention, the value of K is usually set between 5 and 20, and is dynamically adjusted according to system resources and query patterns. The similarity threshold is generally set between 0.7 and 0.9. When the similarity of node encodings exceeds this threshold, they are considered to be neighboring nodes.

[0035] The frequency statistics unit 32 obtains the number of times the node is cached according to the access frequency of each node. The access frequency can be implemented by a simple counter, or a time decay function can be used to give higher weight to the most recent access. In the preferred embodiment of the present invention, an exponential decay function is used to calculate the weighted access frequency: QUOTE , where, QUOTE is the frequency count of the QUOTE th access, QUOTE is the timestamp of the QUOTE th access, QUOTE is the current timestamp, QUOTE is the decay coefficient (usually set between 0.01 and 0.1).

[0036] The data storage unit 33 stores the data of each node in the edge node to implement subgraph caching. The cache capacity is usually set to 5% - 20% of the total data volume, and the specific value is dynamically adjusted according to system resources and access patterns.

[0037] Through this edge caching mechanism of multi-dimensional similarity calculation, the present invention significantly improves the cache hit rate. Compared with the hit rate of 30% - 40% of the traditional LRU or LFU cache policies, the cache mechanism of the present invention can achieve a hit rate of 50% - 70%, significantly reducing the network transmission overhead and query latency.

[0038] The reinforcement learning module 5 includes a query pattern recognition unit 51, a policy selection unit 52, and a learning optimization unit 53.

[0039] The query pattern recognition unit 51 analyzes historical queries and identifies common query patterns and access patterns. In the embodiments of the present invention, query patterns can be divided into four basic types: point query, neighbor query, path query, and subgraph query. The system automatically identifies them by analyzing the structural features of the query statements.

[0040] The policy selection unit 52 selects the optimal query policy based on the query pattern and system state. The query policies include various options such as index priority, cache priority, and partition priority. The policy selection uses the multi-armed bandit algorithm, which can be expressed as: QUOTE , where QUOTE is the selected action (policy), QUOTE is the estimated value of action QUOTE , QUOTE is the exploration parameter (usually set to 12), QUOTE is the current time step, QUOTE is the number of times action QUOTE has been selected.

[0041] The learning and optimization unit 53 continuously updates and optimizes the policy selection model to improve the query efficiency. The learning and optimization uses the Q-learning algorithm, and the update rule is: QUOTE , where QUOTE is the Q value of selecting action QUOTE in state QUOTE , QUOTE is the learning rate (usually set to 0.1 - 0.3), QUOTE is the obtained reward, QUOTE is the discount factor (usually set to 0.9 - 0.99),QUOTE is the maximum Q value of the next state.

[0042] Through this query optimization mechanism based on reinforcement learning, the present invention realizes the dynamic adaptability of the query policy. As the system running time increases, the query efficiency is gradually improved, and for the query of repeated patterns, the performance improvement can reach 3 - 5 times.

[0043] The dynamic graph data management module 1 includes a data update unit 11, a record management unit 12, and a synchronization processing unit 13.

[0044] The data update unit 11 reads the data of each edge, stores the numbers at the source nodes respectively, and stores all nodes distributively according to the paths. In the embodiments of the present invention, the edge data usually includes information such as the source node ID, target node ID, weight, and attributes. The system distributes the data to different storage nodes according to the hash value of the node ID.

[0045] The record management unit 12 numbers the newly inserted edge records and updates the data stored in all nodes. The edge record numbering adopts a combination of timestamp and incrementing sequence number to ensure global uniqueness.

[0046] The synchronization processing unit 13 updates the data of the edge cache module 3 and the path index module 4 simultaneously. To ensure data consistency, the update operation adopts a two-phase commit protocol. First, the update is applied to the main storage, and after success, it is synchronously updated to the cache and index.

[0047] Through the coordinated work of the dynamic graph data management module 1, the present invention realizes the efficient management and update of dynamic graph data. The system can process thousands to tens of thousands of edge update operations per second, meeting the requirements of high-concurrency scenarios.

[0048] The graph encoding module 6 includes a node encoding unit 61, a topological relationship encoding unit 62, and an encoding optimization unit 63.

[0049] The node encoding unit 61 assigns a unique code to each node in the graph to represent the node attributes. In the embodiments of the present invention, the node encoding adopts an integer encoding method. For nodes with special meanings (such as central nodes or boundary nodes), specific ranges of codes can be assigned for subsequent processing.

[0050] The topological relationship encoding unit 62 stores the topological relationships between nodes through hash buckets. The encoding of topological relationships considers factors such as the connection type, direction, and weight between nodes, providing a rich representation of the graph structure.

[0051] The encoding optimization unit 63 dynamically adjusts the encoding scheme to improve the compression efficiency and query performance. The encoding optimization adopts a local re-encoding strategy. When nodes in a certain area are frequently updated or queried, the encoding of this area is optimized to improve the local performance.

[0052] The distributed storage and query optimization method for the dynamic graph data of the present invention includes the following steps: Step 1: Compress the graph data and store the encoded graph data as a graph partition key value in the nodes.

[0053] In this step, the system first compresses and encodes the entire graph data. The compression process adopts the aforementioned two stages of initialization encoding and update encoding, and stores the neighbor relationships of nodes through hash buckets to achieve efficient compression. The compressed graph data uses the node encoding as the partition key value and is stored in distributed nodes.

[0054] Step 2: Calculate the hash bucket for each node according to the graph partition, calculate the similarity threshold for each encoding using the graph data encoding, and find the K nearest neighbor nodes for each encoding according to the similarity threshold.

[0055] In this step, the system calculates the hash bucket for each node as the storage index of the node data. At the same time, the system calculates the similarity between node encodings, finds the K most similar nodes for each node, and provides a basis for subsequent cache optimization. The similarity calculation uses the aforementioned cosine similarity formula, the K value is usually set to 5 - 20, and the similarity threshold is set between 0.7 - 0.9.

[0056] Step 3: Obtain the number of times each node is cached according to the access frequency of each node, and store the data of each node in the edge node.

[0057] In this step, the system counts the access frequency of each node, calculates the weighted access frequency using the aforementioned exponential decay function, and determines the cache priority of the node according to the access frequency. The nodes with high priority and their neighboring nodes are stored in the edge cache nodes to achieve cache optimization at the subgraph level.

[0058] Step 4: According to the number of graph path queries, record the frequent query paths of each node and generate the subgraph data cached in the edge node.

[0059] In this step, the system analyzes the historical query records, identifies the frequent query path patterns, and preferentially caches the subgraph data related to these paths in the edge nodes. This cache strategy based on query patterns greatly improves the cache hit rate and reduces the network transmission overhead.

[0060] Step 5: According to the amount of compressed data when storing the graph data, take the compression ratio of the dynamic graph data as the objective function and the edge cache size as the limit to select the data compression ratio.

[0061] In this step, the system dynamically adjusts the data compression ratio according to the actual data characteristics and system resources. The compression ratio objective function can be expressed as: QUOTE , where, QUOTE is the compression ratio, QUOTE is the size of the compressed data, QUOTE It is the size of the original data. The system will select the optimal compression ratio while ensuring query performance, usually between 40% and 60%.

[0062] Step 6: Establish a dynamic graph data path index and calculate the target path according to the query path to achieve efficient querying.

[0063] In this step, the system constructs a searchable coding tree as the path index to support efficient path querying. The query process is executed according to the aforementioned path index query mechanism, retrieving the path encoding layer by layer, calculating the node similarity, and quickly locating the target path.

[0064] Through the collaborative work of the above six steps, the method of the present invention realizes the efficient storage and fast query of dynamic graph data. The system can dynamically adjust the compression strategy, cache strategy, and query strategy according to the data characteristics and query patterns to provide optimal performance.

[0065] The following uses a specific embodiment to illustrate the working process and effect of the present invention.

[0066] Suppose there is a social network graph data containing 1 million nodes and 10 million edges, and it is necessary to frequently process user relationship queries and path discovery tasks.

[0067] 1. Graph data compression and storage: First, the system compresses and encodes the graph data. Using depth-first traversal, integer encodings (0 - 999999) are assigned to each node, and corresponding hash buckets are created. The size of the hash table is set to 2 20 , and quadratic probing is used to handle collisions. By updating the encoding process, the system achieves a compression ratio of 50%, reducing the original storage space requirement of 20GB to 10GB.

[0068] 2. Path index construction: The system constructs a searchable coding tree as the path index. The order of the B+ tree is set to 5, and the tree height is 4, which can effectively cover all possible paths. For each non-leaf node, its hash bucket stores the path encodings of up to 100 child nodes; the leaf nodes store the actual path data. This index structure reduces the path query response time from 200ms of the traditional method to 20ms.

[0069] 3. Edge cache optimization: The system calculates the similarity of node encodings and sets the similarity threshold to 0.8. For each node, it finds the K = 10 most similar nodes. Meanwhile, the system counts the node access frequencies, with the decay coefficient λ set to 0.05. Based on the weighted access frequencies, the system allocates cache space for the frequently accessed nodes and their neighboring nodes, with the total cache size being 15% of the original data (about 1.5GB). This caching strategy achieves a hit rate of 65%, significantly reducing the network transmission overhead.

[0070] 4. Query strategy optimization: The system optimizes the query strategy based on a reinforcement learning mechanism. The learning rate α is set to 0.2, the discount factor γ is set to 0.95, and the exploration parameter c is set to 1.5. The system identifies four main query patterns: user relationship query, shortest path query between two points, community discovery query, and recommended relationship query. After a period of learning and optimization, the system can automatically select the optimal strategy for different query patterns, and the average query response time is reduced by 75%.

[0071] 5. Overall system performance: Combining the above optimization measures, when the system processes social network graph data with 1 million nodes and 10 million edges, it has the following performance metrics: Data compression ratio: 50%; Path query response time: 20ms; Cache hit rate: 65%; Query performance improvement: 4 times; System throughput improvement: 35%; Load balancing degree improvement: 60%.

[0072] This embodiment verifies the superior performance of the present invention in processing large-scale dynamic graph data, and is particularly suitable for application scenarios in fields such as social networks, knowledge graphs, and network topologies.

[0073] Simulation comparison between the embodiment and the comparative example of the distributed storage and query optimization system for dynamic graph data: Test environment: The simulation experiment is carried out in the following environment: Hardware environment: 10-node distributed cluster, each node is configured with an Intel Xeon E5-2680 v4 CPU (14 cores), 128GB of memory, and a 10Gbps network connection Software environment: Based on the Linux CentOS 7.6 operating system, Java 1.8, Hadoop 3.1.2 Test data set: To comprehensively evaluate the system performance, 3 different types of large-scale graph data are selected: Social network graph data set: A user relationship graph extracted from social network platforms (such as Weibo, WeChat, etc.), containing 10 million nodes and 100 million edges Knowledge Graph Dataset: A subset of the knowledge graph extracted from DBpedia, containing 5 million nodes and 70 million edges. Network Topology Dataset: A simulated topology graph of the Internet backbone, containing 1 million nodes and 15 million edges. Example: A complete implementation solution of the present invention, including key technologies such as graph data compression, searchable coding tree path indexing, query optimization based on reinforcement learning, and multi-dimensional similarity edge caching.

[0074] Comparative Example 1: A distributed graph data storage and query system based on Neo4j, using standard graph data storage and Cypher query language.

[0075] Comparative Example 2: A distributed graph computing system based on Spark GraphX, using the Pregel model for graph data processing.

[0076] Evaluation metrics include: Storage efficiency: Storage space occupancy (GB); Data compression ratio (%); Data import time (minutes); Query performance: Point query response time (milliseconds); Neighbor query response time (milliseconds); Path query response time (milliseconds); Subgraph query response time (milliseconds); Query throughput (queries per second); System scalability: Linear scalability (performance improvement ratio when nodes are expanded); Load balance degree (%, standard deviation of load between nodes / average load); Dynamic update ability: Number of edge updates per second (edges / second); Update latency (milliseconds); Query consistency after update (%); Simulation method: Storage efficiency test: Import the complete dataset into each system and record the storage space occupancy and import time.

[0077] Query performance test: For each query type, generate 1000 random queries and record the average response time and throughput.

[0078] Scalability test: Start from 2 nodes and gradually increase to 10 nodes, measuring the performance changes.

[0079] Dynamic update test: Simulate 10,000 edge updates per second and measure the system's processing capacity and consistency.

[0080] The simulation results are as follows: Table 1: Comparison of storage efficiency of three schemes on different datasets Table 2: Comparison of query performance of three schemes on social network datasets Table 3: Comparison of query performance of three schemes on knowledge graph datasets Table 4: Comparison of query performance of three schemes on network topology datasets Table 5: Comparison of system scalability of three schemes (social network datasets) Table 6: Comparison of dynamic update capabilities of three schemes Index Embodiment of the present invention Comparative Example 1 Comparative Example 2 Number of edge updates per second (edges / second) 42500 12800 18500 Update latency (ms) 24.6 85.3 67.2 Query consistency after update (%) 99.8 92.5 95.3 The analysis of the simulation results is as follows: The embodiments of the present invention show significant advantages in terms of storage efficiency. As can be seen from Table 1, on three different types of graph datasets, the storage space occupied by the embodiments of the present invention is significantly lower than that of the comparative schemes. The compression ratio reaches 53.8% - 57.4%, which is much higher than 10.5% - 14.8% of Comparative Example 1 and 22.3% - 25.6% of Comparative Example 2. This is mainly due to the innovative hash bucket encoding mechanism adopted by the graph data compression module of the present invention, which realizes efficient compression through depth traversal and optimal hash bucket search.

[0081] In terms of data import time, although the embodiments of the present invention need to perform additional encoding and compression processing, due to the adoption of an efficient parallel processing mechanism, the import time is reduced by 42.8% and 36.0% compared with Comparative Example 1 and Comparative Example 2 respectively, significantly improving the initialization efficiency of the system.

[0082] In terms of query performance, the embodiments of the present invention perform excellently in all query types. Especially in complex path queries and subgraph queries, the performance advantages of the embodiments of the present invention are more obvious. For example, in the path query on the social network dataset, the average response time of the embodiments of the present invention is 21.5ms, while those of Comparative Example 1 and Comparative Example 2 are 187.6ms and 135.8ms respectively, and the performance improvements reach 8.7 times and 6.3 times respectively.

[0083] This advantage mainly stems from the searchable coding tree structure adopted by the path index module of the present invention, as well as the ability of the reinforcement learning module to dynamically select the optimal query strategy based on historical query information. In addition, the edge caching module significantly improves the cache hit rate through multi-dimensional similarity calculation and sub-graph level caching, further reducing the query response time.

[0084] The query throughput metric more intuitively reflects the overall system performance. In the embodiments of the present invention, the query throughputs on three datasets are 3,650, 4,120, and 4,580 queries per second respectively, which are 3.5 times higher than that of Comparative Ratio 1 on average and 2.9 times higher than that of Comparative Ratio 2.

[0085] Table 5 shows the performance improvement of the three schemes when the nodes are expanded. When the number of nodes increases from 2 to 4, the query throughput improvement ratio of the embodiments of the present invention is 1.92, approaching the theoretical optimal value of 2, indicating that the system has excellent horizontal scaling ability. In contrast, the improvement ratios of Comparative Ratio 1 and Comparative Ratio 2 are 1.65 and 1.78 respectively, with lower expansion efficiency. As the number of nodes further increases, the expansion efficiency of all schemes decreases, but the embodiments of the present invention always maintain the highest expansion efficiency.

[0086] In terms of load balancing degree, the embodiments of the present invention maintain a low load imbalance degree under various node configurations. Even in the 10-node configuration, the load imbalance degree is only 10.5%, far lower than 32.4% of Comparative Ratio 1 and 25.6% of Comparative Ratio 2. This benefits from the dynamic adaptive multi-dimensional graph partitioning algorithm of the present invention, which can intelligently adjust the partitioning strategy according to the data access pattern to ensure system load balance.

[0087] In terms of dynamic update ability, the embodiments of the present invention also perform excellently. It can process 42,500 edge updates per second, which is 3.3 times that of Comparative Ratio 1 and 2.3 times that of Comparative Ratio 2. The update delay is only 24.6 ms, which is reduced by 71.2% and 63.4% compared with Comparative Ratio 1 and Comparative Ratio 2 respectively. More importantly, the embodiments of the present invention can still maintain a query consistency of 99.8% in a high-concurrency update environment, significantly higher than the comparative schemes.

[0088] This advantage stems from the efficient incremental update mechanism adopted by the dynamic graph data management module of the present invention, as well as the close collaborative work with the path index and cache modules. The system can quickly update the data block numbers of storage nodes after edge data updates, synchronously update the index and cache, ensuring data consistency while minimizing the update overhead.

[0089] According to the simulation results, in the knowledge graph dataset, the embodiments of the present invention exhibit the optimal comprehensive performance. The specific configuration is as follows: The graph data compression adopts depth-first traversal, and the node coding adopts integer coding; The size of the hash table is set to 2 24 , and the quadratic probing method is used for the hash function; The searchable coding tree adopts a B+ tree structure with the order set to 5; The similarity threshold is set to 0.85, and the K value (number of neighboring nodes) is set to 15; The decay coefficient λ for calculating the access frequency is set to 0.05; The cache capacity is set to 12% of the total data volume; The learning rate α of the reinforcement learning module is 0.2, and the discount factor γ is 0.95; The exploration parameter c of the multi-armed bandit algorithm is set to 1.5; Under this configuration, the system demonstrates excellent storage efficiency, query performance, and scalability, and is particularly suitable for processing dynamic graph data such as knowledge graphs.

[0090] Based on the above analysis, the distributed storage and query optimization system for dynamic graph data of the present invention has the following significant advantages compared with the prior art: High storage efficiency: The average compression rate reaches over 55%, which is 30 - 45 percentage points higher than the comparison scheme; Excellent query performance: The path query and subgraph query performance are improved by 6 - 9 times, and the query throughput is improved by 2.9 - 3.5 times; Outstanding scalability: The performance improvement is nearly linear when the number of nodes increases, and the load imbalance degree always remains at a low level; Strong dynamic update ability: It can process more than 40,000 edge updates per second while maintaining extremely high query consistency; Through the innovative design and collaborative work of the graph data compression module, path index module, edge cache module, reinforcement learning module, and graph coding module, the present invention successfully solves the key problems such as low storage efficiency, slow query response, and unbalanced resource utilization of large-scale dynamic graph data, and provides an efficient solution for graph data processing in fields such as social networks, knowledge graphs, and network topologies.

[0091] It should be noted that: The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A distributed storage and query optimization system for dynamic graph data, characterized in that: include: Dynamic graph data management module, used to manage the latest graph data stored on each node, update the data block number of the storage node according to the edge data, and update the path index module and cache module; A graph data compression module is used to compress graph data and store it into a data block. One data block corresponds to one storage node. Edge cache module, used to store frequently queried subgraphs; The path index module is used to maintain the index information of the graph path and quickly query the graph data; Reinforcement learning module, used to record historical query information and dynamically select query strategies; The graph encoding module is used to provide an encoding scheme for the graph structure.

2. The system according to claim 1, characterized in that The compression coding of the graph data compression module includes: Initialize the code: traverse and store the entire graph data, and use the code of each node as the bucket value of the hash bucket corresponding to the node; Update the code: Calculate the number of neighbor nodes of the current node, and search for the optimal hash bucket from the hash bucket based on the number of neighbor nodes of the node. When searching for an empty bucket, determine whether the bucket value of the hash bucket is greater than the current node code, and replace it in the coding order to achieve compression.

3. The system according to claim 2, characterized in that The methods for traversing and storing the entire graph data include: Visit each node of the graph data in the form of depth-first traversal or breadth-first traversal, where each node corresponds to a hash bucket and a code; The encoding of the current node is used to represent the attributes of the current node, and the hash bucket corresponding to the current node is used to represent the neighbor nodes of the current node; When updating the neighboring nodes of the current node, the bucket value is stored as the bucket value corresponding to the node number while ensuring that the bucket value of the hash bucket corresponding to the current node is the minimum code.

4. The system according to claim 1, characterized in that The path index module includes: A searchable coding tree structure is used to establish a hash bucket for each node. The key value of each hash bucket stores the path corresponding to the current node. The hash bucket corresponding to each non-leaf node in the searchable coding tree stores the path code of its child nodes. The distance between the node and the code of the current node is recorded as 1. The path code of each leaf node is used as the path of the current node. The query path part is used to retrieve the path code layer by layer to obtain the target path when querying the path.

5. The system according to claim 4, characterized in that The query path portion includes: Retrieve the last node of the path from the path index table; Take the last node as the parent node and get all the child nodes corresponding to the parent node; Calculate the similarity between the encoding of each child node and the last node, and obtain the leaf node with the largest similarity as the current node; When the code of the current node corresponds to the code of the parent node, update the current node and use the current node as the parent node; The parent node of the current node is continuously updated to the current node until the parent node number matches the next node number of the path, and then the remaining nodes in the path are obtained to generate the query result.

6. The system according to claim 1, characterized in that The edge cache module includes: A similarity calculation unit, used to calculate a similarity threshold of each code using the graph data code, and find K neighboring nodes of each code according to the similarity threshold; A frequency statistics unit, used to obtain the number of times each node is cached according to the access frequency of the node; The data storage unit is used to store the data of each node in the edge node to realize subgraph caching.

7. The system according to claim 1, characterized in that The reinforcement learning module includes: A query pattern recognition unit, which is used to analyze historical queries and identify common query patterns and access patterns; A strategy selection unit, for selecting an optimal query strategy based on a query mode and a system state; The learning optimization unit is used to continuously update and optimize the strategy selection model to improve query efficiency.

8. The system according to claim 1, characterized in that The dynamic graph data management module includes: The data update unit is used to read the data of each edge, store the numbers in the source node respectively, and distribute the data to all nodes according to the path; A record management unit, used to number newly inserted edge records and update the data stored in all nodes; The synchronization processing unit is used to simultaneously update the data of the edge cache module and the path index module.

9. The system according to claim 1, characterized in that The graph encoding module comprises: Node coding unit, used to assign a unique code to each node in the graph, indicating the node attributes; A topological relationship encoding unit, used to store the topological relationship between nodes through a hash bucket; The encoding optimization unit is used to dynamically adjust the encoding scheme to improve compression efficiency and query performance.

10. A method for distributed storage and query optimization of dynamic graph data, using the system according to any one of claims 1 to 9, characterized in that: include: Compress the graph data and store the graph data encoding as the graph partition key value in the node; Calculate the hash bucket of each node according to the graph partition, calculate the similarity threshold of each code using the graph data code, and find the K neighboring nodes of each code according to the similarity threshold; The number of times each node is cached is obtained based on the access frequency of each node, and the data of each node is stored in the edge node; According to the number of graph path queries, the frequent query paths of each node are recorded to generate subgraph data cached by edge nodes; According to the amount of data compressed when storing graph data, the compression rate of dynamic graph data is used as the objective function, and the edge cache size is used as the limit to select the data compression ratio; Establish a dynamic graph data path index and calculate the target path based on the query path to achieve efficient query.

Citation Information

Cited By

  • Distribution measurement method and device based on similarity dynamic compression, equipment and medium

    CN120455306A

  • Query request processing method and device, electronic equipment, medium and program product

    CN120631941A