Distributed Data Storage Network Generation Method Based on Knowledge Graph
By building a semantic similarity matrix and dynamic shard adjustment mechanism, optimizing network topology and query paths, the problems of data sharding static and topological static in the existing technology are solved, and the data management efficiency and query performance of distributed storage systems are improved.
Patent Information
- Application Number
- CN202510423344.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-07
AI Technical Summary
The existing distributed data storage network based on knowledge graphs lacks dynamic adjustment capabilities in data sharding, the static nature of network topology optimization leads to high transmission overhead, and the query path optimization fails to fully utilize semantic information, resulting in limited improvement in query efficiency.
By constructing a semantic similarity matrix, a distributed data shard is generated, and an initial network topology is generated based on semantic correlation, the storage system status is dynamically monitored and the network topology is optimized, the query path is optimized using the knowledge graph semantic path, and a collaborative optimization framework is used for global optimization.
It realizes semantic consistency of data distribution and load balancing of storage nodes, reduces data transmission overhead across nodes, improves the communication efficiency and query performance of the system, and ensures the efficiency and stability of the network topology.
Smart Images

Figure CN119938942B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed storage, and specifically to a method for generating a distributed data storage network based on a knowledge graph. Background Art
[0002] As the infrastructure of modern large-scale data processing systems, a distributed data storage network can support efficient distributed data management and access. In application scenarios based on knowledge graphs, knowledge graphs form a data network with strong semantic relevance by modeling entities and their relationships, providing a semantic basis for the optimization of distributed storage networks. The method for generating a distributed data storage network based on a knowledge graph optimizes data sharding, network topology, and query paths using the semantic structure of the graph, aiming to achieve efficient utilization of storage resources and fast response for data access.
[0003] Existing methods for generating a distributed data storage network based on a knowledge graph can improve the efficiency of distributed data storage to a certain extent by using fixed graph partitioning strategies and static network topologies. These methods rely on the global structural characteristics of the knowledge graph to optimize data sharding and distribution, and achieve low cross-node data transmission overhead in a distributed environment.
[0004] However, there are some deficiencies in the existing technology. Its data sharding lacks the ability of dynamic adjustment and is difficult to adapt to the frequent updates of the knowledge graph; the network topology optimization is relatively static and it is difficult to effectively reduce the transmission overhead in a complex dynamic environment; at the same time, the query path optimization and caching mechanism do not fully utilize semantic information, resulting in limited improvement in query efficiency. Summary of the Invention
[0005] In view of the deficiencies of the existing technology, the present invention provides a method for generating a distributed data storage network based on a knowledge graph, which solves the problems of the lack of dynamic adjustment ability of data sharding, the high transmission overhead caused by the static nature of network topology optimization, and the insufficient utilization of semantic information in query path optimization and caching mechanism in the existing technology.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for generating a distributed data storage network based on a knowledge graph, including the following steps:
[0007] S1. Extract entities and their semantic relationships from the knowledge graph to construct a semantic similarity matrix;
[0008] S2. Based on the semantic similarity matrix, construct a weighted graph and use a graph partitioning algorithm to generate distributed data shards;
[0009] S3. Allocate the generated data shards to storage nodes;
[0010] S4. Generate an initial network topology based on the semantic relevance among data shards;
[0011] S5. Dynamically monitor the storage system status and optimize the network topology;
[0012] S6. Optimize the query path based on the semantic path of the knowledge graph;
[0013] S7. Perform global optimization on data distribution, network topology, and query efficiency through a collaborative optimization framework.
[0014] Preferably, the S1 includes:
[0015] Use a graph embedding algorithm to map entities and relationships in the knowledge graph to a low-dimensional vector space;
[0016] Calculate the semantic similarity through the Euclidean distance between node vectors;
[0017] Construct a semantic similarity matrix based on the semantic similarity.
[0018] Preferably, the S2 includes:
[0019] Construct a weighted graph, where the nodes represent entities in the knowledge graph, and the weights of the edges are determined by the semantic similarity between entities;
[0020] Use the Louvain algorithm to perform modularity segmentation on the weighted graph to generate several data shards;
[0021] The division of data shards meets the requirements of maximizing semantic relevance and load balancing of storage nodes.
[0022] Preferably, the S3 includes:
[0023] According to the shard semantic clustering results, allocate data shards with similar semantics to the same storage node;
[0024] Based on the capacity and load status of the storage nodes, adjust the allocation priority of data shards;
[0025] Ensure global load balancing during the data shard allocation process.
[0026] Preferably, the S4 includes:
[0027] Calculate the data sharing degree among storage nodes, and construct the edge weights of the network according to the sharing degree;
[0028] Adopt a weighted minimum spanning tree algorithm to generate an initial network topology;
[0029] Ensure that the generated network topology can minimize the cross-node data transmission overhead.
[0030] Preferably, the S5 includes:
[0031] Dynamically monitor the load, data migration volume, and network latency of storage nodes;
[0032] Define the objective function of the topology optimization problem based on the current network state;
[0033] Adjust the network topology through reinforcement learning methods, including adding edges, deleting edges, or adjusting edge weights;
[0034] Evaluate the cost of data migration that may be caused by topology adjustment and select the adjustment strategy with the minimum cost.
[0035] Preferably, the S6 includes:
[0036] Calculate the priority of the query path using the semantic information of the knowledge graph;
[0037] Plan the query path based on the priority and generate the optimal path using a heuristic search algorithm;
[0038] Build a distributed cache during the query process to store high-frequency query results to reduce duplicate calculations.
[0039] Preferably, the S7 includes:
[0040] Build a multi-objective optimization model and set the optimization objectives as storage load balancing, network latency minimization, and query efficiency maximization;
[0041] Use an evolutionary algorithm to solve the objective function and generate suggestions for data distribution adjustment and network topology optimization;
[0042] Perform real-time adjustment of storage nodes and network connections based on the optimization results.
[0043] Preferably, the optimization objectives of the multi-objective optimization model include:
[0044] The load balancing degree of storage nodes, measured by the maximum value of the node load ratio;
[0045] Network latency, calculated by the transmission latency between storage nodes and the data transmission volume;
[0046] Query efficiency, calculated by the average response time of the query path.
[0047] Preferably, this method supports the dynamic update of the knowledge graph and includes the following steps:
[0048] Dynamically calculate the semantic similarity of newly added or changed entities and update the semantic similarity matrix;
[0049] Adjust the affected shards, allocate the newly added or changed entities to the most relevant shards, and ensure the semantic relevance and load balancing of the shards;
[0050] Optimize the network topology and dynamically adjust the connection structure between storage nodes to reduce data transmission overhead;
[0051] Update the query path and cache policy to ensure the efficiency of query results;
[0052] Monitor the system performance after the update. If the performance degrades, trigger global optimization to restore the stability and efficiency of the system.
[0053] The present invention provides a method for generating a distributed data storage network based on a knowledge graph, which has the following beneficial effects:
[0054] 1. By introducing a semantic similarity matrix and a dynamic sharding adjustment mechanism, the present invention realizes the semantic consistency of data distribution and the load balancing of storage nodes, and at the same time dynamically adapts to the update of the knowledge graph, significantly improving the data management efficiency and storage resource utilization rate of the distributed storage system.
[0055] 2. By optimizing the network topology and dynamically adjusting the connection structure between storage nodes, the present invention reduces the cross-node data transmission overhead, improves the communication efficiency of the system, and ensures the efficiency and stability of the network topology in a dynamic environment.
[0056] 3. Using heuristic query path planning and distributed cache strategy, the present invention reduces the query latency and the frequency of repeated calculations, realizes efficient query in a complex distributed storage network, and enhances the response speed and overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the specification of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] Please refer to the attached Figure 1 , the embodiments of the present invention provide a method for generating a distributed data storage network based on a knowledge graph, including the following steps:
[0060] S1. Extract entities and their semantic relationships from the knowledge graph to construct a semantic similarity matrix. The constructed semantic similarity matrix can quantify the semantic relevance between entities, provide an accurate semantic information basis for subsequent data sharding, ensure that entities with strong semantic relevance can be aggregated together, and improve the storage and query efficiency;
[0061] S2. Based on the semantic similarity matrix, construct a weighted graph and use the graph partitioning algorithm to generate distributed data shards. By partitioning the data based on semantic similarity, entities with high semantic relevance can be aggregated into the same shard, which helps reduce cross-shard queries and improve the semantic consistency of data storage.
[0062] S3. Allocate the generated data shards to storage nodes. Reasonable allocation of shards can make full use of the resources of storage nodes, avoid uneven storage load, and ensure fast access to shard data, thereby improving the performance of the overall storage system.
[0063] S4. Generate an initial network topology based on the semantic correlation between data shards. The generated initial network topology can reduce cross-node data transmission, improve the response efficiency of storage and queries, and lay a foundation for dynamic optimization.
[0064] S5. Dynamically monitor the status of the storage system and optimize the network topology. Dynamic optimization can adapt to changes in the storage system and ensure that the network topology can still maintain an efficient operating state under load fluctuations or changes in access patterns.
[0065] S6. Optimize the query path based on the semantic path of the knowledge graph. By optimizing the semantic path, query latency and cross-node transmission overhead can be reduced, and query efficiency can be improved.
[0066] S7. Globally optimize data distribution, network topology, and query efficiency through a collaborative optimization framework. The collaborative optimization framework can globally improve the storage and query performance of the system in a dynamic environment and ensure the robustness and efficiency of the distributed storage network.
[0067] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, S1 includes:
[0068] Use a graph embedding algorithm to map the entities and relationships in the knowledge graph to a low-dimensional vector space. The embedding algorithm selection: Use an embedding algorithm (such as TransE, TransR, Node2Vec) to map entities and relationships to a vector space:
[0069] TransE assumes , optimize the following loss function:
[0070] ;
[0071] Among them, is the set of positive samples, is the set of negative samples, is a hyperparameter used to control the interval between positive and negative samples. The embedding algorithm can reduce complex high-dimensional semantic relationships to numerical vectors, providing an efficient and intuitive basis for subsequent calculations while preserving the semantic characteristics and topological information in the knowledge graph;
[0072] Calculate semantic similarity through the Euclidean distance between node vectors. The semantic similarity calculation formula:
[0073] Given two entity nodes , Its semantic similarity is calculated through the Euclidean distance of the embedding vectors:
[0074] ;
[0075] Among them, , are the embedding vectors of nodes , respectively, and is a hyperparameter that controls the distance. By calculating semantic similarity, it quantifies the semantic relevance between entities, providing an accurate semantic basis for subsequent sharding optimization and network topology generation, ensuring that the system can better reflect the internal relationships of the knowledge graph;
[0076] Construct a semantic similarity matrix based on semantic similarity. Matrix construction:
[0077] Store the calculated semantic similarity values in the form of a symmetric matrix. The elements of matrix S are defined as:
[0078] ;
[0079] Among them, represents the semantic similarity between nodes and . The semantic similarity matrix structurally stores the semantic relationships between entities in the knowledge graph, providing a direct input for subsequent weighted graph construction and data sharding. At the same time, through sparsification processing, it optimizes storage and computing resources, improving the efficiency and scalability of the system.
[0080] Please refer to Appendix Figure 1 , in a preferred embodiment of the present invention, S2 includes:
[0081] Construct a weighted graph. The nodes represent the entities in the knowledge graph, and the weights of the edges are determined by the semantic similarity between entities. The graph construction method includes:
[0082] Each entity node in the knowledge graph is mapped to a node in the weighted graph;
[0083] Use the values in the semantic similarity matrix S as the weights of the edges , the definition of an edge is as follows:
[0084] ;
[0085] Among them, is the similarity threshold. Edges with weights less than the threshold are not constructed, thus reducing unnecessary edge connections. By constructing a weighted graph, the semantic information between entities in the knowledge graph can be intuitively presented in the form of a graph structure. At the same time, through sparsification, the redundant edges of the graph are reduced, improving the processing efficiency of the graph and providing a good structural foundation for subsequent data sharding;
[0086] Use the Louvain algorithm to perform modular partitioning on the weighted graph to generate several data shards. The goal of the Louvain algorithm is to find the community structure in the weighted graph by maximizing the modularity Q of the graph. The modularity is defined as:
[0087] ;
[0088] Among them, is the edge weight; and are the degrees of nodes i and j respectively; m is the total edge weight; indicates whether nodes i and j belong to the same shard;
[0089] Steps of the algorithm execution:
[0090] Phase 1: Initially, each node is regarded as an independent community;
[0091] Traverse the node set, move the node to the adjacent community, and calculate the modularity gain :
[0092] ;
[0093] Among them, represents the internal edge weight of the community, represents the total weight of the community.
[0094] Phase 2: Based on the local optimization results, merge the nodes in the same community into a supernode, update the graph structure, and repeat Phase 1 until the modularity converges;
[0095] Through modular partitioning, the Louvain algorithm can aggregate entity nodes with high semantic relevance into the same data shard, ensuring semantic consistency within the shard and laying a foundation for efficient querying in the distributed storage network;
[0096] The division of data sharding meets the requirements of maximizing semantic relevance and load balancing of storage nodes. Through the strategies of maximizing semantic relevance and load balancing of shards, it can not only ensure the semantic consistency within data shards, but also make full use of storage node resources, avoid data skew problems in distributed systems, and improve the storage and access efficiency of the overall system.
[0097] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, S3 includes:
[0098] According to the shard semantic clustering results, allocate data shards with similar semantics to the same storage node. Allocating shards according to the semantic clustering results can significantly reduce the transmission of data with high semantic relevance across storage nodes, improve query efficiency, and at the same time reduce network communication overhead;
[0099] Based on the capacity and load status of storage nodes, adjust the allocation priority of data shards. Dynamically adjusting the allocation priority based on node capacity and load status can achieve efficient utilization of storage node resources, avoid excessive load on a single node, and ensure the stability and reliability of the system;
[0100] Guarantee global load balancing during the data shard allocation process. Through the global load balancing strategy, it can effectively avoid resource waste or overload problems of storage nodes in the system, improve the overall stability and performance of the system, and at the same time reduce data access latency caused by load imbalance.
[0101] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, S4 includes:
[0102] Calculate the data sharing degree between storage nodes, and construct the edge weights of the network according to the sharing degree. Through the calculation of data sharing degree and the construction of a weighted graph, it can quantify the data interaction intensity between storage nodes, provide accurate edge weight information for subsequent network topology generation, and ensure the rationality and efficiency of the topology;
[0103] Use the weighted minimum spanning tree algorithm to generate the initial network topology. The initial network topology generated by the weighted minimum spanning tree algorithm can connect all storage nodes with the lowest cross-node transmission cost, provide an efficient communication path for data interaction in the system, and at the same time reduce network complexity;
[0104] Ensure that the generated network topology can minimize the cross-node data transmission overhead. Through optimization and dynamic adjustment, it can ensure that the network topology always minimizes the cross-node transmission overhead, improve the data transmission efficiency of the system, and reduce the response time of data query and storage operations.
[0105] Please refer to the appendix Figure 1 In a preferred embodiment of the present invention, S5 includes:
[0106] Dynamically monitor the load, data migration volume, and network latency of storage nodes, dynamically monitor the load, data migration volume, and network latency, and provide real-time status information for network topology optimization, enabling topology adjustment to accurately reflect the current system performance bottleneck;
[0107] Define the objective function of the topology optimization problem based on the current network state. The objective function that comprehensively considers network latency, load balancing, and migration cost can guide the topology optimization to adjust towards the optimal global performance direction, improving the system operation efficiency and stability;
[0108] Adjust the network topology through reinforcement learning methods, including adding edges, deleting edges, or adjusting edge weights. Reinforcement learning can dynamically adjust the network topology, achieve adaptive optimization under different states, effectively reduce the manual adjustment cost, and improve the topology adjustment efficiency;
[0109] Evaluate the cost of data migration that may be caused after topology adjustment, and select the adjustment strategy with the minimum cost. The migration cost evaluation and incremental migration mechanism can significantly reduce the performance overhead caused by topology adjustment, avoid resource waste caused by frequent migrations, and ensure the continuity of the system at the same time.
[0110] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, S6 includes:
[0111] Calculate the priority of the query path by using the semantic information of the knowledge graph. By calculating the query path priority using the semantic information of the knowledge graph, the query range can be significantly reduced, the high-priority paths can be placed in the front, the query efficiency can be improved, and the system burden can be reduced;
[0112] Plan the query path based on the priority, and use the heuristic search algorithm to generate the optimal path. The heuristic search algorithm combined with the semantic priority can efficiently generate the optimal query path in a complex distributed storage network, reduce the number of hops and cost of cross-node queries, thereby improving the query performance;
[0113] Build a distributed cache during the query process to store high-frequency query results to reduce repeated calculations. The introduction of the distributed cache can reduce the repeated calculations of queries and the frequency of cross-node data access, reduce the network load, and improve the response speed and overall performance of the system.
[0114] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, S7 includes:
[0115] Build a multi-objective optimization model, and set the optimization objectives as storage load balancing, network latency minimization, and query efficiency maximization. The construction of the multi-objective optimization model can simultaneously focus on the storage, network, and query performance of the system, provide the optimization objectives of global performance, and provide comprehensive guidance for system decision-making;
[0116] Use an evolutionary algorithm to solve the objective function, generate suggestions for data distribution adjustment and network topology optimization. The evolutionary algorithm uses NSGA-II (Non-dominated Sorting Genetic Algorithm) to solve multi-objective optimization problems. The evolutionary algorithm can efficiently solve complex multi-objective optimization problems, generate optimization schemes that balance different objectives, and ensure the global optimality of system performance in a dynamic environment;
[0117] Based on the optimization results, make real-time adjustments to the storage nodes and network connections. By making real-time adjustments to the storage nodes and network connections, it can dynamically adapt to changes in system requirements, quickly respond to performance bottlenecks, and improve the robustness and adaptability of the system.
[0118] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, the optimization objectives of the multi-objective optimization model include:
[0119] The load balance degree of the storage nodes, measured by the maximum value of the node load ratio. To avoid resource skew of the storage nodes, the load balance objective is defined as minimizing the load difference between nodes. By calculating the maximum load ratio of the storage nodes Measure the balance degree:
[0120] ;
[0121] Wherein, is the amount of data currently stored in node ; is the total storage capacity of node ;
[0122] Regularly calculate the load balance degree B of the system:
[0123] ;
[0124] If (preset threshold), trigger the load balance optimization mechanism, and restore the balance degree to the ideal range by adjusting the shard distribution. By optimizing the load balance objective, it can avoid the resource skew problem of the storage nodes, ensure that all nodes evenly utilize resources, and thus improve the overall stability and scalability of the system;
[0125] Network latency, calculated by the transmission latency and data transmission volume between storage nodes. To reduce the latency caused by cross-node data access, the network latency objective is defined as the sum of the transmission latencies between nodes:
[0126] ;
[0127] Wherein, represents node and transmission delay between them; represents the amount of data transmitted on edge (i, j). By optimizing the network latency target, the average time for data to access across nodes can be reduced, and the data interaction efficiency of the distributed storage system and the response speed of user queries can be improved;
[0128] Query efficiency, calculated by the average response time of the query path. The query efficiency target is defined as minimizing the average response time of the query path ;
[0129] ;
[0130] wherein, represents the response time of query q, including path search time and data transmission time; Q is the query set;
[0131] Optimize the query path using the semantic information of the knowledge graph, and preferentially select paths with high semantic relevance and low transmission cost:
[0132] ;
[0133] wherein, represents the physical distance of the path;
[0134] Preferentially use cache hit paths in path planning to reduce repeated calculations and cross-node access;
[0135] By optimizing the query efficiency target, the response time of the query path can be reduced, the user experience can be improved, and the system load and communication cost can be reduced.
[0136] Please refer to the appendix Figure 1 , in a preferred embodiment of the present invention, the method supports dynamic update of the knowledge graph and includes the following steps:
[0137] Dynamically calculate the semantic similarity of newly added or changed entities, and update the semantic similarity matrix. By dynamically updating the semantic similarity matrix, it can ensure that the semantic relationships between newly added or changed entities and other entities in the knowledge graph are accurately reflected in the storage system, providing a basis for subsequent shard adjustment and query optimization;
[0138] Adjust the affected shards, and allocate the newly added or changed entities to the most relevant shards to ensure the semantic relevance and load balance of the shards. Dynamically adjusting the shard allocation can ensure that the newly added or changed entities are allocated to the most suitable shards, maintaining both the semantic consistency within the shards and avoiding uneven node loads, improving the robustness and query efficiency of the system;
[0139] Optimize the network topology, dynamically adjust the connection structure between storage nodes to reduce data transmission overhead. By optimizing the network topology, it is possible to reduce the cross-node data interaction overhead caused by the addition or change of entities, and improve the transmission efficiency and topology stability of the system;
[0140] Update the query path and cache policy to ensure the efficiency of query results. According to the sharding allocation of newly added or changed entities, recalculate the query path priority between it and other entities:
[0141] ;
[0142] where t is the target entity, represents the physical path distance. By dynamically updating the query path and cache policy, it is possible to significantly reduce the cross-node hops and repeated calculations of queries, and improve the query efficiency and system response speed;
[0143] Monitor the system performance after the update. If the performance drops, trigger global optimization to restore the stability and efficiency of the system. Through the global optimization mechanism, it is possible to quickly restore stability when the system performance drops, and ensure the efficient operation of the system after the dynamic update of the knowledge graph.
[0144] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for generating a distributed data storage network based on a knowledge graph, characterized in that It includes the following steps: S1. Extract entities and their semantic relationships from the knowledge graph to construct a semantic similarity matrix; S2. Based on the semantic similarity matrix, construct a weighted graph and use a graph partitioning algorithm to generate distributed data shards; S3. Allocate the generated data shards to storage nodes; S4. Generate an initial network topology according to the semantic relevance between data shards; S5. Dynamically monitor the storage system status and optimize the network topology; S5 includes: Dynamically monitor the load, data migration volume, and network latency of storage nodes; Define the objective function of the topology optimization problem based on the current network status; Adjust the network topology through reinforcement learning methods, including adding edges, deleting edges, or adjusting edge weights; Evaluate the cost of data migration that may be caused after topology adjustment, and select the adjustment strategy with the minimum cost; S6. Optimize the query path based on the semantic path of the knowledge graph; S6 includes: Calculate the priority of the query path using the semantic information of the knowledge graph; Plan the query path based on the priority, and use a heuristic search algorithm to generate the optimal path; Construct a distributed cache during the query process to store high-frequency query results to reduce duplicate calculations; S7. Globally optimize data distribution, network topology, and query efficiency through a collaborative optimization framework; S7 includes: Construct a multi-objective optimization model, and set the optimization objectives as storage load balancing, network latency minimization, and query efficiency maximization; Use an evolutionary algorithm to solve the objective function and generate suggestions for data distribution adjustment and network topology optimization; Perform real-time adjustment of storage nodes and network connections based on the optimization results.
2. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, wherein The above S1 includes: Use a graph embedding algorithm to map entities and relationships in the knowledge graph to a low-dimensional vector space; Calculate semantic similarity through the Euclidean distance between node vectors; Construct a semantic similarity matrix based on the semantic similarity.
3. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, wherein The above S2 includes: Construct a weighted graph, where nodes represent entities in the knowledge graph, and the weights of edges are determined by the semantic similarity between entities; Use the Louvain algorithm to perform modular partitioning on the weighted graph to generate several data shards; The division of data shards meets the requirements of maximizing semantic relevance and load balancing of storage nodes.
4. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, wherein The above S3 includes: According to the shard semantic clustering results, allocate data shards with similar semantics to the same storage node; Based on the capacity and load status of storage nodes, adjust the allocation priority of data shards; Ensure global load balancing during the data shard allocation process.
5. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, wherein The above S4 includes: Calculate the data sharing degree between storage nodes, and construct the edge weights of the network according to the sharing degree; Use the weighted minimum spanning tree algorithm to generate an initial network topology; Ensure that the generated network topology can minimize the cross-node data transmission overhead.
6. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, wherein The optimization objectives of the multi-objective optimization model include: The load balancing degree of storage nodes, measured by the maximum value of the node load ratio; Network latency, calculated by the transmission latency and data transmission volume between storage nodes; Query efficiency, calculated by the average response time of the query path.
7. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, wherein This method supports the dynamic update of the knowledge graph and includes the following steps: Dynamically calculate the semantic similarity of newly added or changed entities and update the semantic similarity matrix; Adjust the affected shards, allocate newly added or changed entities to the most relevant shards to ensure semantic relevance and load balancing of the shards; Optimize the network topology and dynamically adjust the connection structure between storage nodes to reduce data transmission overhead; Update the query path and cache policy to ensure the efficiency of query results; Monitor the system performance after the update. If the performance degrades, trigger global optimization to restore the stability and efficiency of the system.
Citation Information
Patent Citations
Rapid construction and storage system for electric power threat intelligence knowledge graph
CN117453922A
Optimization method and device based on intelligent distributed network topology and electronic equipment
CN119603162A