Distributed data storage network generation method based on knowledge graph

By building a semantic similarity matrix and dynamic sharding adjustment mechanism, the network topology and query paths are optimized, and the problem of data sharding static, network topology static and query path optimization in the existing technology is solved, efficient data management and storage resource utilization are achieved, and the system's communication and query efficiency is improved.

CN119938942AActive Publication Date: 2025-05-06BEIJING HANXINSHENG TECH CO LTD

Patent Information

Application Number
CN202510423344.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-06
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

The existing distributed data storage network generation method based on knowledge graphs has problems such as lack of dynamic adjustment capabilities for data sharding, high transmission overhead resulting in static network topology optimization, and insufficient use of semantic information by query path optimization and caching mechanisms.

Method used

By extracting entities and their semantic relationships from the knowledge graph, constructing semantic similarity matrix, weighted graph construction and graph division algorithms generate distributed data shards, and dynamically monitor the state of the storage system to optimize network topology. At the same time, query paths are optimized based on knowledge graph semantic paths, and data distribution, network topology and query efficiency are globally optimized through a collaborative optimization framework.

Benefits of technology

The semantic consistency of data distribution and load balancing of storage nodes are realized, and the update of knowledge graphs is dynamically adapted to the data management efficiency and storage resource utilization of distributed storage systems are significantly improved, and the overhead of cross-node data transmission is reduced, and the communication efficiency and query efficiency of the system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938942A_ABST
    Figure CN119938942A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of distributed storage, and discloses a distributed data storage network generation method based on a knowledge graph, comprising the following steps: S1, extracting entities and semantic relationships thereof from the knowledge graph, and constructing a semantic similarity matrix; s2, on the basis of the semantic similarity matrix, weighted graph construction is carried out, and a graph division algorithm is used to generate distributed data fragments; s3, distributing the generated data fragments to storage nodes; s4, generating an initial network topology according to the semantic correlation among the data fragments; s5, dynamically monitoring the state of the storage system and optimizing the network topology; s6, optimizing a query path based on a knowledge graph semantic path; and S7, performing global optimization on data distribution, network topology and query efficiency through a collaborative optimization framework. Through dynamic fragmentation adjustment, network topology optimization and semantic-driven query path planning, the storage resource utilization rate, the data transmission efficiency and the query performance of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of distributed storage technology, and specifically to a method for generating a distributed data storage network based on a knowledge graph. Background Art

[0002] As the infrastructure of modern large-scale data processing systems, distributed data storage networks can support efficient distributed data management and access. In application scenarios based on knowledge graphs, knowledge graphs form data networks with strong semantic associations by modeling entities and their relationships, providing a semantic basis for the optimization of distributed storage networks. The distributed data storage network generation method based on knowledge graphs uses the semantic structure of graphs to optimize data sharding, network topology, and query paths, aiming to achieve efficient use of storage resources and rapid response to data access.

[0003] Existing methods for generating distributed data storage networks based on knowledge graphs can improve the efficiency of distributed data storage to a certain extent by using fixed graph partitioning strategies and static network topologies. These methods rely on the global structural characteristics of knowledge graphs to shard and distribute data, and achieve low cross-node data transmission overhead in a distributed environment.

[0004] However, the existing technology has some shortcomings. Its data sharding lacks dynamic adjustment capabilities and is difficult to adapt to the frequent updates of knowledge graphs; network topology optimization is relatively static and it is difficult to effectively reduce transmission overhead in a complex dynamic environment; at the same time, query path optimization and caching mechanisms do not fully utilize semantic information, resulting in limited improvement in query efficiency. Summary of the invention

[0005] In view of the shortcomings of the prior art, the present invention provides a method for generating a distributed data storage network based on a knowledge graph, which solves the problems in the prior art that data sharding lacks dynamic adjustment capabilities, the static nature of network topology optimization leads to high transmission overhead, and the query path optimization and caching mechanism do not fully utilize semantic information.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A method for generating a distributed data storage network based on a knowledge graph, comprising the following steps: S1. Extract entities and their semantic relations from the knowledge graph and construct a semantic similarity matrix; S2. Based on the semantic similarity matrix, a weighted graph is constructed and a graph partitioning algorithm is used to generate distributed data shards. S3, distribute the generated data shards to storage nodes; S4, generating an initial network topology based on the semantic correlation between data shards; S5, dynamically monitor storage system status and optimize network topology; S6. Optimize query path based on knowledge graph semantic path; S7. Globally optimize data distribution, network topology and query efficiency through a collaborative optimization framework.

[0007] Preferably, the S1 comprises: Use graph embedding algorithms to map entities and relationships in knowledge graphs to low-dimensional vector spaces; The semantic similarity is calculated by the Euclidean distance between node vectors; Construct a semantic similarity matrix based on semantic similarity.

[0008] Preferably, S2 includes: Construct a weighted graph, where nodes represent entities in the knowledge graph and edge weights are determined by the semantic similarity between entities; Use Louvain algorithm to modularize the weighted graph and generate several data slices; The division of data shards meets the requirements of maximizing semantic relevance and balancing the load of storage nodes.

[0009] Preferably, S3 includes: According to the shard semantic clustering results, data shards with similar semantics are allocated to the same storage node; Adjust the allocation priority of data shards based on the capacity and load status of storage nodes; Ensure global load balancing during data shard allocation.

[0010] Preferably, S4 includes: Calculate the data sharing degree between storage nodes and construct the edge weight of the network based on the sharing degree; The weighted minimum spanning tree algorithm is used to generate the initial network topology; Ensure that the generated network topology minimizes the cross-node data transmission overhead.

[0011] Preferably, S5 includes: Dynamically monitor storage node load, data migration volume, and network latency; Define the objective function of the topology optimization problem based on the current network state; Adjust the network topology through reinforcement learning methods, including adding edges, deleting edges, or adjusting edge weights; Evaluate the cost of data migration that may be caused by topology adjustment and select the adjustment strategy with the lowest cost.

[0012] Preferably, S6 includes: Use the semantic information of the knowledge graph to calculate the priority of the query path; Plan the query path based on priority and use heuristic search algorithm to generate the optimal path; Build a distributed cache during the query process to store high-frequency query results to reduce repeated calculations.

[0013] Preferably, the S7 includes: Build a multi-objective optimization model and set the optimization objectives to balance storage load, minimize network latency, and maximize query efficiency; Use evolutionary algorithms to solve the objective function and generate data distribution adjustment and network topology optimization suggestions; Make real-time adjustments to storage nodes and network connections based on optimization results.

[0014] Preferably, the optimization objectives of the multi-objective optimization model include: The load balance of storage nodes is measured by the maximum node load ratio; Network latency, calculated by the transmission delay between storage nodes and the amount of data transmitted; Query efficiency, calculated by the average response time of the query path.

[0015] Preferably, the method supports dynamic updating of the knowledge graph and includes the following steps: Dynamically calculate the semantic similarity of newly added or changed entities and update the semantic similarity matrix; Adjust the affected shards and assign the new or changed entities to the most relevant shards to ensure the semantic relevance and load balancing of the shards; Optimize network topology and dynamically adjust the connection structure between storage nodes to reduce data transmission overhead; Update query paths and cache strategies to ensure efficient query results; Monitor system performance after the update. If performance degrades, trigger global optimization to restore system stability and efficiency.

[0016] The present invention provides a method for generating a distributed data storage network based on a knowledge graph. It has the following beneficial effects: 1. The present invention achieves semantic consistency of data distribution and load balancing of storage nodes by introducing a semantic similarity matrix and a dynamic sharding adjustment mechanism, while dynamically adapting to the update of the knowledge graph, significantly improving the data management efficiency and storage resource utilization of the distributed storage system.

[0017] 2. By optimizing the network topology and dynamically adjusting the connection structure between storage nodes, the present invention reduces the data transmission overhead across nodes, improves the communication efficiency of the system, and ensures the efficiency and stability of the network topology in a dynamic environment.

[0018] 3. By utilizing heuristic query path planning and distributed caching strategies, the present invention reduces query latency and repeated calculation frequency, realizes efficient query in complex distributed storage networks, and enhances the response speed and overall performance of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0020] The following will be combined with the drawings in the specification of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0021] Please see attached Figure 1 , an embodiment of the present invention provides a method for generating a distributed data storage network based on a knowledge graph, comprising the following steps: S1. Extract entities and their semantic relationships from the knowledge graph and construct a semantic similarity matrix. The constructed semantic similarity matrix can quantify the semantic relevance between entities, provide an accurate semantic information basis for subsequent data sharding, ensure that entities with strong semantic relevance can be aggregated together, and improve storage and query efficiency; S2. Based on the semantic similarity matrix, a weighted graph is constructed and a graph partitioning algorithm is used to generate distributed data shards. By partitioning data based on semantic similarity, entities with high semantic relevance can be aggregated into the same shard, which helps reduce cross-shard queries and improve the semantic consistency of data storage. S3. Allocate the generated data shards to storage nodes. Reasonable allocation of shards can fully utilize storage node resources, avoid storage load imbalance, and ensure fast access to shard data, thereby improving overall storage system performance. S4. Generate an initial network topology based on the semantic correlation between data shards. The generated initial network topology can reduce cross-node data transmission, improve storage and query response efficiency, and lay the foundation for dynamic optimization. S5. Dynamically monitor the storage system status and optimize the network topology. Dynamic optimization can adapt to changes in the storage system and ensure that the network topology can maintain efficient operation when the load fluctuates or the access mode changes. S6. Optimize query paths based on knowledge graph semantic paths. Optimizing semantic paths can reduce query latency and cross-node transmission overhead, thus improving query efficiency. S7. Global optimization of data distribution, network topology and query efficiency is performed through a collaborative optimization framework. The collaborative optimization framework can globally improve the storage and query performance of the system in a dynamic environment, ensuring the robustness and efficiency of the distributed storage network.

[0022] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, S1 includes: Use graph embedding algorithms to map entities and relationships in the knowledge graph to low-dimensional vector space. Embedding algorithm selection: Use embedding algorithms (such as TransE, TransR, Node2Vec) to map entities and relationships to vector space: TransE Assumptions , optimize the following loss function: ; in, is the positive sample set, is the set of negative samples, It is a hyperparameter used to control the interval between positive and negative samples. The embedding algorithm can reduce the dimension of complex high-dimensional semantic relationships into numerical vectors, providing an efficient and intuitive basis for subsequent calculations while retaining the semantic characteristics and topological information in the knowledge graph. The semantic similarity is calculated by the Euclidean distance between node vectors. The semantic similarity calculation formula is: Given two entity nodes , Their semantic similarity is calculated by the Euclidean distance of the embedding vectors: ; in, , Node , The embedding vector of It is a hyperparameter that controls the distance. It quantifies the semantic relevance between entities by calculating semantic similarity, providing an accurate semantic basis for subsequent sharding optimization and network topology generation, ensuring that the system can better reflect the intrinsic relationship of the knowledge graph. Construct a semantic similarity matrix based on semantic similarity. The matrix is ​​constructed as follows: The calculated semantic similarity values ​​are stored in the form of a symmetric matrix, and the elements of the matrix S are defined as: ; in, Representation Node and The semantic similarity matrix stores the semantic relationship between entities in the knowledge graph in a structured manner, providing direct input for subsequent weighted graph construction and data sharding. At the same time, it optimizes storage and computing resources through sparse processing, thereby improving the efficiency and scalability of the system.

[0023] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, S2 includes: Construct a weighted graph, where nodes represent entities in the knowledge graph and edge weights are determined by the semantic similarity between entities. The graph is constructed in the following ways: Each entity node in the knowledge graph are mapped as nodes in a weighted graph; The value in the semantic similarity matrix S is used as the weight of the edge , the edge is defined as: ; in, The similarity threshold is used, and edges smaller than the threshold are not constructed, thereby reducing unnecessary edge connections. By constructing a weighted graph, the semantic information between entities in the knowledge graph can be intuitively presented in the form of a graph structure. At the same time, the redundant edges of the graph are reduced through sparse processing, the processing efficiency of the graph is improved, and a good structural foundation is provided for subsequent data sharding. The Louvain algorithm is used to modularize the weighted graph and generate several data fragments. The goal of the Louvain algorithm is to find the community structure in the weighted graph by maximizing the modularity Q of the graph. The modularity is defined as: ; in, is the edge weight; and are the degrees of nodes i and j respectively; m is the total edge weight; Indicates whether nodes i and j belong to the same shard; Algorithm execution steps: Phase 1: Initially, each node is considered as an independent community; Traverse the node set, move nodes to adjacent communities, and calculate modularity gain : ; in, represents the edge weight within the community, Represents the total weight of the community.

[0024] Phase 2: Based on the local optimization results, nodes in the same community are merged into super nodes, the graph structure is updated and phase 1 is repeated until the modularity converges; The Louvain algorithm can aggregate entity nodes with high semantic relevance into the same data shard through modular segmentation, ensuring semantic consistency within the shard and laying the foundation for efficient query of distributed storage networks. The division of data shards meets the requirements of maximizing semantic relevance and balancing the load of storage nodes. Through the maximization of shard semantic relevance and load balancing strategy, it can not only ensure the semantic consistency within the data shard, but also make full use of storage node resources, avoid data skew problems in distributed systems, and improve the storage and access efficiency of the overall system.

[0025] Please see attached Figure 1 In a preferred embodiment of the present invention, S3 includes: According to the shard semantic clustering results, data shards with similar semantics are allocated to the same storage node. Allocating shards according to the semantic clustering results can significantly reduce the transmission of data with high semantic relevance across storage nodes, improve query efficiency, and reduce network communication overhead; Adjust the allocation priority of data shards based on the capacity and load status of storage nodes. Dynamically adjust the allocation priority based on node capacity and load status to achieve efficient utilization of storage node resources, avoid excessive load on a single node, and ensure system stability and reliability. Ensure global load balancing during data shard allocation. Through the global load balancing strategy, it can effectively avoid resource waste or overload problems of storage nodes in the system, improve the overall stability and performance of the system, and reduce data access delays caused by load imbalance.

[0026] Please see attached Figure 1 In a preferred embodiment of the present invention, S4 includes: Calculate the data sharing degree between storage nodes and construct the edge weight of the network based on the sharing degree. Through data sharing degree calculation and weighted graph construction, the data interaction intensity between storage nodes can be quantified, providing accurate edge weight information for subsequent network topology generation, ensuring the rationality and efficiency of the topology. The weighted minimum spanning tree algorithm is used to generate the initial network topology. The initial network topology generated by the weighted minimum spanning tree algorithm can connect all storage nodes with the lowest cross-node transmission cost, provide an efficient communication path for system data interaction, and reduce network complexity. Ensure that the generated network topology can minimize the cross-node data transmission overhead. Through optimization and dynamic adjustment, it can ensure that the network topology always minimizes the cross-node transmission overhead, improve the system's data transmission efficiency, and reduce the response time of data query and storage operations.

[0027] Please see attached Figure 1 In a preferred embodiment of the present invention, S5 includes: Dynamically monitor the load, data migration volume and network latency of storage nodes, and provide real-time status information for network topology optimization, so that topology adjustment can accurately reflect the current system performance bottleneck; The objective function of the topology optimization problem is defined based on the current network status. By comprehensively considering the objective function of network delay, load balancing and migration cost, the topology optimization can be guided to adjust to the direction of global performance optimization, thereby improving the system operation efficiency and stability. Adjust the network topology through reinforcement learning methods, including adding edges, deleting edges, or adjusting edge weights. Reinforcement learning can dynamically adjust the network topology and achieve adaptive optimization under different states, effectively reducing manual adjustment costs and improving topology adjustment efficiency. The cost of data migration that may be caused by topology adjustment is evaluated, and the adjustment strategy with the lowest cost is selected. The migration cost evaluation and incremental migration mechanism can significantly reduce the performance overhead caused by topology adjustment, avoid resource waste caused by frequent migration, and ensure the continuity of the system.

[0028] Please see attached Figure 1 In a preferred embodiment of the present invention, S6 includes: By using the semantic information of the knowledge graph to calculate the priority of the query path, the query scope can be significantly reduced, high-priority paths can be placed in the front, query efficiency can be improved, and system burden can be reduced; The query path is planned based on priority, and the optimal path is generated by using a heuristic search algorithm. The heuristic search algorithm combined with semantic priority can efficiently generate the optimal query path in a complex distributed storage network, reduce the number of hops and costs of cross-node queries, and thus improve query performance; Build a distributed cache during the query process to store high-frequency query results to reduce repeated calculations. The introduction of distributed cache can reduce repeated query calculations and cross-node data access frequency, reduce network load, and improve system response speed and overall performance.

[0029] Please see attached Figure 1 In a preferred embodiment of the present invention, S7 includes: Build a multi-objective optimization model and set the optimization goals to balance storage load, minimize network latency, and maximize query efficiency. The construction of the multi-objective optimization model can focus on the system's storage, network, and query performance at the same time, provide global performance optimization goals, and provide comprehensive guidance for system decision-making; Use evolutionary algorithms to solve the objective function and generate data distribution adjustment and network topology optimization suggestions. The evolutionary algorithm uses NSGA-II (non-dominated sorting genetic algorithm) to solve multi-objective optimization problems. The evolutionary algorithm can efficiently solve complex multi-objective optimization problems, generate optimization solutions that balance different objectives, and ensure the global optimality of system performance in a dynamic environment; Based on the optimization results, storage nodes and network connections are adjusted in real time. By adjusting storage nodes and network connections in real time, it is possible to dynamically adapt to changes in system requirements, quickly respond to performance bottlenecks, and improve the robustness and adaptability of the system.

[0030] Please refer to the attached Figure 1 In a preferred embodiment of the present invention, the optimization objectives of the multi-objective optimization model include: The load balance of storage nodes is measured by the maximum load ratio of the nodes. In order to avoid resource skewness of storage nodes, the load balancing goal is defined as minimizing the load difference between nodes. Measuring balance: ; in, For Node The amount of data currently stored; For Node Total storage capacity; Regularly calculate the system's load balance B: ; like (preset threshold), triggering the load balancing optimization mechanism, and restoring the balance to the ideal range by adjusting the shard distribution. By optimizing the load balancing target, the resource tilt problem of storage nodes can be avoided, ensuring that all nodes use resources evenly, thereby improving the overall stability and scalability of the system; Network latency is calculated by storing the transmission latency and data transmission volume between nodes. To reduce the latency caused by cross-node data access, the network latency target is defined as the sum of the transmission latency between nodes: ; in, Representation Node and Transmission delay between Represents the amount of data transmitted on edge (i, j). By optimizing the network latency target, it can reduce the average time for data access across nodes, improve the data interaction efficiency of the distributed storage system and the response speed of user queries; Query efficiency is calculated by the average response time of the query path. The query efficiency goal is defined as minimizing the average response time of the query path. ; ; in, represents the response time of query q, including path search time and data transmission time; Q is the query set; The semantic information of the knowledge graph is used to optimize the query path, giving priority to paths with high semantic relevance and low transmission cost: ; in, Indicates the physical distance of the path; In path planning, cache hit paths are given priority to reduce repeated calculations and cross-node accesses; By optimizing the query efficiency target, the response time of the query path can be reduced, the user experience can be improved, and the system load and communication costs can be reduced.

[0031] Please see attached Figure 1 In a preferred embodiment of the present invention, the method supports dynamic updating of the knowledge graph and includes the following steps: Dynamically calculate the semantic similarity of newly added or changed entities and update the semantic similarity matrix. By dynamically updating the semantic similarity matrix, it can ensure that the semantic relationship between the newly added or changed entities and other entities in the knowledge graph is accurately reflected in the storage system, providing a basis for subsequent sharding adjustments and query optimization. Adjust the affected shards and assign the newly added or changed entities to the most relevant shards to ensure the semantic relevance and load balance of the shards. Dynamically adjusting shard allocation can ensure that the newly added or changed entities are assigned to the most suitable shards, which not only maintains the semantic consistency within the shards, but also avoids unbalanced node loads, thus improving the robustness and query efficiency of the system. Optimize network topology and dynamically adjust the connection structure between storage nodes to reduce data transmission overhead. By optimizing the network topology, the cross-node data interaction overhead caused by adding or changing entities can be reduced, and the transmission efficiency and topology stability of the system can be improved. Update the query path and cache strategy to ensure the efficiency of query results. Recalculate the query path priority with other entities based on the shard allocation of the newly added or changed entity: ; Among them, t is the target entity, Indicates the physical path distance. By dynamically updating the query path and cache strategy, it can significantly reduce the number of cross-node hops and repeated calculations of the query, thereby improving query efficiency and system response speed. Monitor system performance after the update. If the performance degrades, trigger global optimization to restore system stability and efficiency. Through the global optimization mechanism, stability can be quickly restored when system performance degrades, ensuring efficient operation of the system after the knowledge graph is dynamically updated.

[0032] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for generating a distributed data storage network based on a knowledge graph, characterized in that: The following steps are involved: S1. Extract entities and their semantic relations from the knowledge graph and construct a semantic similarity matrix; S2. Based on the semantic similarity matrix, a weighted graph is constructed and a graph partitioning algorithm is used to generate distributed data shards. S3, distribute the generated data shards to storage nodes; S4, generating an initial network topology based on the semantic correlation between data shards; S5, dynamically monitor storage system status and optimize network topology; S6. Optimize query path based on knowledge graph semantic path; S7. Globally optimize data distribution, network topology and query efficiency through a collaborative optimization framework.

2. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The S1 includes: Use graph embedding algorithms to map entities and relationships in knowledge graphs to low-dimensional vector spaces; The semantic similarity is calculated by the Euclidean distance between node vectors; Construct a semantic similarity matrix based on semantic similarity.

3. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The S2 includes: Construct a weighted graph, where nodes represent entities in the knowledge graph and edge weights are determined by the semantic similarity between entities; Use Louvain algorithm to modularize the weighted graph and generate several data slices; The division of data shards meets the requirements of maximizing semantic relevance and balancing the load of storage nodes.

4. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The S3 includes: According to the shard semantic clustering results, data shards with similar semantics are allocated to the same storage node; Adjust the allocation priority of data shards based on the capacity and load status of storage nodes; Ensure global load balancing during data shard allocation.

5. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The S4 includes: Calculate the data sharing degree between storage nodes and construct the edge weight of the network based on the sharing degree; The weighted minimum spanning tree algorithm is used to generate the initial network topology; Ensure that the generated network topology minimizes the cross-node data transmission overhead.

6. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The S5 includes: Dynamically monitor storage node load, data migration volume, and network latency; Define the objective function of the topology optimization problem based on the current network state; Adjust the network topology through reinforcement learning methods, including adding edges, deleting edges, or adjusting edge weights; Evaluate the cost of data migration that may be caused by topology adjustment and select the adjustment strategy with the lowest cost.

7. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The S6 includes: Use the semantic information of the knowledge graph to calculate the priority of the query path; Plan the query path based on priority and use heuristic search algorithm to generate the optimal path; Build a distributed cache during the query process to store high-frequency query results to reduce repeated calculations.

8. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The S7 includes: Build a multi-objective optimization model and set the optimization objectives to balance storage load, minimize network latency, and maximize query efficiency; Use evolutionary algorithms to solve the objective function and generate data distribution adjustment and network topology optimization suggestions; Make real-time adjustments to storage nodes and network connections based on optimization results.

9. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The optimization objectives of the multi-objective optimization model include: The load balance of storage nodes is measured by the maximum node load ratio; Network latency, calculated by the transmission delay between storage nodes and the amount of data transmitted; Query efficiency, calculated by the average response time of the query path.

10. The method for generating a distributed data storage network based on a knowledge graph according to claim 1, characterized in that: The method supports dynamic updating of knowledge graph and includes the following steps: Dynamically calculate the semantic similarity of newly added or changed entities and update the semantic similarity matrix; Adjust the affected shards and assign the new or changed entities to the most relevant shards to ensure the semantic relevance and load balancing of the shards; Optimize network topology and dynamically adjust the connection structure between storage nodes to reduce data transmission overhead; Update query paths and cache strategies to ensure efficient query results; Monitor system performance after the update. If performance degrades, trigger global optimization to restore system stability and efficiency.

Citation Information

Patent Citations

  • Construction method for multidimensional data-oriented semantic indexing peer-to-peer network

    CN101853283A

  • Semantic data storage and retrieval method and device based on maximum area grid

    CN112148830A

  • Efficient query method for large-scale graph data

    CN116383247A

  • Rapid construction and storage system for electric power threat intelligence knowledge graph

    CN117453922A

  • Data space construction method and system based on knowledge graph

    CN118193491A

Cited By

  • Data caching method, device and equipment, network equipment and storage medium

    CN120499203A

  • Database multi-table query optimization method and device

    CN120723806A

  • Graph database Leader fragment distribution method in multi-Zone scene

    CN120785742A

  • Distributed data storage and access processing method

    CN120785908A

  • A distributed data storage and access processing method

    CN120785908B