Distributed data identifier analysis method and device based on bidirectional routing
By constructing a distributed data identification resolution method with a bidirectional routing table, the problem of low efficiency of data identification resolution in large-scale heterogeneous networks is solved, fast and reliable data positioning and query are achieved, and the robustness and scalability of the system are improved.
Patent Information
- Application Number
- CN202510966144.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies have difficulty in achieving rapid and accurate positioning of data resources in large-scale heterogeneous networks, resulting in low efficiency and insufficient reliability in data identification and resolution. In particular, high latency and resolution failures are prone to occur in complex paths or dynamic topologies, affecting the robustness and scalability of the system.
A distributed data identification resolution method based on bidirectional routing is adopted. The first data segment in the data identification locator uniquely corresponds to the target data, and the second data segment uniquely corresponds to the target node. The forward and reverse routing tables are constructed, and the fast resolution and fault tolerance of data are achieved through the matching and forwarding mechanism.
It improves the efficiency and reliability of data identification queries, can quickly locate target data in complex networks, and switch addressing mechanisms when routing table matching fails, ensuring the continuity and robustness of queries, reducing resource waste, and improving the system's fault tolerance.
Smart Images

Figure CN120821753A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of distributed storage, and specifically relates to a distributed data identification resolution method, device, equipment and storage medium based on bidirectional routing. Background Art
[0002] With the widespread adoption of distributed storage and data sharing technologies, how to quickly and accurately locate data resources in large-scale heterogeneous networks has become a research hotspot. In particular, in scenarios such as the Internet of Things, edge computing, and trusted data sharing, the global uniqueness and rapid resolution of data identifiers directly impact the system's collaborative efficiency and service quality.
[0003] Existing identity resolution technologies are primarily implemented using centralized or hierarchical architectures. For example, naming schemes based on prefix and suffix structures implement hierarchical identity resolution through global registries and local service nodes, or domain name resolution services convert identifiers to specific server addresses.
[0004] However, existing technologies generally have difficulty maintaining efficient addressing in complex paths or dynamic topologies, resulting in limited efficiency in data identification resolution. At the same time, when target node information is missing or the number of path hops increases, the system is prone to high latency or even the risk of resolution failure, affecting overall robustness, and has limited scalability and fault tolerance. Summary of the Invention
[0005] The present application aims to provide a distributed data identification resolution method, apparatus, device and storage medium based on bidirectional routing, which at least solves the problem of insufficient efficiency and reliability of data identification query.
[0006] In a first aspect, an embodiment of the present application discloses a distributed data identification resolution method based on bidirectional routing, which is applied to a first node in a distributed storage network, comprising: In response to an identification query instruction for first target data, matching a first data identifier locator corresponding to the first target data in the identification query instruction with a first routing table of the first node; a first data segment in the first data identifier locator uniquely corresponds to the first target data, and a second data segment in the first data identifier locator uniquely corresponds to a first target node storing the first target data; If the first data identifier locator successfully matches the first routing table, forwarding the identifier query instruction to the second node according to the first routing table; If the first data identifier locator fails to match the first routing table, matching the first data identifier locator with the second routing table of the first node; If the first data identifier locator successfully matches the second routing table, forwarding the identifier query instruction to a third node according to the second routing table; In a case where the first data identifier locator fails to be matched with the first routing table and the second routing table, the first node is determined as the first target node.
[0007] In a second aspect, an embodiment of the present application further discloses a distributed data identification resolution device based on bidirectional routing, which is applied to a first node in a distributed storage network, including: a first matching module, configured to, in response to an identification query instruction for first target data, match a first data identifier locator corresponding to the first target data in the identification query instruction with a first routing table of the first node; wherein a first data segment in the first data identifier locator uniquely corresponds to the first target data, and a second data segment in the first data identifier locator uniquely corresponds to a first target node storing the first target data; a first forwarding module, configured to forward the identification query instruction to a second node according to the first routing table when the first data identifier locator successfully matches the first routing table; a second matching module, configured to match the first data identifier locator with a second routing table of the first node if the first data identifier locator fails to match the first routing table; a second forwarding module, configured to forward the identification query instruction to a third node according to the second routing table when the first data identifier locator successfully matches the second routing table; The query confirmation module is configured to determine the first node as the first target node if the first data identifier locator fails to match the first routing table and the second routing table.
[0008] In a third aspect, an embodiment of the present application further discloses an electronic device comprising a processor and a memory, wherein the memory stores programs or instructions that can be run on the processor, and when the programs or instructions are executed by the processor, the steps of the method described in the first aspect are implemented.
[0009] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] In summary, in the embodiment of the present application, the unique correspondence between the first data segment in the data identifier locator and the target data is firstly utilized so that the node can accurately identify the identification instruction of the target data, thereby ensuring that the parsing instruction has a clear addressing target and providing a semantically clear starting point for subsequent routing selection; at the same time, the second data segment is uniquely corresponding to the target node, further enhancing the adaptability of the locator to the physical network topology. This structural encoding lays the foundation for subsequent efficient matching and helps to improve the accuracy of the overall parsing instruction; further, in the routing matching process, if the first data identifier locator is successfully matched in the first routing table, it can be directly passed. The system quickly reaches the next hop node in the direction of the target data through the reverse path to achieve rapid data resolution; when the first routing table fails to match, the system automatically switches to the second routing table and starts the forward path addressing mechanism, so that the system still has the ability to resolve the target in abnormal situations such as the node distance is far, the node is missing, the network is interrupted, or the number of path hops increases abnormally, effectively avoiding the query failure problem caused by the singleness of the path; further, when both the forward and reverse tables are not hit, it can be determined that the current node is the target node where the data is stored, and the query process is terminated, while reducing the waste of network resources. It provides a small range of data storage nodes with resolution autonomy. Therefore, based on the method of the embodiment of the present application, by constructing a dual-table matching mechanism including a forward routing table and a reverse routing table, the redundant design and fault tolerance improvement of the resolution path are achieved, the performance bottleneck caused by the single routing structure is effectively alleviated, and the query efficiency and reliability of the data identification are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the attached figure: Figure 1 This is a flowchart of a distributed data identification resolution method based on bidirectional routing provided by an embodiment of the present application; Figure 2 This is a schematic diagram of an identification query process in an embodiment of the present application; Figure 3 This is a flowchart of another distributed data identification resolution method based on bidirectional routing provided by an embodiment of the present application; Figure 4 This is a schematic diagram of an identification indexing process in an embodiment of the present application; Figure 5 This is a flowchart of another method for distributing data identification resolution based on bidirectional routing provided by an embodiment of the present application; Figure 6 It is a program execution process of a cache decision process under an embodiment of the present application; Figure 7 It is the program execution process of identifying the query under the cache decision process of the embodiment of the present application; Figure 8This is a block diagram of a distributed data identification resolution device based on bidirectional routing provided by an embodiment of the present application; Figure 9 This is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0013] The terms "first," "second," and the like in the application documents of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the application documents of this application represents at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0014] Based on the above process, Figure 1 As shown, a distributed data identification resolution method based on bidirectional routing is provided in an embodiment of the present application, which is applied to the first node in a distributed storage network.
[0015] The method may include the following steps: Step 101: In response to an identification query instruction for first target data, a first data identifier locator corresponding to the first target data in the identification query instruction is matched with a first routing table of a first node.
[0016] The first data segment in the first data identifier locator uniquely corresponds to the first target data, and the second data segment in the first data identifier locator uniquely corresponds to the first target node storing the first target data.
[0017] In some embodiments of the present application, in order to start the identification and positioning process for a specific data resource, the first node responds to the identification query instruction for the first target data, and matches the first data identification locator corresponding to the first target data in the identification query instruction with the first routing table of the first node. This process ensures that the query information is accurately mapped to the expected path by performing structured matching on the first data segment and the second data segment contained in the locator. The first data segment is used to uniquely identify the data resource itself, and the second data segment is used to identify the location of the target data storage node. The two together constitute the basis for accurate addressing. After the matching is completed, the node can quickly determine whether the target data can be obtained through reverse routing based on the matching results, thereby providing a criterion for the selection of subsequent addressing paths, which helps to achieve rapid initialization of data identification resolution instructions and improved addressing accuracy in distributed networks.
[0018] In a specific example, when a user wants to access a data resource identified by "pku.cs / yyds801", the system generates a structured locator (Digital Object Locator, DOL) for this identifier through a segmented hash mapping rule. Its format is "xxxxx-yyy-00-zzzzzz" and has the following structure: "xxxxx" is the space coding segment, which occupies the first 32 bits and is generated by the data space name (such as "pku") through a hash algorithm, indicating the logical space area of the access identity resolution network; "yyy" is the node coding segment, which occupies the middle 24 bits and is generated by hashing the node information (such as "cs") and is used to distinguish different data nodes under the subject; "00" is a fixed delimiter used to mark the boundary between the prefix and suffix; "zzzzzz" is the data suffix coding segment, which occupies 64 bits and is generated by hashing the part that identifies the user data resource (such as "yyds801") and is used to identify a specific data entity. After receiving the query, the first node identifies the spatial and node codes in the locator and matches them against its locally maintained forward routing table to determine whether a neighboring node exists along the corresponding path. Once a match is made, the system can then decide whether to forward the request. This, based on a clear query target, triggers the path addressing process, improving data access target aggregation and transmission efficiency.
[0019] Step 102: When the first data identifier locator successfully matches the first routing table, forward the identifier query instruction to the second node according to the first routing table.
[0020] In some embodiments of the present application, in order to complete the effective routing of the parsing instruction after identifying the target data path, the first node will forward the identification query instruction to the second node according to the reverse routing table when the first data identifier locator successfully matches the first routing table (i.e., the reverse routing table). This process determines the node most relevant to the target data and completes the forwarding of the message by searching the Bloom Filter bit vector set containing the target identifier in the reverse routing table. The reverse routing table records the mapping relationship between the data identifier set and the corresponding node, and can provide fast backtracking path support when a specific node does not initiate addressing. When the match is successful and the forwarding is completed, the system can reversely trace back to the data storage node in a complex network topology, which helps to avoid path stagnation and improve the continuity and stability of the query.
[0021] like Figure 2 The diagram below illustrates the transmission relationship between nodes in a distributed storage network. The thick black arrows in the diagram represent the forwarding mechanism for a specific data identifier in the first routing table, and the thick white arrows represent the forwarding mechanism for a specific data identifier in the second routing table. In a specific example, node C receives a query instruction from a user, aiming to locate a data resource represented by "xxxxx-yyy-00-zzzzzz." The system first identifies the prefix field of the locator and confirms that the data resource is associated with a specific data space and node information. Node C then checks its local first routing table (i.e., the reverse routing table) and finds that the Bloom filter set of a routing entry contains the data identifier. Based on this, node C determines that node B owns the target data or is responsible for the transit path. Based on the matching result in the reverse routing table, node C forwards the query instruction from node C to node B. This path allows the query request to be quickly traced back to the potential target node B without relying on a network-wide broadcast. This ensures that the system can smoothly process the query and shorten the response link even in nonlinear physical topologies or complex network structures. This reverse path query provides an important supplementary resolution capability when the system does not have a complete forward addressing path.
[0022] Step 103: When the first data identifier locator fails to match the first routing table, the first data identifier locator is matched with the second routing table of the first node.
[0023] In some embodiments of the present application, in order to continue the data identification parsing process when the reverse path match fails, the first node will match the first data identification locator with the second routing table it maintains (i.e., the forward routing table) to attempt to continue the addressing operation in a direction that is logically closer to the target data. The matching process involved in this step is usually based on logical distance judgment under the Distributed Hash Table (DHT) mechanism, using routing algorithms such as Kademlia to select appropriate neighbor nodes through XOR calculation of the distance between nodes, and achieve rapid convergence of the optimal path within O(logN) complexity. The forward routing table is divided into multiple K buckets according to distance, and each bucket records information about other nodes within a specific range from the current node. When the match is successful, it can provide candidate paths for subsequent instruction forwarding, enhancing the system's addressing capability and path flexibility when reverse routing is unavailable, thereby reducing the network query failure rate and improving overall robustness.
[0024] like Figure 2 In a specific example, as shown in the figure, if node F fails to find the target identifier when attempting to query the identifier, it determines that there is no valid match in the routing table. At this point, node F immediately invokes its secondary routing table (forward routing table) and searches the corresponding K bucket for a data node with a closer logical distance based on the prefix hash value of "xxxxx-yyy-00-zzzzzz". Assuming that there is a candidate node E in the routing table whose identifier is closer to the target data in the hash space, node F selects this candidate node for next-hop forwarding, thus continuing the identifier query process and ensuring addressing continuity and information forwarding efficiency within the network.
[0025] Step 104: When the first data identifier locator successfully matches the second routing table, the identifier query instruction is forwarded to the third node according to the second routing table.
[0026] In some embodiments of the present application, in order to continue the identification query process after a match fails in the first routing table (i.e., the reverse routing table), the first node, upon successfully matching the first data identifier locator with the second routing table (i.e., the forward routing table), forwards the identification query instruction to a third node according to the routing policy of the forward routing table. After a successful match and forwarding, the identification query instruction can be advanced along a path that logically approaches the target node, thereby improving addressing convergence speed and query success rate, and enhancing the system's path redundancy and business continuity in scenarios where reverse path support is lost.
[0027] As shown in Figure 2, in a specific example, assuming that the first node is node E, after node E fails to match the identification query instruction with the local reverse routing table, it starts the forward path matching process. The system uses the hash value corresponding to "xxxxx" in the data identification locator "xxxxx-yyy-00-zzzzzz" to match in the forward routing table, and finds the third node D as the next hop forwarding object in the K bucket with a relatively close logical distance. Node E can encapsulate the request and send it to node D through the DOIP protocol to continue the identification and positioning process of the target data. This jump can supplement the addressing link when the path does not form a closed loop, which helps the system maintain the integrity and efficiency of identification resolution in complex or partially failed network structures.
[0028] Step 105 : When the first data identifier locator fails to match the first routing table and the second routing table, the first node is determined as the first target node.
[0029] In some embodiments of the present application, the resolution mechanism provided by the present application ensures that, if a first data identifier locator fails to match both the first routing table (reverse routing table) and the second routing table (forward routing table), the current first node itself is uniquely identified as the first target node corresponding to the data identifier. This determination mechanism is based on the principle of uniqueness of the locator structure. That is, if any routing path cannot be forwarded, it indicates that the data may be locally stored or its routing information has not yet been widely disseminated. Therefore, the current node has potential data ownership or interpretation rights, thereby triggering the termination of data resolution.
[0030] In a specific example, when a user initiates a query request for the locator "xxxxx-yyy-00-zzzzzz", the instruction is forwarded to a certain node A. Node A first attempts to find a set of recorded data identifiers through the reverse routing table, but does not find a Bloom filter bit vector that matches the identifier. It then continues to search for neighboring nodes with a closer logical distance in the forward routing table, but also fails to find a suitable forwarding target. At this point, the system can finally determine that the data corresponding to the identifier may originate from node A's own data space, or that an effective path has not yet been established to propagate its index information. Node A then uses itself as the response node for the query target, completing the landing process of the identifier query.
[0031] In summary, in the embodiment of the present application, the unique correspondence between the first data segment in the data identifier locator and the target data is firstly utilized so that the node can accurately identify the identification instruction of the target data, thereby ensuring that the parsing instruction has a clear addressing target and providing a semantically clear starting point for subsequent routing selection; at the same time, the second data segment is uniquely corresponding to the target node, further enhancing the adaptability of the locator to the physical network topology. This structural encoding lays the foundation for subsequent efficient matching and helps to improve the accuracy of the overall parsing instruction; further, in the routing matching process, if the first data identifier locator is successfully matched in the first routing table, it can be directly passed. The system quickly reaches the next hop node in the direction of the target data through the reverse path to achieve rapid data resolution; when the first routing table fails to match, the system automatically switches to the second routing table and starts the forward path addressing mechanism, so that the system still has the ability to resolve the target in abnormal situations such as the node distance is far, the node is missing, the network is interrupted, or the number of path hops increases abnormally, effectively avoiding the query failure problem caused by the singleness of the path; further, when both the forward and reverse tables are not hit, it can be determined that the current node is the target node where the data is stored, and the query process is terminated, while reducing the waste of network resources. It provides a small range of data storage nodes with resolution autonomy. Therefore, based on the method of the embodiment of the present application, by constructing a dual-table matching mechanism including a forward routing table and a reverse routing table, the redundant design and fault tolerance improvement of the resolution path are achieved, the performance bottleneck caused by the single routing structure is effectively alleviated, and the query efficiency and reliability of the data identification are improved.
[0032] Figure 3 This is another distributed data identification resolution method based on bidirectional routing provided in an embodiment of the present application, which is applied to the first node in a distributed storage network.
[0033] The method may include the following steps: Step 201: In response to second target data stored in a first node, determine a second target node for storing a second data identifier locator, and generate an identifier index instruction.
[0034] In some embodiments of the present application, in order to enable the newly stored data to be accurately parsed and located subsequently, after receiving the second target data, the first node will generate its corresponding data identifier locator based on the identification structure of the target data, and determine the second target node responsible for storing the locator. The selection of this node is usually based on a distributed hash map (such as SHA256) and a logical distance calculation to ensure that the mapping value of the generated identifier locator in the hash space is close to the identifier of the second target node. Subsequently, the first node will construct an identifier index instruction for the locator and start the index transfer process to provide a basis for the subsequent construction of the reverse routing path. The identifier index instruction is a structured message format for disseminating data identification information. Its core is to announce the storage path of the identifier to other nodes in the network, providing a dynamic basis for the formation of the reverse routing table and the construction of routing items.
[0035] In a specific example, Figure 4 As shown, node A receives the second target data. The system performs a segmented hash mapping on it, generating the locator "xxxxx-yyy-00-zzzzzz." Based on the distribution of the prefix segment of this locator in the hash space, the system determines that the node with the closest logical distance is node D, thus determining D as the second target node. Node A then generates an identification index instruction and forwards it according to the order of neighbor nodes in the second routing table (forward routing table). As shown in the figure, the index instruction is forwarded in a progressive manner, following the path indicated by the thick black arrows: A → B → C → D. During instruction forwarding, each node along the way constructs or updates its local reverse routing table, forming an identification tracking path from node B back to node A. This mechanism not only enables the forward propagation of the identification path but also lays the routing foundation for reverse addressing during the subsequent identification resolution process.
[0036] Optionally, step 201 includes the following sub-steps: Sub-step 2011: Generate a first data segment in a second data identifier locator of the second target data according to a hash value of the data identifier of the second target data.
[0037] In some embodiments of the present application, in order to construct a structured identifier for uniquely representing the second target data, the system calculates a hash value based on the original data identifier of the data and uses the result as the first data segment in the second data identifier locator. This process generates a fixed-length output by executing a cryptographic hash algorithm (such as SHA-256) on the original identifier, and intercepts the portion used to identify the data body as the first data segment. The cryptographic hash algorithm can map input of any length into a fixed-length, collision-resistant code, thereby achieving a desensitizing, unique, and secure identifier conversion effect. By generating this data segment, the system not only improves the discreteness and retrieval accuracy of the identifier space, but also lays the foundation for the complete construction and route matching of subsequent locators.
[0038] In a specific example, a user wishes to register a new data resource named "pku.cs / yyds801." The system hashes the string "yyds801" using the SHA-256 algorithm, obtaining a 256-bit hash value. 64 bits of this hash value are then extracted as the first data segment "zzzzzz" in the second data identifier locator. This "zzzzzz" encoded segment is subsequently used for routing table matching and locator assembly. As the core field identifying the data entity, it ensures uniqueness and stability within the distributed network.
[0039] Sub-step 2012: Match the first data segment in the second data identifier locator with the second routing table of the first node to obtain node information of the second target node.
[0040] In some embodiments of the present application, to select a suitable target node for the second target data and generate a complete locator structure, the system matches the first data segment (i.e., the data suffix encoding segment) in the generated second data identifier locator with the second routing table (i.e., the forward routing table) maintained by the first node to obtain the information of the second target node with the closest logical distance. This matching operation is typically based on the exclusive OR (XOR) logical distance between the hash value of the first data segment and the node identifiers in the routing table, using the Kademlia algorithm to search for the nearest neighbor node from several K buckets. The first data segment is a 64-bit encoding segment generated by hashing the original data identifier using SHA-256, which is used to represent a specific data resource. Its uniqueness and irreversibility ensure the accuracy and security of the resolution path. After the match is completed, the system can obtain the node identifier and network address information of the target node, providing a precise positioning basis for the subsequent generation of the second data segment and initiation of the identifier indexing process.
[0041] In a specific example, the data uploaded by the user is identified as "pku.cs / yyds801." The system hashes it to obtain a 64-bit hash value of "zzzzzz," which serves as the first data segment of the locator. The current node is node A. The system matches "zzzzzz" with the K-bucket structure in the second routing table maintained by node A, and through logical distance judgment, finds node B, which is closest to "zzzzzz" in the hash space, as the second target node. The system then extracts node B's identification and address information, providing a path basis for the subsequent steps of constructing a complete locator and the direction of index instruction transmission. Through this process, the system completes the first-stage binding of identification and storage path, improving the convergence and reachability of subsequent addressing targets.
[0042] Sub-step 2013: Generate a second data segment in a second data identifier locator of the second target data according to the hash value of the node identifier in the node information of the second target node.
[0043] In some embodiments of the present application, in order to complete the structural encoding of the second data identifier locator and establish a unique binding relationship with the target storage node, the system calculates its hash value based on the node identifier in the node information of the second target node, and uses the result as the second data segment in the second data identifier locator. This process generates a stable, irreversible, fixed-length code as one of the structural segments in the identification path by executing a cryptographic hash algorithm (such as SHA-256) on the node identifier (usually a node unique ID or network identifier). The hash value of the node identifier not only maps the node's position in the logical network space, but also has the ability to uniquely calibrate the node's identity, thereby giving the locator's second data segment clear network directivity. By generating this field, the system establishes a bidirectional structural index of "data resource-to-belonging node", providing consistency constraints for the subsequent parsing and verification of the addressing path.
[0044] In a specific example, the system extracts the node identifier of the second target node B obtained by the previous matching as "Node B ==4001", then hash the node identifier using the SHA-256 algorithm and extract a coded segment from the hash result as the second data segment "xxxxx-yyy". This coded segment, together with the previously generated data segment "zzzzzz", forms the complete second data identifier locator "xxxxx-yyy-00-zzzzzz". Through this structure, the system achieves precise binding of data to the node it belongs to while retaining the abstractness and security of the identifier, laying the foundation for target node confirmation in the bidirectional routing mechanism.
[0045] Sub-step 2014: Generate a third data segment in the second data identifier locator of the second target data according to the hash value of the data space of the second target data.
[0046] In some embodiments of the present application, in order to construct a routable identification structure for the data space, the system calculates a hash value based on the information of the data space to which the second target data belongs (such as the data space name or code), and uses the hash value as the third data segment in the second data identifier locator. This process usually uses a secure hash algorithm (such as SHA-256) to convert the data space name into a unique and stable hash code to represent the logical space area to which the data belongs. The generation of the data space hash value not only enables cross-domain identification, but also serves as the basis for distributed path selection and spatial clustering in the identity resolution network. By generating this third data segment, the system introduces a spatial logical dimension into the identification structure, providing a structural basis for subsequent node identification, path matching, and cross-domain management.
[0047] In a specific example, a data identifier "pku.cs / yyds801" belongs to the data space named "pku." The system uses "pku" as input, calculates its hash value using the SHA-256 algorithm, and extracts the first several bits of the code as the third data segment "xxxxx." This field serves as the first segment of the data identifier locator "xxxxx-yyy-00-zzzzzz," representing the logical space to which the data belongs. This provides an identification basis for subsequent routing table classification and space aggregation, and strengthens the ability to manage consistent data attribution across nodes.
[0048] Sub-step 2015: Generate an identification index instruction based on at least one of the generated first data segment, second data segment, and third data segment.
[0049] In some embodiments of the present application, in order to complete the expression and propagation mechanism of the identification path based on known identification information, the system will use at least one of the generated first data segment (zzzzzz), second data segment (yyy), and third data segment (xxxxx) to construct an identification index instruction. This process can form a structured instruction message by embedding the selected data segment into a standardized index instruction template. The generated index instruction contains complete or partial data identifier locator information as the basis for forwarding and indexing records between nodes. The first data segment is used to identify a specific data item, the second data segment carries node identification information, and the third data segment reflects the space to which the data belongs. Through dynamic combination, instruction design and network adaptation of different granularities can be achieved. In this way, the system can support the fine propagation and distributed recording of index information, enhancing the flexibility and structural integrity of the data location path.
[0050] In a specific example, node A has obtained the first data segment "zzzzzz" of a data identifier "pku.cs / yyds801," belonging to the data space "pku," corresponding to the third data segment "xxxxx," whose target storage node is node B. The node identifier is hashed to obtain the second data segment "yyy." The system concatenates these three segments into the format "xxxxx-yyy-00-zzzzzz" and encapsulates them into an identifier index instruction. If the network policy only requires the storage path index, the system can also construct a structurally abbreviated instruction based on "zzz" and "yyy." This instruction is then sent to downstream nodes in the network, triggering the generation of a reverse routing table and assisting in the subsequent query route construction.
[0051] Step 202, in response to the identification index instruction for the second target data, the routing entry of the Kademlia routing table established according to the second data identifier locator corresponding to the second target data in the identification index instruction is determined as the routing entry corresponding to the second data identifier locator in the second routing table of the first node.
[0052] In some embodiments of the present application, in order to establish a forward addressing path during the identification information propagation process, when responding to an identification index instruction for the second target data, the first node will identify the optimal routing entry in the corresponding Kademlia routing table based on the second data identification locator contained in the instruction, and incorporate the routing entry into the second routing table (i.e., the forward routing table) maintained by itself. This process is based on the XOR distance calculation logic between the hash value of the prefix segment in the locator and the current node ID. By selecting nodes with a closer logical distance, the routing entries in the Kademlia structure are supplemented and dynamically updated, so that the node has a more efficient jump path selection capability in subsequent resolution operations. The Kademlia routing table organizes routing entries in multiple K-buckets, each K-bucket is responsible for storing node information within a distance interval, ensuring the stability of the routing lookup process and the controllability of the number of hops in a distributed network environment.
[0053] In a specific example, suppose that node A initiates an identification index instruction for the data "pku.cs / yyds801" and forwards the instruction to the target node B through the path in the second routing table. During the forwarding process, the system recognizes that the locator is close to node A in the logical space and belongs to the coverage range of a K bucket in its Kademlia structure. Node A compares the target locator in this instruction with its local K bucket record, and after confirming the correspondence between node B and the locator, it registers B's network address, node identifier and related information as the corresponding routing item in the second routing table. This operation not only enhances the forward addressing capability of node A, but also improves the addressing connectivity of the entire network, providing a path selection basis for subsequent query processes initiated based on the locator.
[0054] The forward routing table maintenance method adopted in this application is based on the K-bucket mechanism of the Kademlia routing structure. This mechanism combines the logical distance design and efficient update strategy of the Distributed Hash Table (DHT) to achieve fast adaptation and low-complexity maintenance in multiple scenarios such as node joining, identity propagation, and path indexing: Each node maintains a forward routing table consisting of multiple K-buckets. The entire node ID space is divided into 160 logical distance intervals, and the i-th K-bucket records the distance from the node ID to the node ID. Information about other nodes within the range (if the K-bucket exceeds its capacity, the K-bucket can be adjusted based on preset removal conditions (such as data activity)); each K-bucket can store up to K routing entries, ensuring that the routing table size is controlled and covers the entire network; when the node discovers a new node (for example, after receiving a message from it or indirectly learning about it through other nodes), it calculates the XOR distance between the node's ID and its own ID and uses this to determine the K-bucket it should enter. If the target K-bucket is not yet full, the new node information is directly added. If the K-bucket is full, the Kademlia protocol prioritizes nodes with more active responses - it will attempt to ping the oldest node and replace it if it does not respond; otherwise, the original entry will be retained; during the forwarding process of executing the identity index instruction (for example, node A transmits the identity index to node B via the forward path), the passing nodes will record the source node or forward node information in the corresponding K-bucket based on the logical distance of the target identity locator. (For example, when a node receives an index instruction in a path, it registers the next-hop node information in the K bucket corresponding to its logical distance from the target identifier, thereby establishing a forwarding point for future forward queries.) This forward routing table has strong network adaptability. In scenarios where nodes frequently go online and offline, identifiers frequently change, or the network topology dynamically adjusts, a replacement strategy based on response activity and distance ensures that the overall structure is not frequently reconstructed, thereby maintaining the continuity of the query path.
[0055] Step 203 : In response to the identification index instruction for the second target data obtained from the fifth node, a routing entry for the second target data in the first routing table is generated according to the node information of the fifth node.
[0056] In some embodiments of the present application, in order to dynamically construct a reverse addressing path during the identification index transmission process, after receiving the identification index instruction for the second target data from the fifth node, the first node will generate and update the corresponding routing item in its first routing table (i.e., the reverse routing table) based on the node information of the fifth node. This operation enables this node to record the logical association between the fifth node and the second data identifier, and provide reverse path support for subsequent query requests based on the identifier. Each routing item in the first routing table includes a node identifier, address information, and a set of data identifiers contained therein, and the identification information is recorded through a Bloom filter to support a space-efficient, dynamically scalable storage structure. Through this mechanism, even if the data identifier is not initially issued by the current node, a complete reverse transmission link can be gradually established during the index message propagation process, thereby enhancing the traceability of the parsing process and the query landing capability.
[0057] like Figure 4As shown, in a specific example, node C receives an identification index instruction from the fifth node B, which contains locator information related to the target data "xxxxx-yyy-00-zzzzzz". The system identifies the fifth node B as the transmitter of the previous hop index message, and then generates or updates a new routing entry in the reverse routing table of node C, which records B's node ID, network address and the Bloom filter bit vector related to the identifier. This entry is used to indicate that "the target data locator can be obtained by tracing back from the direction of node B", so that when a node issues a query instruction for the data identifier to C in the future, the query can be routed back to the direction of B based on the reverse record, reducing the complexity of the path search and improving the implementation efficiency.
[0058] Step 204 : In response to the identification query instruction for the first target data, matching the first data identifier locator corresponding to the first target data in the identification query instruction with the first routing table of the first node.
[0059] The first data segment in the first data identifier locator uniquely corresponds to the first target data, and the second data segment in the first data identifier locator uniquely corresponds to the first target node storing the first target data.
[0060] The method shown in this step has been described in step 101 and will not be repeated here.
[0061] Step 205: When the first data identifier locator successfully matches the first routing table, the identifier query instruction is forwarded to the second node according to the first routing table.
[0062] The method shown in this step has been described in step 102 and will not be repeated here.
[0063] Step 206: If the first data identifier locator fails to match the first routing table, match the first data identifier locator with the second routing table of the first node.
[0064] The method shown in this step has been described in step 103 and will not be repeated here.
[0065] Step 207: When the first data identifier locator successfully matches the second routing table, the identifier query instruction is forwarded to the third node according to the second routing table.
[0066] The method shown in this step has been described in step 104 and will not be repeated here.
[0067] Step 208: If the first data identifier locator fails to match the first routing table and the second routing table, the first node is determined as the first target node.
[0068] The method shown in this step has been explained in step 105 and will not be repeated here.
[0069] Figure 5 This is another distributed data identification resolution method based on bidirectional routing provided in an embodiment of the present application, which is applied to the first node in a distributed storage network.
[0070] The method may include the following steps: Step 301: In response to an identification index instruction, matching a second data identifier locator corresponding to second target data in the identification index instruction with a second routing table of a first node.
[0071] In some embodiments of the present application, in order to establish a forward addressing path suitable for subsequent parsing tasks, when the first node receives an identification index instruction, it extracts a second data identification locator corresponding to the second target data from the instruction, and matches the locator with the second routing table (i.e., the forward routing table) maintained by this node. The matching process calculates the logical distance between the prefix hash value of the locator and the node ID of each routing item to find the optimal matching bucket under the Kademlia structure to determine whether the current node has the ability to continue forwarding the index instruction. If the match is successful, it means that the current node already has a forwarding path that is logically closer to the data target, which helps the data index message to accurately and quickly approach the target storage node.
[0072] like Figure 3 As shown, in a specific example, node C receives an identification index instruction containing a locator of "xxxxx-yyy-00-zzzzzz". The system uses the prefix "xxxxx" of the locator as the logical space information and performs distance matching with each K bucket in the second routing table maintained by node C. If a K bucket is found to contain a node that is closer to the locator, such as the next-hop node D that matches "xxxxx" in the figure, node C confirms that the match is successful. In subsequent steps, the index instruction can be forwarded to the downstream node based on the routing path, which helps to continuously improve the mapping relationship between data identification and network structure.
[0073] Step 302: When the second data identifier locator successfully matches the second routing table, forward the identifier index instruction to the fourth node according to the second routing table.
[0074] In some embodiments of the present application, in order to effectively transmit the identification index instruction in a direction that is logically closer to the target data storage location, when the second data identification locator successfully matches the second routing table (forward routing table), the system will forward the index instruction to the fourth node based on the node information matched in the forward routing table. This process is based on the logical distance principle of the Kademlia structure. By selecting the next-hop neighbor node with a smaller XOR distance, the identification information is gradually forwarded to the target node, ensuring the orderly advancement and update of the index path structure. Through the establishment of this path, the system can complete the structural propagation of data in the early stage of identification resolution, providing a path support and efficient routing foundation for subsequent reverse query and identification positioning.
[0075] like Figure 3 As shown, in a specific example, when node C matches the locator "xxxxx-yyy-00-zzzzzz" in its forward routing table, it determines that the fourth node is the logically adjacent node D on the current query path. The system uses the node address information in the routing item to forward the identification index instruction from node C to the fourth node D, thereby continuing the identification path from the source node to the target storage node. This not only builds an index propagation chain, but also provides node references along the way for the accumulation of records in the reverse routing table. Ultimately, this transmission link will improve the system's addressability and query response speed for the data identifier.
[0076] Step 303: In response to the identification query instruction for the first target data, match the first data identifier locator corresponding to the first target data in the identification query instruction with the first routing table of the first node.
[0077] The first data segment in the first data identifier locator uniquely corresponds to the first target data, and the second data segment in the first data identifier locator uniquely corresponds to the first target node storing the first target data.
[0078] The method shown in this step has been described in step 101 and will not be repeated here.
[0079] Step 304: When the first data identifier locator successfully matches the first routing table, the identifier query instruction is forwarded to the second node according to the first routing table.
[0080] The method shown in this step has been described in step 102 and will not be repeated here.
[0081] Step 305: If the first data identifier locator fails to match the first routing table, the first data identifier locator is matched with the second routing table of the first node.
[0082] The method shown in this step has been described in step 103 and will not be repeated here.
[0083] Step 306: If the first data identifier locator successfully matches the second routing table, forward the identifier query instruction to the third node according to the second routing table.
[0084] The method shown in this step has been described in step 104 and will not be repeated here.
[0085] Step 307: If the first data identifier locator fails to match the first routing table and the second routing table, the first node is determined as the first target node.
[0086] The method shown in this step has been explained in step 105 and will not be repeated here.
[0087] Step 308: When the first data identifier locator successfully matches the first routing table or the second routing table, generate or update identifier query history information for the first target data.
[0088] In some embodiments of the present application, in order to improve the system's request management capabilities for data identifiers and optimize subsequent caching strategies, when the first data identifier locator successfully matches the first routing table (reverse routing table) or the second routing table (forward routing table) of the first node, the system will immediately generate or update an identifier query history information for the first target data. This historical information records the interaction trajectory of the data identifier in the addressing process of the current node, and may include structured metadata such as access time, number of hops, and path nodes. The continuous accumulation of this information can serve as an important basis for subsequent caching strategy formulation and data priority evaluation, supporting the system to dynamically identify hot data in query-intensive scenarios, thereby improving overall performance and query response efficiency.
[0089] In a specific example, node C receives a query request for the data identifier "xxxxx-yyy-00-zzzzzz." Upon searching, it finds a match for this identifier in its reverse routing table, pointing to node B. After the system successfully completes the redirect, it adds or updates a record for this identifier in node C's local query history, including information such as the forwarding timestamp, transmission direction, and downstream node identifier. If this identifier has been forwarded or matched multiple times, the access frequency field is also updated accordingly. This historical information provides a useful reference for determining whether the data should be cached on this node.
[0090] Step 309: Determine a cache strategy for the first target data according to the identified query history information, and execute the cache strategy.
[0091] In some embodiments of the present application, in order to improve the data access efficiency of the system and the load balancing capability between nodes, the system will determine and execute a corresponding caching strategy for the first target data based on the existing identification query history information. The caching strategy usually combines dimensional factors such as access frequency, routing hops, and node activity to evaluate the priority score of the target data. On this basis, the system determines whether the current node should cache the data, and decides whether to add, replace, or delay processing of the data content based on the current cache resource situation. The caching mechanism consists of two parts: one is based on the data content and indicator records maintained by the hash table, and the other is based on the priority queue stored in the ordered set. The two are used together to support efficient retrieval and dynamic replacement. When the execution strategy is completed, the system can prioritize the retention of high-value, high-access frequency data under limited memory conditions, improve the cache hit rate and reduce redundant query costs.
[0092] In a specific example, after node C completes parsing and responding to the target data identified by "xxxxx-yyy-00-zzzzzz," the system reads the query history information for that identifier and discovers that the data has been accessed frequently, has a large number of path hops, and is queried from multiple neighboring nodes. The system assesses this data as having high cache value. If the cache is not yet full, node C directly writes the data and its associated metadata into the cache. If the cache is full, node C compares it with the currently lowest-priority data and decides whether to replace it. Through this strategy, the system implements dynamic caching decisions driven by historical query behavior, enhancing the real-time nature of data access and the efficiency of system resource utilization.
[0093] Optionally, in order to determine a cache strategy for the first target data according to the identification query history information, step 309 includes the following sub-steps: Sub-step 3091, determines the cache priority evaluation value of the first target data based on at least one of the activity evaluation value of the second data segment in the first data identifier locator, the access frequency of the first target data, and the distance between each node that queries the first target data and the first node, based on the identification query history information.
[0094] In some embodiments of the present application, in order to quantitatively evaluate the cache priority of the first target data at the current node, the system will determine multiple feature values closely related to the cache requirements based on the existing identification query history information, including the activity evaluation value of the second data segment in the first data identifier locator, the access frequency of the first target data, and at least one of the distances between each node that queries the data and the current node. This process calculates the corresponding feature weights in sequence by reading the metadata related to node interaction, access trajectory and hop path in the historical records, and aggregates them to generate a cache priority evaluation value. Among them, the second data segment is a structure field obtained by hashing the identification information of the node to which the data belongs, which is used to indicate the node identity code in the storage path; its activity evaluation value reflects the request forwarding frequency and response capability of the target node in the recent network. After executing this step, the system can establish a cache priority reference index that integrates access behavior characteristics and network structure attributes, providing a basic basis for the optimization selection of subsequent cache strategies.
[0095] In a specific example, node C records that the first target data "xxxxx-yyy-00-zzzzzz" has been queried 9 times in the past hour, and comes from 3 neighboring nodes in different logical areas, and the average number of query path hops is 2. The system reads the second data segment "yyy" corresponding to the identifier, and based on the frequency of occurrence of node "yyy" in other data identifiers and query traffic in history, it evaluates that its activity is relatively high. Subsequently, the system uses access frequency, average path distance, and node activity as input features, and aggregates and calculates that the cache priority evaluation value of the target data is 0.87. This result will be used in the next step to participate in the priority sorting of cache entries and to determine whether to retain or replace the cache.
[0096] In a further embodiment of the present application, to achieve a quantitative assessment of the cache value of the first target data, the system defines a cache priority evaluation value calculation formula, which comprehensively scores the necessity of caching by aggregating multiple influencing factors. The general form of the formula is as follows:
[0097] Where: P represents the cache priority evaluation value; F is the access frequency of the target data, which can be expressed as the number of hits per unit time; D is the average logical distance between the access node and the current node, usually measured by the number of hops or XOR distance; A is the activity evaluation value of the node corresponding to the second data segment, reflecting the interaction frequency of the node in the overall network; w1, w2, and w3 are the weight coefficients corresponding to each factor, which are configured or dynamically adjusted according to the system policy.
[0098] Sub-step 3092: determining a cache strategy according to the ranking position of the cache priority evaluation value of the first target data among the multiple cache priority evaluation values stored in the first node.
[0099] In some embodiments of the present application, in order to give priority to retaining the most valuable data under limited cache resources, the system will determine the corresponding cache strategy based on the ranking position of the cache priority evaluation value of the first target data in the multiple evaluation values stored in the first node. This process compares the priority of the evaluation value of the current cache object with the multiple cache items already existing in the system, and determines whether to perform cache write, replacement or rejection operations based on its ranking position in the cache queue. The cache strategy may include branch paths such as direct insertion, elimination of low priority items or temporary storage. The evaluation value sorting structure usually adopts the form of ordered set or priority queue management, combined with cache strategies such as Least Recently Used (LRU) or Least Frequently Used (LFU), to balance the hit rate and resource utilization efficiency in dynamic adjustment. After executing this step, the system can implement precise cache control strategies based on the sorting mechanism, improve the residence rate of hot data and reduce access latency.
[0100] In a specific example, node C calculates a cache priority of 0.87 for the first target data item "xxxxx-yyy-00-zzzzzz." The node's cache already contains five items, with corresponding priorities of 0.93, 0.89, 0.85, 0.81, and 0.78. After inserting item 0.87 into the ordered queue, the system discovers that the cache is full, but the target data item is ranked third, higher than the item originally ranked last (0.78). The system then decides to eliminate the lowest-priority item and write "xxxxx-yyy-00-zzzzzz" into the cache, thus implementing a replacement caching strategy driven by sorting position. This strategy ensures that the cache always serves high-value data access, enhancing the node's local responsiveness.
[0101] Step 310 : Record the total cache duration of the first target data in the first node, and remove the cached first target data from the cache queue when the total cache duration exceeds a preset life cycle threshold.
[0102] In some embodiments of the present application, in order to ensure efficient use of cache resources and avoid long-term occupation of invalid data, the system will continuously record the total cache duration of the data after the first target data is cached in the first node, and dynamically compare it with the preset lifetime threshold (Time To Live, TTL). This process is usually implemented by setting a cache write timestamp and periodically checking the survival status of the cache item. When it is detected that the survival duration of a certain data exceeds the lifetime, the system automatically removes it from the cache queue. This time threshold-based management mechanism can effectively control the activity level of data items in the cache, prevent hot spots from still occupying space after they subside, and improve the real-time and flow efficiency of the cache structure.
[0103] In a specific example, node C writes the first target data identified as "xxxxx-yyy-00-zzzzzz" into the cache according to the policy, and sets its life cycle threshold to 24 hours. The system records the starting timestamp for the cache item and monitors its cumulative time in the cache through scheduled tasks or request triggers. When it is found that the data has been cached for more than 24 hours and there has been no access or update behavior during this period, the system will synchronously delete it from the cache hash table and priority sorting structure, thereby freeing up cache space for subsequent higher-priority data to ensure the timeliness and accuracy of resource utilization.
[0104] like Figure 6 As shown, this is the program execution process of a cache decision process under an embodiment of the present application. The cache decision process follows the "priority value driven + path backup" mechanism, realizes dynamic scheduling and priority reservation of cache resources, and improves the response efficiency of node service hotspot data and the overall access performance of the system: R1. Node data collection: The node monitors and receives data resources incoming from the network, triggering the cache strategy judgment process; R2. Determine whether it is target data: The system determines whether the data belongs to the predefined first target data type based on the data identifier, access source, or parsing path. If yes, the system proceeds to the next step; otherwise, the process ends. R3. Calculate priority: The system calls the priority evaluation algorithm and calculates the cache priority evaluation value of the data by combining indicators such as access frequency, query path distance, and target node activity; R4: Is the data priority high? The evaluation value is compared with the priority values of the node's existing cache. If the current data is at the front of the cache queue (for example, above a certain threshold or at least above the end), it enters the primary cache process, otherwise it enters the backup decision process. R5. Determine whether it can be used as a backup data cache: If the primary cache conditions are not met, the system will determine whether it is suitable to store the data in the backup cache area based on the policy for standby transfer or subsequent policy evaluation; if the judgment result is no, the data will not be cached.
[0105] like Figure 7 As shown, the program execution process of the identification query under the cache decision process of an embodiment of the present application is shown. The entire query process follows the hierarchical path selection strategy of "cache priority, storage supplement, and network backup". This process takes into account system resource utilization and path traceability while ensuring response speed: S1. Accepting query request: The current node receives an identification query request from an upstream node or user, triggering the query process; S2. Query cache: The system first searches the local cache for a data item that matches the target identifier. S3. Cache query results: If the cache hits, the query results are returned directly and the process ends; otherwise, proceed to the next step; S4. Check local storage: The system searches for the target data in the local persistent storage as a supplementary path in case of cache misses. S5. Local storage query result: If the local storage hits, the result is returned; otherwise, the query continues to the network; S6, routing query suitable node: The system selects the node with the closest logical distance according to the forward routing table and forwards the query request; S7. Mark the reverse direction of message transmission: During the forwarding process, the system records the reverse path information between the current node and the upstream node for subsequent response return: S8, query completion node: The query is completed at the target node or intermediate node, the system marks the query as complete, and returns the result through the reverse path.
[0106] In summary, in the embodiment of the present application, the unique correspondence between the first data segment in the data identifier locator and the target data is firstly utilized so that the node can accurately identify the identification instruction of the target data, thereby ensuring that the parsing instruction has a clear addressing target and providing a semantically clear starting point for subsequent routing selection; at the same time, the second data segment is uniquely corresponding to the target node, further enhancing the adaptability of the locator to the physical network topology. This structural encoding lays the foundation for subsequent efficient matching and helps to improve the accuracy of the overall parsing instruction; further, in the routing matching process, if the first data identifier locator is successfully matched in the first routing table, it can be directly passed. The system quickly reaches the next hop node in the direction of the target data through the reverse path to achieve rapid data resolution; when the first routing table fails to match, the system automatically switches to the second routing table and starts the forward path addressing mechanism, so that the system still has the ability to resolve the target in abnormal situations such as the node distance is far, the node is missing, the network is interrupted, or the number of path hops increases abnormally, effectively avoiding the query failure problem caused by the singleness of the path; further, when both the forward and reverse tables are not hit, it can be determined that the current node is the target node where the data is stored, and the query process is terminated, while reducing the waste of network resources. It provides a small range of data storage nodes with resolution autonomy. Therefore, based on the method of the embodiment of the present application, by constructing a dual-table matching mechanism including a forward routing table and a reverse routing table, the redundant design and fault tolerance improvement of the resolution path are achieved, the performance bottleneck caused by the single routing structure is effectively alleviated, and the query efficiency and reliability of the data identification are improved.
[0107] refer to Figure 8 , which shows a distributed data identification resolution device 40 based on bidirectional routing provided by an embodiment of the present application, applied to a first node in a distributed storage network, including: A first matching module 401 is configured to, in response to an identification query instruction for first target data, match a first data identifier locator corresponding to the first target data in the identification query instruction with a first routing table of a first node; wherein a first data segment in the first data identifier locator uniquely corresponds to the first target data, and a second data segment in the first data identifier locator uniquely corresponds to a first target node storing the first target data; A first forwarding module 402 is configured to forward the identification query instruction to the second node according to the first routing table if the first data identifier locator successfully matches the first routing table; A second matching module 403 is configured to match the first data identifier locator with the second routing table of the first node if the first data identifier locator fails to match the first routing table; A second forwarding module 404 is configured to forward the identification query instruction to a third node according to the second routing table if the first data identification locator successfully matches the second routing table; The query confirmation module 405 is configured to determine the first node as the first target node when the first data identifier locator fails to match the first routing table and the second routing table.
[0108] Optionally, the distributed data identification resolution device 40 based on bidirectional routing further includes: The first table building module is configured to generate a routing entry for the second target data in the first routing table according to the node information of the fifth node in response to an identification index instruction for the second target data obtained from the fifth node.
[0109] Optionally, the distributed data identification resolution device 40 based on bidirectional routing further includes: The second table building module is used to respond to the identification index instruction of the second target data and determine the routing entry of the Kademlia routing table established according to the second data identification locator corresponding to the second target data in the identification index instruction as the routing entry corresponding to the second data identification locator in the second routing table of the first node.
[0110] Optionally, the distributed data identification resolution device 40 based on bidirectional routing further includes: The index initiating module is configured to determine a second target node for storing a second data identifier locator in response to the second target data stored in the first node, and generate an identifier index instruction.
[0111] Optionally, the index initiation module includes: a first hash submodule, configured to generate a first data segment in a second data identifier locator of the second target data according to a hash value of the data identifier of the second target data; a node confirmation submodule, configured to match the first data segment in the second data identifier locator with the second routing table of the first node to obtain node information of the second target node; A second hash submodule, configured to generate a second data segment in a second data identifier locator of the second target data according to a hash value of the node identifier in the node information of the second target node; The index initiating submodule is used to generate an identification index instruction according to the generated first data segment and second data segment.
[0112] Optionally, the distributed data identification resolution device 40 based on bidirectional routing further includes: an index matching module, configured to, in response to the identification index instruction, match a second data identifier locator corresponding to the second target data in the identification index instruction with the second routing table of the first node; The index forwarding module is configured to forward the identification index instruction to the fourth node according to the second routing table when the second data identification locator successfully matches the second routing table.
[0113] Optionally, the distributed data identification resolution device 40 based on bidirectional routing further includes: A history recording module, configured to generate or update identification query history information for the first target data when the first data identifier locator successfully matches the first routing table or the second routing table; The cache strategy module is used to determine a cache strategy for the first target data according to the identification query history information and execute the cache strategy.
[0114] Optionally, the cache strategy module includes: a priority evaluation submodule, configured to determine a cache priority evaluation value for the first target data based on at least one of an activity evaluation value for the second data segment in the first data identifier locator, an access frequency to the first target data, and a distance between each node that queries the first target data and the first node; The cache strategy submodule is configured to determine a cache strategy according to a ranking position of the cache priority evaluation value of the first target data among the plurality of cache priority evaluation values stored in the first node.
[0115] Optionally, the distributed data identification resolution device 40 based on bidirectional routing further includes: The cycle management module is used to record the total cache time of the first target data cached in the first node, and remove the cached first target data from the cache queue when the total cache time exceeds a preset life cycle threshold.
[0116] In summary, in the embodiment of the present application, the unique correspondence between the first data segment in the data identifier locator and the target data is firstly utilized so that the node can accurately identify the identification instruction of the target data, thereby ensuring that the parsing instruction has a clear addressing target and providing a semantically clear starting point for subsequent routing selection; at the same time, the second data segment is uniquely corresponding to the target node, further enhancing the adaptability of the locator to the physical network topology. This structural encoding lays the foundation for subsequent efficient matching and helps to improve the accuracy of the overall parsing instruction; further, in the routing matching process, if the first data identifier locator is successfully matched in the first routing table, it can be directly passed. The system quickly reaches the next hop node in the direction of the target data through the reverse path to achieve rapid data resolution; when the first routing table fails to match, the system automatically switches to the second routing table and starts the forward path addressing mechanism, so that the system still has the ability to resolve the target in abnormal situations such as the node distance is far, the node is missing, the network is interrupted, or the number of path hops increases abnormally, effectively avoiding the query failure problem caused by the singleness of the path; further, when both the forward and reverse tables are not hit, it can be determined that the current node is the target node where the data is stored, and the query process is terminated, while reducing the waste of network resources. It provides a small range of data storage nodes with resolution autonomy. Therefore, based on the method of the embodiment of the present application, by constructing a dual-table matching mechanism including a forward routing table and a reverse routing table, the redundant design and fault tolerance improvement of the resolution path are achieved, the performance bottleneck caused by the single routing structure is effectively alleviated, and the query efficiency and reliability of the data identification are improved.
[0117] Reference Figure 9 , is a block diagram of an electronic device 500 according to another embodiment of the present invention. For example, electronic device 500 may be provided as a server. Electronic device 500 may include one or more of the following components: a processing component 502, a memory 504, a power supply component 506, a multimedia component 508, an audio component 510, an input / output (I / O) interface 512, a sensor component 514, and a communication component 516.
[0118] Electronic device 500 includes a processing component 502, which further includes one or more processors, and a memory resource represented by memory 504 for storing instructions executable by processing component 502, such as applications. The application stored in memory 504 may include one or more modules, each corresponding to a set of instructions. In addition, processing component 502 is configured to execute instructions to perform the methods provided in the embodiments of the present application.
[0119] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 502 may include one or more modules to facilitate interaction between the processing component 502 and other components.
[0120] The memory 504 is used to store various types of data to support operations in the electronic device 500. The memory 504 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0121] The power supply assembly 506 provides power to the various components of the electronic device 500. The power supply assembly 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 500.
[0122] Multimedia component 508 includes an interface that provides an output interface between electronic device 500 and a user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). In some embodiments, multimedia component 508 includes a front-facing camera and / or a rear-facing camera.
[0123] The audio component 510 is used to output and / or input audio signals. The received audio signals can be further stored in the memory 504 or transmitted via the communication component 516.
[0124] The input / output (I / O) interface 512 provides an interface between the processing component 502 and the peripheral interface module.
[0125] Sensor assembly 514 includes one or more sensors for providing various status assessments for electronic device 500. For example, sensor assembly 514 can detect the open / closed state of electronic device 500 and the relative positioning of components. Sensor assembly 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 514 can also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications.
[0126] The communication component 516 is used to facilitate wired or wireless communication between the electronic device 500 and other devices. The electronic device 500 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof.
[0127] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to implement the methods provided in the embodiments of the present application.
[0128] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, which can be executed by the processor 520 of the electronic device 500 to perform the above method.
[0129] The electronic device 500 may also operate based on an operating system stored in the memory 504, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0130] It should be noted that, for the sake of simplicity, the method embodiments of the present application are described as a series of action combinations. However, those skilled in the art should be aware that the embodiments of the present application are not limited by the order of the actions described, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0131] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the application disclosed herein. This application is intended to encompass any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art that are not disclosed herein. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims claimed below.
[0132] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the claims.
Claims
1. A distributed data identification resolution method based on bidirectional routing, characterized in that: A first node applied to a distributed storage network includes: In response to an identification query instruction for first target data, matching a first data identifier locator corresponding to the first target data in the identification query instruction with a first routing table of the first node; a first data segment in the first data identifier locator uniquely corresponds to the first target data, and a second data segment in the first data identifier locator uniquely corresponds to a first target node storing the first target data; If the first data identifier locator successfully matches the first routing table, forwarding the identifier query instruction to the second node according to the first routing table; If the first data identifier locator fails to match the first routing table, matching the first data identifier locator with the second routing table of the first node; If the first data identifier locator successfully matches the second routing table, forwarding the identifier query instruction to a third node according to the second routing table; In a case where the first data identifier locator fails to be matched with the first routing table and the second routing table, the first node is determined as the first target node.
2. The distributed data identification resolution method based on bidirectional routing according to claim 1, characterized in that: The distributed data identification resolution method based on bidirectional routing also includes: In response to the second target data stored in the first node, a second target node for storing a second data identifier locator is determined, and an identifier index instruction is generated.
3. The distributed data identification resolution method based on bidirectional routing according to claim 2, characterized in that: The step of determining, in response to the second target data stored in the first node, a second target node for storing the second data identifier locator and generating the identifier index instruction comprises: generating a first data segment in a second data identifier locator of the second target data according to a hash value of the data identifier of the second target data; Matching the first data segment in the second data identifier locator with the second routing table of the first node to obtain node information of the second target node; generating a second data segment in a second data identifier locator of the second target data according to a hash value of the node identifier in the node information of the second target node; The identification index instruction is generated according to the generated first data segment and the second data segment.
4. The distributed data identification resolution method based on bidirectional routing according to claim 1, characterized in that: The distributed data identification resolution method based on bidirectional routing also includes: In response to the identification index instruction, matching a second data identifier locator corresponding to the second target data in the identification index instruction with a second routing table of the first node; In a case where the second data identifier locator successfully matches the second routing table, the identification index instruction is forwarded to the fourth node according to the second routing table.
5. The distributed data identification resolution method based on bidirectional routing according to claim 1, characterized in that: The distributed data identification resolution method based on bidirectional routing also includes: In response to the identification index instruction for the second target data obtained from the fifth node, a routing entry for the second target data in the first routing table is generated according to the node information of the fifth node.
6. The distributed data identification resolution method based on bidirectional routing according to claim 1, characterized in that: The distributed data identification resolution method based on bidirectional routing also includes: If the first data identifier locator successfully matches the first routing table or the second routing table, generating or updating identification query history information for the first target data; A cache strategy for the first target data is determined according to the identification query history information, and the cache strategy is executed.
7. The distributed data identification resolution method based on bidirectional routing according to claim 6, characterized in that: The determining of a cache strategy for the first target data according to the identification query history information includes: determining a cache priority evaluation value for the first target data based on at least one of an activity evaluation value for the second data segment in the first data identifier locator, an access frequency to the first target data, and a distance between each node that queries the first target data and the first node according to the identifier query history information; The cache strategy is determined according to a ranking position of the cache priority evaluation value of the first target data among the multiple cache priority evaluation values stored in the first node.
8. A distributed data identification resolution device based on bidirectional routing, characterized in that: A first node applied to a distributed storage network includes: a first matching module, configured to, in response to an identification query instruction for first target data, match a first data identifier locator corresponding to the first target data in the identification query instruction with a first routing table of the first node; wherein a first data segment in the first data identifier locator uniquely corresponds to the first target data, and a second data segment in the first data identifier locator uniquely corresponds to a first target node storing the first target data; a first forwarding module, configured to forward the identification query instruction to a second node according to the first routing table when the first data identifier locator successfully matches the first routing table; a second matching module, configured to match the first data identifier locator with a second routing table of the first node if the first data identifier locator fails to match the first routing table; a second forwarding module, configured to forward the identification query instruction to a third node according to the second routing table when the first data identifier locator successfully matches the second routing table; The query confirmation module is configured to determine the first node as the first target node if the first data identifier locator fails to match the first routing table and the second routing table.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the distributed data identification resolution method based on bidirectional routing according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The method comprises a processor, a memory and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the distributed data identification resolution method based on bidirectional routing as described in any one of claims 1 to 7 are implemented.