A decentralized resource discovery and retrieval system based on double-coordinate system indexing and semantic-aware routing
Patent Information
- Application Number
- CN202610988997.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-29
AI Technical Summary
该类方式虽然能够支持语义相关资源检索,但在去中心化分布式节点环境中,若将各节点资源信息集中维护于统一索引节点,容易产生索引维护成本高、节点自治性降低以及单点瓶颈等问题;若通过广播或泛洪方式进行语义查询,则会造成较大的网络通信开销,难以适用于大规模分布式节点环境
[0161]与现有技术相比,本发明的有益效果在于:本发明实现了去中心化环境下精确寻址与语义近似发现的协同处理,在不集中存储资源原始数据且避免全网广播的前提下,提高了分布式资源发现的准确性、召回能力和通信效率。
Smart Images

Figure CN122838465A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer networks and distributed information retrieval technology, and in particular to a resource discovery and retrieval system running in a decentralized distributed node environment. Specifically, it relates to a distributed resource discovery and retrieval technology that constructs a dual-coordinate index of resources based on identifier coordinates and semantic coordinates, and combines semantic-aware routing to achieve precise resource addressing, semantic approximation discovery, and backup tracing. Background Technology
[0002] With the development of distributed networks, edge computing, and decentralized applications, various data resources, service resources, model resources, and knowledge resources are gradually being deployed across multiple autonomous nodes. Existing resource discovery technologies typically employ distributed addressing methods based on resource identifiers, Uniform Resource Identifiers (URIs), hash keys, or node identifiers to locate the responsible node corresponding to a resource within the logical addressing space. While this approach is suitable for precise query scenarios with clearly defined resource identifiers, its addressing basis primarily stems from structured identification information, making it difficult to effectively process unstructured semantic queries that only contain natural language descriptions, topic requirements, or functional requirements.
[0003] Existing semantic retrieval technologies typically rely on centralized indexes, vector databases, or unified retrieval platforms to achieve resource matching through semantic vector similarity calculations. While these methods can support semantically relevant resource retrieval, in decentralized distributed node environments, centralizing resource information across nodes on a unified index node can easily lead to problems such as high index maintenance costs, reduced node autonomy, and single-point bottlenecks. Furthermore, using broadcast or flooding methods for semantic queries results in significant network communication overhead, making them unsuitable for large-scale distributed node environments.
[0004] Furthermore, existing distributed resource discovery technologies typically separate precise identifier addressing from semantic approximate retrieval, lacking an indexing mechanism capable of simultaneously expressing the precise location and semantic features of resources within the same decentralized addressing space. Therefore, when a query request contains both structured identifier information and unstructured semantic information, existing technologies struggle to adaptively select precise, semantic, or hybrid query paths based on the query payload composition, easily leading to query path mismatches, insufficient resource retrieval, or increased query costs.
[0005] Meanwhile, existing distributed routing mechanisms mostly rely on node identifiers, topological distances, or pre-defined routing tables for next-hop selection, lacking a mechanism to guide query packets hop-by-hop forwarding using the semantic features of local node resources. In semantic query scenarios, query packets may fail to approach the target resource node along semantically relevant directions, leading to repeated forwarding, partial stagnation, or insufficient search results. Even when some technologies introduce semantic similarity calculations, they typically lack node semantic profiling and semantic proximity routing mechanisms that are integrated with the distributed logical addressing space.
[0006] Furthermore, semantic routing is characterized by approximation and uncertainty. When a query packet fails to obtain a satisfactory result after passing through several nodes, expanding the query scope further can easily increase communication overhead; conversely, terminating the query directly may result in the omission of relevant resources. Existing technologies lack a controlled backup tracing mechanism that can supplement the discovery of candidate resources within a limited scope and avoid network-wide broadcasting and repeated access when semantic cooperative addressing results are insufficient.
[0007] Therefore, it is necessary to provide a resource discovery and retrieval system suitable for decentralized distributed node environments, so as to achieve the collaborative processing of precise resource addressing and semantic approximation discovery without centralized storage of raw resource data, and to perform controlled backup tracing when semantic query results are insufficient, thereby improving the accuracy, recall capability and communication efficiency of distributed resource discovery. Summary of the Invention
[0008] In view of the above, this invention provides a decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantic-aware routing.
[0009] To achieve the above objectives, the present invention provides the following technical solution:
[0010] A decentralized resource discovery and retrieval system based on dual-coordinate indexing and semantic-aware routing operates in a closed logical addressing space composed of multiple distributed nodes. It includes a dual-coordinate indexing construction module, a semantic anchor registration module, a node semantic profiling module, a query intent awareness module, a semantic collaborative addressing module, a hop-by-hop retrieval accumulation module, a semantic anchor tracing module, and a result fusion output module. Each module operates collaboratively according to the following process:
[0011] (1) The dual-coordinate index construction module obtains the resource identifier and resource semantic representation of the resource entity, performs deterministic mapping on the resource identifier, generates identifier coordinates for precise addressing, and maps the resource semantic representation through multiple sets of semantic projection functions to generate multiple semantic coordinates for semantic approximation discovery, and maps the identifier coordinates and multiple semantic coordinates to the same closed logical addressing space.
[0012] (2) The semantic anchor registration module determines the corresponding semantic anchor node based on multiple semantic coordinates, and registers a semantic pointer containing the resource semantic coordinates, resource identifier coordinates and the address of the node to which the resource belongs to the semantic anchor node for subsequent semantic anchor tracing.
[0013] (3) The node semantic profiling module generates one or more semantic centroids and corresponding semantic centroid coordinates of a node based on the semantic representation of the local hosted resources of each distributed node. It establishes a set of semantic neighbor candidate nodes based on the semantic centroid coordinates and centroid weights, and provides the set of semantic neighbor candidate nodes to the semantic collaborative addressing module.
[0014] (4) The query intent perception module parses the precise identifier field and unstructured semantic field in the query payload, determines the precise query mode, semantic query mode or mixed query mode based on the effectiveness of the precise identifier field and the semantic focus of the unstructured semantic field, and starts the corresponding query path.
[0015] (5) In precise query mode, the identification responsibility node is located and precise addressing results are obtained based on the precise identification field; in semantic query mode, the semantic collaborative addressing module generates query semantic coordinates based on the unstructured semantic field, merges the topology routing candidate nodes of the current node with the semantic neighbor candidate nodes, and selects the next hop node based on the proximity of the candidate node to the query semantic coordinates. The hop-by-hop retrieval accumulation module performs local resource retrieval and result deduplication accumulation at the nodes traversed by the query packet; in mixed query mode, precise query and semantic query are executed simultaneously; when the accumulated result meets the preset result requirements, semantic collaborative addressing stops; when semantic collaborative addressing terminates and the accumulated result does not meet the preset result requirements, the semantic anchor tracing module locates the query semantic anchor node based on the query semantic coordinates, obtains the semantic pointer, and initiates a tracing request to the node to which the corresponding resource belongs.
[0016] (6) The result fusion output module deduplicatizes the precise addressing results, hop-by-hop cumulative results and tracing results to generate the final resource discovery results.
[0017] Furthermore, the dual-coordinate index construction module includes:
[0018] First, the system obtains the resource entity. resource content And generate resource semantic representation vectors through semantic mapping. ;
[0019] Next, the system obtains the resource entity. Resource Identifier Primitives Resource identification primitives consist of unique identifiers for each resource. Resource types After fixed field order, unified encoding, field cleaning, null value placeholders, delimiter concatenation, and deterministic serialization, the data is then concatenated together. = ;
[0020] Based on the above resource identifier primitives The identifier coordinates are generated in the closed logical address space through a hash mapping function. Then, based on semantic resource representation Resource entities are generated through semantic projection functions. The L semantic coordinates are represented by the set of projection vectors as follows:
[0021]
[0022] in, , Indicates the first The first group semantic projection function One projection vector; The number of bits in the binary code corresponding to each semantic coordinate;
[0023] For resource entities In the The i-th semantic code under the group projection function is generated as follows:
[0024]
[0025] in, Represents resource entities semantic vectors In the The binarization result of the i-th projection direction; the binarization result is used to characterize the semantic distinction state of the resource semantic representation vector in the corresponding projection direction; then, the i-th... The group projection function obtained Bitwise codes are concatenated into integer coordinates to obtain the resource entity. The Semantic coordinates:
[0026]
[0027] Multiple semantic coordinates are combined to obtain a semantic coordinate set:
[0028]
[0029] Each semantic coordinate All fall into the closed logical address space;
[0030] Identify coordinates Used to locate resource entities Identifying the responsible node, a set of semantic coordinates Used to locate resource entities Multiple semantic anchor nodes, and the dual-coordinate index data does not contain the original resource data;
[0031] In addition, the L-group projection functions are also configured to allow the query intent awareness module to generate query semantic coordinates. And the node semantic centroid coordinates generated by the node semantic profiling module.
[0032] Furthermore, the semantic anchor registration module includes: based on resource entities L semantic coordinates In the closed logical address space, determine L corresponding semantic anchor nodes respectively; and send the corresponding semantic pointers to the semantic anchor nodes respectively.
[0033] Closed logic address space is modulo In a ring-shaped space, the clockwise distance between any two coordinates a and b on the ring is defined as:
[0034]
[0035] For any target coordinate Its corresponding successor responsibility node Defined as node identifiers on the ring not less than the target coordinates. Or, the first node reached after crossing zero, its mathematical expression is:
[0036]
[0037] in, Represents the set of nodes in a closed logical address space; Identifier representing a node; Indicates from the target coordinates Start by moving clockwise through the closed logical address space to reach the node. The distance on the ring; This means returning the ID of the specific node that minimizes the distance, which is... If the distance to the previous node and the distance to the next node are equal, then the next node is chosen.
[0038] For resource entities The semantic coordinates The corresponding semantic anchor nodes are determined as follows:
[0039]
[0040] in, Represents resource entities In the The semantic anchor node corresponding to each semantic coordinate; then the semantic anchor node receives and saves the first system push. Semantic pointers corresponding to each semantic coordinate:
[0041]
[0042] in, Represents resource entities The A semantic coordinate, Represents resource entities The coordinates of the identifier; Indicates the address of the publishing node;
[0043] Semantic pointers are used in subsequent semantic queries when the accumulated results formed by hop-by-hop semantic addressing are insufficient. They allow the query initiating node to locate the corresponding query semantic anchor node based on the query semantic coordinates and obtain candidate resource pointers through the query semantic anchor node to trace the node to which the resource belongs.
[0044] Furthermore, the node semantic profiling module includes: calculating the semantic centroid coordinates of the node based on the semantic vector of the local hosted resources of the current node; and constructing a corresponding set of semantically neighboring candidate nodes based on the semantic centroid coordinates of the node.
[0045] Set the current node Locally hosted k resource entities The set of semantic vectors of k resource entities is = ;
[0046] The node semantic profiling module calculates the current node Single semantic centroid vector:
[0047]
[0048] When node number of semantic centroids =1 indicates a node Using single-semantic centroid profiling, when >1 indicates a node Employing multi-semantic centroid profiling;
[0049] When node When using multi-semantic centroid profiling, the semantic vector set Perform semantic clustering to obtain Semantic clusters:
[0050]
[0051] Then, the semantic centroid vector of each semantic cluster is calculated separately:
[0052]
[0053] Subsequently, using the same L sets of local sensitive projection functions as in the resource semantic representation vector coordinate generation stage, the semantic centroid vector is mapped to L semantic centroid coordinates:
[0054]
[0055] Furthermore, a centroid weight is assigned to each semantic centroid. The semantic centroid weight is used to represent the proportion of the resource set represented by the semantic centroid in the node's locally hosted resources; it is determined in the following manner:
[0056]
[0057] in, This represents the resource entities contained in the t-th semantic cluster. Quantity, k represents the number of nodes Total number of locally hosted resource entities; This indicates that the t-th semantic centroid is located at node . The weights in the semantic profile, and satisfying:
[0058]
[0059] Each candidate node entry in the semantic neighbor candidate node set includes at least:
[0060]
[0061] in, Indicates neighboring candidate nodes Node identifier, Indicates neighboring candidate nodes The access address, Indicates neighboring candidate nodes The number of semantic centroids;
[0062] The set of semantic centroid coordinates is as follows:
[0063]
[0064] The set of centroid weights is:
[0065]
[0066] Among them, when When, candidate node entries degenerate into single centroid entries containing L semantic centroid coordinates; when At that time, the candidate node entries contain the semantic centroid coordinates and their weights corresponding to multiple semantic centroids;
[0067] The node semantic profiling module provides a set of semantically neighboring candidate nodes to the semantic collaborative addressing module, enabling the semantic collaborative addressing module to determine the semantic proximity of the candidate node to the query intent based on the on-ring distance between the semantic centroid coordinates of the candidate node and the query semantic coordinates when selecting the next hop.
[0068] Furthermore, the query intent awareness module employs an adaptive mode switching mechanism, specifically including the following steps:
[0069] ① The query intent awareness module receives the query payload. Then parse the identifier field. The system extracts structured fields and generates standardized identifier key-value pairs; simultaneously, the system also extracts semantic domains. Unstructured text content in Query semantic feature vectors mapped to a high-dimensional continuous space ;
[0070] ② Module calculates the discrete convergence degree. With semantic energy diffusion entropy This is used to quantitatively evaluate the intent focus and traffic splitting benchmark of query load Q;
[0071] Discrete convergence The deterministic calculation process is as follows:
[0072] Completeness retrieval is performed on standardized identifier key-value pairs to construct an existence feature state vector containing N preset metadata baseline dimensions. ; In establishing the state vector If the i-th metadata baseline dimension exists in the input and passes the format validity check, then the state value of that bit is assigned a value. Otherwise, assign a value ;
[0073] Further call to the field convergence weight vector and first-order hard convergence control operator The discrete convergence degree of the identifier is quantified and calculated according to the following formula. :
[0074]
[0075] in, If and only if the state vector When a globally unique identifier or responsibility node identifier field exists that can uniquely lock the one-dimensional coordinates of the closed address space, set =1; If all key positioning elements used to generate precise one-dimensional coordinates are missing, a gated deadlock setting is triggered. =0, at this time Forced convergence to 0;
[0076] Semantic energy diffusion entropy The probabilistic statistical process is as follows: The query semantic feature vector... Input L sets of hash space projection operators, calculate The set of marginal probability distributions under each projection axis Then, the projection space divergence on each projection axis is quantitatively calculated to obtain the semantic energy dispersion entropy. :
[0077]
[0078] The lower the value, the more it tends to converge to the local routing area in the closed address space, which means that the retrieval intent of unstructured text is more explicit;
[0079] The query intent awareness module calculates the discrete convergence of the identifier. and the inverse measure of semantic energy diffusion entropy (1- ), respectively with the preset discrete convergence hard threshold and elastic semantic focus threshold
[0080] Perform interval matching and dynamically distribute the current query load Q to the corresponding routing channel based on different combinations of conditions:
[0081] Condition combination one: When the condition is met When both the identifier domain and semantic domain features are valid, the hybrid cooperative addressing mode is activated.
[0082] Combination of conditions 2: When conditions are met When only semantic domain features are valid, the single semantic approximate routing mode is activated.
[0083] Combination of conditions 3: When conditions are met When the determination is made that only the identifier field feature is valid, the single-label precise addressing mode is activated;
[0084] Combination of conditions 4: When conditions are met If the query characteristics fail to meet the preset routing baseline, the module terminates the initialization routing process and returns an invalid query message to the client.
[0085] Furthermore, during the semantic query process, the semantic collaborative addressing module uses a routing scoring model based on centroid set fusion of semantic distance and path duplication penalty to select the next hop node;
[0086] The node where the current query package is located is denoted as The semantic cooperative addressing module merges the current node's topological routing candidate node set with the semantic neighbor candidate node set to form a candidate node set. ;
[0087] When the query packet reaches the current node At this time, the semantic cooperative addressing module uses the current node's topological routing candidate nodes and semantic neighbor candidate nodes as candidate next-hop nodes; for any candidate next-hop node The semantic collaborative addressing module reads the L semantic centroid coordinates of the candidate node:
[0088]
[0089] When the query initiating node receives the unstructured query payload, it maps the query payload to a query semantic vector. The same LSH projection function as the resource semantic coordinate generation stage is used to generate L semantic coordinates for the query, where L is the preset number of semantic projection groups. Each semantic projection function corresponds to a semantic projection channel, which is used to generate a query semantic coordinate in the closed logical address space.
[0090]
[0091] For the semantic coordinates generated by the l-th semantic projection function, the semantic cooperative addressing module calculates the candidate next-hop node. The on-ring distance between the semantic centroid coordinates of query Q and the query semantic coordinates of query Q:
[0092]
[0093] in, Indicates the candidate next hop node The ring distance between the t-th semantic centroid and the query semantic coordinates generated by query Q in the same semantic projection channel under the L-th semantic projection channel;
[0094] Furthermore, the system employs a temperature coefficient By fusing the ring distances under each semantic projection channel, candidate next-hop nodes are obtained. The fused semantic distance relative to query Q:
[0095]
[0096] in, Indicates the candidate next hop node The fused semantic distance between the query Q and the query Q This is a temperature coefficient used to control the degree of bias in the fused semantic distance towards the minimum distance semantic hash table;
[0097] The semantic cooperative addressing module then calculates the current node. relative to the semantic distance of the L table fusion for query Q And based on candidate next-hop nodes With the current node Based on the fusion of semantic distance differences, candidate next-hop nodes are calculated. Relative semantic progress rate:
[0098]
[0099] in, For smoothing parameters, Indicates starting from the current node Jump to candidate next hop node The relative semantic distance reduction ratio obtained afterwards; when When, it indicates a candidate next-hop node. Relative to the current node Closer to the semantic target of the query;
[0100] The semantic cooperative addressing module determines candidate next-hop nodes based on the visited node records carried in the query packet. Path duplication penalty:
[0101]
[0102] in, This represents the set of visited nodes carried in the query packet;
[0103] The semantic cooperative addressing module calculates candidate next-hop nodes according to the following routing scoring function. Route rating:
[0104]
[0105] in, These are non-negative weighting coefficients; Indicates the candidate next hop node Semantic similarity to query Q; Indicates the candidate next hop node Contribution to effective semantic progress; Indicates a penalty for duplicate paths;
[0106] The semantic cooperative addressing module selects the node with the highest routing score as the next hop node from the candidate next hop nodes; when all candidate next hop nodes satisfy:
[0107]
[0108] The semantic cooperative addressing module determines the current node. It is located in a locally semantically optimal region or a low-yield routing region, triggering an early termination decision or semantic anchor point tracing process for semantic cooperative addressing; among which, This is the minimum effective progress threshold.
[0109] Furthermore, the hop-by-hop retrieval accumulation module adopts a result accumulation mechanism based on cross-node deduplication and incremental benefit evaluation to update the accumulated result set and routing state quantity carried by the query packet during the semantic collaborative addressing process;
[0110] Specifically, when the query mode is a semantic query or a hybrid query, the query packet carries at least a query semantic vector during the semantic cooperative addressing process. Query semantic coordinate set Cumulative result set The set of visited nodes Current jump count h and consecutive low-reward jump count ;
[0111] When the query packet reaches the h-th hop node At that time, based on the query semantic vector and query semantic coordinate set ), at node Obtain the local candidate result set for this hop from the local resource index. ;
[0112] For any candidate result r, the coordinates are identified. Unique identifier for unified resources Determine the unique deduplication key;
[0113] Then, for the local candidate result set of this hop Cumulative result set with the previous hop Perform cross-node deduplication and merging to obtain a candidate merge result set. ;
[0114] The hop-by-hop retrieval cumulative module further calculates the relevance score between the resource result r and the query Q. For the candidate merge result set Sort the results and retain the top K' candidate results to obtain the updated cumulative result set. ;
[0115] The hop-by-hop retrieval accumulation module calculates the incremental revenue for the current hop based on the newly added valid result set; whereby the newly added valid result set for the current hop is defined as:
[0116]
[0117] in, The threshold for the relevance of valid results;
[0118] The incremental profit for this jump is defined as:
[0119]
[0120] in, This indicates the effective contribution of the h-th hop node to the cumulative result set;
[0121] The hop-by-hop retrieval and accumulation module updates the number of consecutive low-yield hops based on the incremental profit of the current hop; the low-yield hop count is incremented by 1 only when the incremental profit of the current hop is lower than the preset incremental profit threshold.
[0122] Then the current node Write to the set of visited nodes And update the current jump count;
[0123] Updated cumulative result set Continuous low-yield jumps The set of visited nodes The current hop count h is written back to the query packet for subsequent semantic collaborative addressing, early termination determination, and semantic anchor tracing.
[0124] Furthermore, the hop-by-hop retrieval accumulation module also includes a semantic cooperative addressing termination determination mechanism. The semantic cooperative addressing termination determination mechanism is used to determine whether to terminate the current semantic cooperative addressing process based on the accumulated result set carried by the query packet, the number of consecutive low-yield hops, the visited node records and the current route hops, and to determine the processing path after termination.
[0125] After the query packet completes the local resource retrieval and cumulative result update at the h-th hop node, the hop-by-hop retrieval and accumulation module reads the addressing status variables carried in the query packet:
[0126]
[0127] in, This represents the cumulative result set formed after the h-th hop; This represents the set of nodes that have been visited. Indicates the current hop count; This indicates the number of consecutive low-yield jumps; This represents the incremental gain of the h-th hop node on the cumulative result set for this hop;
[0128] Then calculate the cumulative satisfaction level after the h-th jump:
[0129]
[0130] in, This represents the cumulative satisfaction level after the h-th hop; Indicates the preset target number of output results. Represents the cumulative result set The number of resource results in the middle. Represents the cumulative result set Compared to the overall relevance quality of query Q, and The weight coefficients are non-negative and satisfy the following conditions:
[0131]
[0132] Among them, overall relevance quality Determine as follows:
[0133]
[0134] in, Indicates from the cumulative result set The top-ranked candidates were selected based on their relevance scores. A collection of resource results;
[0135] when hour, = ; This represents the relevance score between the resource result r and the query Q;
[0136] when When empty, set ;
[0137] The hop-by-hop retrieval accumulation module further calculates the semantic route decay after the h-th hop based on the number of consecutive low-yield hops, visited node records, and the current route hop count.
[0138]
[0139] in, This represents the semantic route decay after the h-th hop; This indicates the maximum number of consecutive low-return jumps that can be preset. This indicates the preset maximum number of route hops. Indicates the path repetition intensity. , , The weight coefficients are non-negative and satisfy the following conditions:
[0140]
[0141] Among them, path repetition intensity Based on the set of visited nodes Determining the degree of node repetition:
[0142]
[0143] in, This indicates that the records are from visited nodes. The set of nodes obtained after deduplication This indicates the number of visited nodes after deduplication;
[0144] Semantic route decay This is used to characterize the degree of benefit decay when continuing to perform semantic cooperative addressing; among them, the higher the number of consecutive low-benefit hops, the higher the path repetition intensity, and the closer the current route hop count is to the preset maximum route hop count, the higher the semantic route decay degree. The higher;
[0145] The hop-by-hop retrieval accumulation module further sets hard termination boundaries; when the number of consecutive low-yield hops reaches the preset upper limit, or the current route hop count reaches the preset maximum route hop count, or the node where the current query packet is located already exists in the record of the visited nodes before it reaches the current node, it is determined that the current semantic cooperative addressing triggers a hard termination boundary. Then, based on the cumulative result satisfaction, semantic route decay, and hard termination boundary, the subsequent processing path of semantic cooperative addressing is determined:
[0146] When the cumulative result satisfaction reaches the preset cumulative result satisfaction threshold, the current semantic collaborative addressing process is stopped, and the cumulative result set is input into the result fusion output module, which then organizes the results to generate semantic query results.
[0147] When the cumulative result satisfaction does not reach the preset cumulative result satisfaction threshold, and the semantic route decay reaches the preset semantic route decay threshold, or a hard termination boundary is triggered, the current semantic cooperative addressing is determined to enter the route decay termination state, the current semantic cooperative addressing process is stopped, and the query packet state is switched to the semantic anchor tracing state.
[0148] When the cumulative result satisfaction does not reach the preset cumulative result satisfaction threshold, and the semantic route decay does not reach the preset semantic route decay threshold, and no hard termination boundary is triggered, it is determined that the current semantic cooperative addressing has not yet reached the termination condition. The updated query packet state is written back to the query packet, and the semantic cooperative addressing module is called again to select the next hop node.
[0149] Furthermore, the semantic anchor tracing module executes a tracing mechanism based on result gaps and tracing budget constraints, which is used to supplement the discovery of candidate resources when semantic cooperative addressing terminates and the accumulated result set does not meet the preset output requirements;
[0150] The semantic anchor tracing module receives the query packet transmitted by the hop-by-hop retrieval accumulation module. The query packet carries at least the query semantic coordinate set, the accumulated result set, the visited node record, and the current route hop count.
[0151] The semantic anchor tracing module reads the cumulative result set formed when semantic cooperative addressing terminates. And calculate the number of result gaps, i.e. the maximum number of semantic pointers, based on the preset target result number K:
[0152]
[0153] The semantic anchor tracing module locates the corresponding query semantic anchor node in the closed logical address space based on the query semantic coordinate set, and sends a backup tracing request to the query semantic anchor node. After receiving the backup tracing request, the query semantic anchor node reads candidate semantic pointers from the locally stored semantic pointer directory, and calculates the tracing score of the candidate semantic pointers based on the on-ring distance between the resource semantic coordinates and the query semantic coordinates in the candidate semantic pointers.
[0154]
[0155] in, Represents a candidate semantic pointer. This represents the semantic coordinates of the resource corresponding to the candidate semantic pointer. Indicates the semantic coordinates of the query. Represents the distance on the ring within the closed logical address space. Represents the coordinate scale of a closed logical address space;
[0156] The query semantic anchor node sorts the candidate semantic pointers according to the traceability score and includes the traceability budget. Within the range, a set of candidate semantic pointers is returned. When the semantic anchor tracing module receives the set of candidate semantic pointers, it generates a deduplication identifier based on the resource identifier coordinates, resource unique identifier, or resource node access location information in the candidate semantic pointers, and matches the candidate semantic pointers with the existing resource results in the current accumulated result set to remove duplicate semantic pointers.
[0157] The semantic anchor tracing module initiates a secondary resource discovery request to the corresponding resource node only for the candidate semantic pointers retained after deduplication, obtains supplementary resource results, and then merges the supplementary resource results with the current cumulative result set before inputting them into the result fusion output module.
[0158] Furthermore, the result fusion output module performs unified fusion output on the identifier addressing result, the cumulative result set formed by semantic collaborative addressing, and the traceability result set formed by the semantic anchor traceability module;
[0159] The result fusion output module generates a result deduplication key based on the resource identifier coordinates and resource unique identifier information, and merges and deduplicates resource results from different sources according to the result deduplication key;
[0160] The result fusion output module merges and organizes the deduplicated resource results to generate the final resource discovery results.
[0161] Compared with the prior art, the beneficial effects of the present invention are as follows: the present invention realizes the collaborative processing of accurate addressing and semantic approximation discovery in a decentralized environment, and improves the accuracy, recall capability and communication efficiency of distributed resource discovery without centralized storage of original resource data and without broadcasting across the entire network. Attached Figure Description
[0162] Figure 1 This is a system overall structure diagram of the present invention.
[0163] Figure 2 This is a system flowchart of the present invention.
[0164] Figure 3 This is a comparison chart of recall rate and communication overhead for each scheme in the embodiments.
[0165] Figure 4 This is a comparison chart of semantic retrieval metrics for each scheme in the embodiments. Detailed Implementation
[0166] The technical solutions in the implementation of the present invention will be clearly and completely described below with reference to the accompanying drawings in the examples of the present invention.
[0167] like Figure 1 and 2 As shown, this invention provides a decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantic-aware routing. The system operates within a closed logical addressing space composed of multiple distributed nodes, and includes the following steps:
[0168] Step 1: Dual-coordinate index construction and semantic anchor registration.
[0169] During the system initialization phase, the system presets the coordinate scale of the closed logical address space to be... Each distributed node generates a node identifier based on its registration information, network address, or unique identifier. This node identifier is then mapped to the closed logical address space. Node access address This indicates the network entry point for the node to receive query packets, registration requests, and tracing requests, including information such as its IP address. Each node maintains a set of candidate nodes for topology routing based on its node identifier.
[0170] The node owning the resource first obtains the locally hosted resource entity. resource content And generate resource semantic representation vectors in a continuous high-dimensional space through semantic mapping. Simultaneously acquire resource entities Resource Identifier Primitives The resource identifier primitive consists of a unique resource identifier. Resource types After some processing, the following was generated by splicing: =
[0171] Next, based on the resource identifier primitive Through a preset hash mapping function, in a modulus of Generate one-dimensional scalar form identifier coordinates in the circular closed logical address space. Used to uniquely lock the resource entity Identify the responsible node in the distributed network. Then, generate resource entities through L sets of semantic projection functions. L semantic coordinates, where the first... The set of projection vectors is represented as:
[0172]
[0173] in, , Indicates the first The first group semantic projection function One projection vector; The number of bits in the binary code corresponding to each semantic coordinate.
[0174] For resource entities In the The i-th semantic code under the group projection function is generated as follows:
[0175]
[0176] in, Represents resource entities semantic vectors In the The binarization result of the i-th projection direction. The binarization result is used to characterize the semantic distinction state of the resource semantic representation vector in the corresponding projection direction.
[0177] Then, the first The group projection function obtained Bitwise codes are concatenated into integer coordinates to obtain the resource entity. The Semantic coordinates:
[0178]
[0179] Ultimately, this is integrated into a semantic coordinate set:
[0180]
[0181] Each semantic coordinate falls independently into the closed logical address space.
[0182] The node to which the resource belongs is based on the resource identifier coordinates. Locate the responsible node for identification and send an identification index registration request to that node.
[0183] When a resource entity joins the network, it is simultaneously assigned both identifier coordinates and multi-dimensional semantic coordinates, achieving a unified approach to deterministic localization and approximate discovery. Since both types of coordinates share the same closed logical address space, the complexity of converting between heterogeneous index spaces is eliminated. Furthermore, multiple semantic coordinates represent resources from multiple angles through different semantic projection channels, significantly reducing the risk of missed detections due to insufficient single semantic representation. Under this mechanism, the original resource data remains local; the system only maintains indexes and pointers in a distributed manner, achieving efficient cross-node resource discovery while preserving node autonomy.
[0184] Step 2, Semantic Anchor Registration
[0185] The system is based on resource entities L semantic coordinates In the closed logical address space, determine L corresponding semantic anchor nodes respectively; and send the corresponding semantic pointers to the semantic anchor nodes respectively.
[0186] The clockwise distance between any two coordinates a and b on the ring is defined as:
[0187]
[0188] For any target coordinate Its corresponding successor responsibility node Defined as the node identifier on the ring being no less than the target coordinates. Or, the first node reached after crossing zero, its mathematical expression is:
[0189]
[0190] in, Represents the set of nodes in a closed logical address space; Identifier representing a node; Indicates from the target coordinates Start by moving clockwise through the closed logical address space to reach the node. The distance on the ring; This means returning the ID of the specific node that minimizes the distance, which is... If the distance to the previous node and the distance to the next node are equal, then the next node is chosen.
[0191] For resource entities The semantic coordinates The semantic anchor registration determines the corresponding semantic anchor node in the following manner:
[0192]
[0193] in, Represents resource entities In the The semantic anchor node corresponding to each semantic coordinate; then the semantic anchor node receives and saves the first system push. Semantic pointers corresponding to each semantic coordinate:
[0194]
[0195] in, Represents resource entities The A semantic coordinate, Represents resource entities The coordinates of the identifier Indicates resource description; Indicates the address of the publishing node.
[0196] Semantic pointers are used in subsequent semantic queries when the accumulated results formed by hop-by-hop semantic addressing are insufficient. They allow the query initiating node to locate the corresponding query semantic anchor node based on the query semantic coordinates and obtain candidate resource pointers through the query semantic anchor node to trace the node to which the resource belongs.
[0197] The multidimensional semantic coordinates of resources are registered to corresponding semantic anchor nodes, constructing a multi-entry backup retrieval path. When the cooperative addressing results are insufficient, the system accesses the anchor point according to the query coordinates to obtain nearby lightweight pointer information. This mechanism only maintains pointers rather than the original data, fully preserving the local node's absolute control over the resource; at the same time, the anchor point is only activated as a backup supplementary entry point, effectively avoiding large-scale broadcasting in the initial stage and reducing the overall communication overhead of the distributed network.
[0198] Step 3: Node semantic profile construction.
[0199] Each distributed node generates a node semantic profile based on the semantic representation of the locally hosted resources of the distributed node, and converts the node semantic profile into node semantic centroid coordinates that are in the same address space as the resource semantic coordinates. The node semantic centroid coordinates are used to construct a set of semantically neighboring candidate nodes.
[0200] In the specific profile construction, assume that any distributed node locally hosts k resource entities. The corresponding set of semantic vectors is represented as The semantic profiling module first calculates the single semantic centroid vector of the locally hosted resource. The calculation formula is as follows:
[0201]
[0202] Then, based on the semantic vector set The degree of topic dispersion determines the nodes number of semantic centroids ,when =1 indicates that node x uses a single semantic centroid profile. >1 indicates a node Multi-semantic centroid profiling is employed.
[0203] When node When using multi-semantic centroid profiling, the node semantic profiling module processes the multi-modal semantic vector set. Perform semantic clustering to obtain Semantic clusters:
[0204]
[0205] Then, the semantic centroid vector of each semantic cluster is calculated separately:
[0206]
[0207] Subsequently, the node semantic profiling module adopts the same L-group locality-sensitive hash projection function as in step 1:
[0208] ,
[0209] Map the semantic centroid vectors to a set of semantic centroid coordinates containing L integers:
[0210]
[0211] Each semantic centroid is dynamically assigned a centroid weight reflecting resource richness, calculated using the following formula:
[0212]
[0213] The set of centroid weights is:
[0214]
[0215] Finally, the current node formats and encapsulates its own generated node identifier, number of semantic centroids, L semantic centroid coordinates corresponding to each centroid, and corresponding centroid weights to construct a complete semantic neighbor candidate node entry. Through network discovery, each node declares its own candidate node entry to its neighboring nodes in a bounded manner, thereby maintaining and updating its unique set of semantic neighbor candidate nodes locally, which is used to provide semantic proximity comparison in subsequent addressing.
[0216] Node semantic profiles are constructed based on local resource distribution, and node topics are represented by single or multiple semantic centroids to prevent semantic shifts caused by averaging multiple topics. Since the node centroid and query payload are in the same closed logical address space, the routing module can directly calculate the spatial distance to assess semantic proximity, thereby guiding query packets to be forwarded to highly relevant nodes first, improving the routing directionality of decentralized addressing.
[0217] Step 4: Query intent awareness and mode switching.
[0218] When a user initiates a query, the query intent awareness module receives the query payload Q. In this embodiment, the query payload can be represented as:
[0219]
[0220] in, This represents the structured identifier field in the query. This represents the unstructured semantic domain in the query.
[0221] The query intent awareness module first parses the identifier field. The system extracts relevant fields that can be used to generate identifier coordinates and generates standardized identifier key-value pairs. Simultaneously, the system will define the semantic domain... The natural language text, keywords, or functional descriptions in the query are mapped to the query semantic vector.
[0222] Subsequently, the query intent awareness module calculates the discrete convergence degree of the identifier. and semantic diffusion entropy It is used to determine whether the query payload is suitable for entering the exact addressing, semantic addressing, or hybrid addressing channel.
[0223] First, the system constructs an existence feature state vector containing N preset metadata baseline dimensions:
[0224]
[0225] Specifically, when the i-th metadata baseline dimension exists in the input and passes the format validity check, the following is set: Otherwise set to .
[0226] The system calls the preset field convergence weight vector:
[0227]
[0228] And set a first-order hard convergence control operator. When a standardized identifier key-value pair contains a globally unique identifier that can uniquely lock the one-dimensional coordinates of a closed logical address space, or a key positioning field that can generate precise identifier coordinates, then set... =1; Set when all key positioning fields used to generate precise identifier coordinates are missing. =0.
[0229] The discrete convergence is calculated using the following formula:
[0230]
[0231] in, The higher the value, the more likely the identifier field in the query is to converge to a specific identifier addressing target.
[0232] For the semantic domain, the system will query the semantic vector. Input L sets of semantic projection functions to obtain the set of query semantic coordinates. Simultaneously, based on the distribution of the query semantic vector in each projection space, the set of marginal probability distributions is calculated. And calculate the semantic diffusion entropy:
[0233]
[0234] To facilitate threshold determination, in specific implementation, it can be... Normalize to the interval [0,1]. The lower the value, the more concentrated the query semantics, meaning the more suitable the query semantics are for entering the semantic routing channel.
[0235] The query intent awareness module will With preset discrete convergence threshold Compare, and put 1− With preset semantic focus threshold The comparison yielded the following mode switching results:
[0236] When satisfied When both the identifier field and the semantic field are valid, the hybrid cooperative addressing mode is activated;
[0237] When satisfied When the condition is met, it is determined that only the semantic domain is valid, and the single semantic approximate routing mode is activated.
[0238] When satisfied When the condition is met, only the identifier field is determined to be valid, and the single-label precise addressing mode is activated.
[0239] When satisfied If the system determines that the query load does not meet the routing initialization conditions, it terminates the query initialization process and returns an invalid query message to the client.
[0240] Through the aforementioned query intent perception process, the system can adaptively select precise query, semantic query, or hybrid query path based on the validity of structured identifier information and unstructured semantic information in the query payload. For queries containing only valid identifier fields, the system avoids initiating unnecessary semantic routing; for queries containing only semantic descriptions, the system can directly enter semantic collaborative addressing; for queries containing both identifier and semantic fields, the system can initiate identifier addressing path and semantic addressing path in parallel.
[0241] Therefore, the query execution path can be matched with the information composition of the query payload, avoiding misleading semantic queries into precise addressing paths or vice versa. Simultaneously, the hybrid query mode can balance deterministic hits and semantic expansion discovery, improving the completeness of resource retrieval in complex query scenarios.
[0242] Step 5: In semantic or mixed queries, the query packet reaches the current node. Then, the semantic cooperative addressing module reads the current node's set of topological routing candidate nodes and set of semantic neighbor candidate nodes, and merges them to form a candidate next-hop set. .
[0243] For any candidate next-hop node The system reads its semantic centroid coordinate set. And the centroid weight set. For the l-th semantic coordinate group and the t-th semantic centroid, calculate the candidate node. The distance on the ring between the query Q and the query Q:
[0244]
[0245] in, It represents the clockwise distance in a closed logical address space.
[0246] Then, the system uses a temperature coefficient The distances between multiple semantic tables and multiple semantic centroids are fused to obtain candidate nodes. The fused semantic distance relative to query Q:
[0247]
[0248] in, Used to control the degree to which the fusion distance is biased towards smaller semantic distances; The smaller,
[0249] Indicates candidate nodes The closer it is to the query target semantically.
[0250] The semantic cooperative addressing module further calculates the current node. The fused semantic distance relative to query Q And calculate candidate next-hop nodes. Relative semantic progress rate:
[0251]
[0252] in, For smoothing parameters. When When, it indicates a jump from the current node to a candidate node. Afterwards, the query packet can converge toward the semantic target.
[0253] To prevent query packets from repeatedly accessing nodes, the system uses the set of already visited nodes carried in the query packet as a reference. Calculate path duplication penalty:
[0254]
[0255] The semantic cooperative addressing module calculates candidate next-hop nodes according to the following routing scoring function. Route rating:
[0256]
[0257] in, The non-negative weight coefficient is used by the system to select the node with the highest routing score from the candidate next-hop set as the next-hop node.
[0258] When all candidate next-hop nodes satisfy When the system determines that the current node is in a locally semantically optimal region or a low-yield routing region, then... This is the minimum effective progress threshold. At this point, the semantic cooperative addressing module triggers an early termination decision or a semantic anchor tracing process.
[0259] Step 6, hop-by-hop retrieval accumulation. When the query packet reaches the h-th hop node... At that time, the hop-by-hop retrieval accumulation module reads the query semantic vector from the query packet. Query semantic coordinate set Previous hop cumulative result set The set of visited nodes Current jump count h and consecutive low-reward jump count .
[0260] Current node Perform a local search in the local resource index to obtain the local candidate result set for this hop:
[0261]
[0262] Local retrieval can be achieved based on the similarity of resource semantic representation vectors, the proximity of resource semantic coordinates, resource type filtering, keyword matching, or a combination thereof.
[0263] For any resource result r, the system identifies the coordinates. and Uniform Resource Identifier It determines a unique deduplication key.
[0264] Then, the hop-by-hop retrieval accumulation module performs local candidate result set analysis on the current hop. Cumulative result set with the previous hop Perform deduplication and merging to obtain a candidate merge result set.
[0265] Then, the system calculates the relevance score between the resource result r and the query Q. Sort the candidate merge result set and retain the top K′ candidate results to obtain the updated cumulative result set.
[0266] To determine whether the current node makes a valid contribution to the search results, the system calculates the set of newly added valid results for this hop:
[0267]
[0268] in, The threshold for the relevance of valid results.
[0269] The incremental profit for this jump is defined as:
[0270]
[0271] in, This indicates the effective contribution of the h-th hop node to the cumulative result set.
[0272] The hop-by-hop retrieval and accumulation module updates the number of consecutive low-yield hops based on the incremental revenue of the current hop; the low-yield hop count is incremented by 1 only when the incremental revenue of the current hop is lower than the preset incremental revenue threshold.
[0273] The hop-by-hop retrieval cumulative module then retrieves the current node. Write to the set of visited nodes And update the current jump count.
[0274] Updated cumulative result set Continuous low-yield jumps The set of visited nodes The current hop count h is written back to the query packet for subsequent semantic routing selection, termination determination, and semantic anchor tracing.
[0275] Through a hop-by-hop retrieval accumulation mechanism, the query packet does not only perform retrieval at the final node during semantic collaborative addressing, but also triggers local resource retrieval at each hop node, and merges the current hop result with the historical accumulated results to remove duplicates. This fully utilizes the local resources of the nodes traversed along the query path, improving the phased recall capability during the semantic query process.
[0276] Furthermore, the system uses a unified resource identifier or resource identifier coordinates to establish deduplication criteria, reducing duplicate results when different nodes return the same resource. By sorting by relevance and retaining preceding candidate results, the system can control the size of the result set carried by the query packet, preventing the result set from expanding indefinitely with the number of hops. Through the incremental benefit evaluation of this hop, the system can determine whether the current hop has generated valid new results, providing a basis for whether to continue routing or initiate tracing.
[0277] Step 7: After the query packet completes the local retrieval and cumulative result update of the h-th hop node, the system calculates the cumulative result satisfaction and semantic route decay to determine whether to continue semantic cooperative addressing.
[0278] First, the system calculates the cumulative result satisfaction after the h-th hop:
[0279]
[0280] in, This represents the cumulative satisfaction level after the h-th hop; Indicates the preset target number of output results. Represents the cumulative result set The number of resource results in the middle. Represents the cumulative result set Compared to the overall relevance quality of query Q, and The weight coefficients are non-negative and satisfy the following conditions:
[0281] Overall relevance quality Determine as follows:
[0282]
[0283] in, Indicates from the cumulative result set The top-ranked candidates were selected based on their relevance scores. A collection of resource results.
[0284] Then, the system calculates the path duplication intensity based on the number of consecutive low-yield hops, visited node records, and the current route hop count:
[0285]
[0286] in, Represents the set of visited nodes The set of nodes obtained after deduplication. This is the set of nodes obtained when no duplicate accesses occur in the query package. The ratio of h approaches 1. The value approaches 0; when a query packet accesses an existing node multiple times, Increase.
[0287] Further calculate semantic route decay:
[0288]
[0289] in, This indicates a preset upper limit on the number of consecutive low-yield jumps. This indicates the preset maximum number of route hops. , , The weights are non-negative and satisfy the following conditions:
[0290]
[0291] Semantic route decay The higher the value, the lower the benefit of continuing to perform semantic cooperative addressing.
[0292] Therefore, the system further sets a hard termination boundary. When the number of consecutive low-yield hops reaches a preset upper limit, or the current route hop count reaches a preset maximum route hop count, or the node where the current query packet resides already exists in the record of visited nodes before it reaches the current node, it is determined that the current semantic cooperative addressing triggers a hard termination boundary. Then, based on the cumulative result satisfaction, semantic route decay, and the hard termination boundary, the subsequent processing path of semantic cooperative addressing is determined:
[0293] When the cumulative result satisfaction reaches a preset cumulative result satisfaction threshold, the current semantic collaborative addressing process is stopped, and the cumulative result set is input into the result fusion output module, which then organizes the results to generate semantic query results.
[0294] When the cumulative result satisfaction does not reach the preset cumulative result satisfaction threshold, and the semantic route decay reaches the preset semantic route decay threshold, or when the hard termination boundary is triggered, it is determined that the current semantic cooperative addressing has entered the route decay termination state, the current semantic cooperative addressing process is stopped, and the query packet state is switched to the semantic anchor tracing state.
[0295] When the cumulative result satisfaction does not reach the preset cumulative result satisfaction threshold, and the semantic route decay does not reach the preset semantic route decay threshold, and the hard termination boundary is not triggered, it is determined that the current semantic cooperative addressing has not yet reached the termination condition. The updated query packet status is written back to the query packet, and the semantic cooperative addressing module is called again to select the next hop node.
[0296] Step 8: When semantic collaborative addressing terminates and the accumulated result set still does not meet the preset output requirements, the semantic anchor tracing module starts the controlled backup tracing mechanism.
[0297] The semantic anchor tracing module reads the set of semantic coordinates, cumulative result set, visited node records, and current route hop count from the query package.
[0298] First, based on the target number of output results K and the current cumulative number of results... The calculated gap number is the maximum number of semantic pointers.
[0299]
[0300] The semantic anchor tracing module locates the corresponding query semantic anchor node based on the query semantic coordinate set. Subsequently, the semantic anchor tracing module sends a backup tracing request to the aforementioned query semantic anchor node. The query semantic anchor node reads a set of candidate semantic pointers from its local semantic pointer directory and calculates a tracing score based on the ring distance between the resource semantic coordinates and the query semantic coordinates in the candidate semantic pointers.
[0301]
[0302] in, Represents a candidate semantic pointer. This represents the semantic coordinates of the resource corresponding to the candidate semantic pointer. Indicates the semantic coordinates of the query. Represents the distance on the ring within the closed logical address space. It represents the coordinate scale of the closed logical address space.
[0303] The query semantic anchor node sorts the candidate semantic pointers according to the traceability score and returns a set of candidate semantic pointers within the traceability range.
[0304] After receiving the candidate semantic pointer set, the semantic anchor tracing module generates a deduplication key based on the resource identifier coordinates, resource unified identifier, and resource-owning node access address in the candidate semantic pointers. Then, the system matches the candidate semantic pointers with existing resource results in the current accumulated result set, eliminating duplicate semantic pointers. For the candidate semantic pointers retained after deduplication, the semantic anchor tracing module initiates a secondary resource discovery request to the corresponding resource-owning node. The resource-owning node verifies and returns supplementary resource results in its local resource index based on the query semantic vector, query semantic coordinates, resource identifier coordinates, or resource unified identifier in the request.
[0305] The input result is merged with the current cumulative result set and then output to the fusion module.
[0306] The semantic anchor tracing mechanism serves as a controlled supplementary measure when results are insufficient. The system uses query coordinates to access semantic anchor nodes to obtain pointers to nearby resources, and then precisely initiates a tracing request to the node to which the resource belongs. This process is constrained by both semantic coordinates and tracing budget, effectively avoiding blind broadcasting. The tracing module ultimately directly penetrates the pointers to obtain the entity supplementary resource and sends it to the fusion output module, achieving resource completion.
[0307] Step 9: The result fusion output module receives one or more of the following result sets:
[0308] a) Represents the set of precisely identified addressing results.
[0309] b) Represents the cumulative result set formed by semantic cooperative addressing;
[0310] c) represents the set of traceability results formed by semantic anchor tracing.
[0311] The system first generates a deduplication key for results from all sources, and then merges and deduplicates results from multiple sources. When the same resource result comes from multiple paths in precise addressing, semantic collaborative addressing, and semantic anchor tracing, the system retains only one resource result record, while also retaining its source tag, relevance score, resource summary, resource node address, and resource identifier coordinates.
[0312] Finally, the system generates the final resource discovery results according to the query pattern:
[0313] When the query mode is exact query, the exact identifier addressing result will be output first;
[0314] When the query mode is semantic query, the output includes semantic cumulative results and traceability supplementary results;
[0315] When the query mode is a mixed query, the resource results that are precisely addressed are retained first, and then the semantic accumulation results and traceability results are supplemented.
[0316] Through the result fusion output module, the system can unify results from different query paths into a single output structure. For records pointing to the same resource entity in the exact addressing results, semantic accumulation results, and tracing results, the system performs deduplication and merging based on the resource's unique identifier, resource identifier coordinates, or the address of the node to which the resource belongs, thus avoiding duplicate resources in the final result.
[0317] When the same resource is hit by multiple paths simultaneously, the system can retain its multi-source hit information to characterize that the resource not only meets the precise identification conditions but also has relevance to the query semantics. Therefore, the final output not only includes the resource discovery result itself but also reflects the source path and hit basis of the resource result, improving the stability and interpretability of the output.
[0318] After completing the above process, to verify the resource discovery effect of the present invention, a comparative experiment was conducted under the same test conditions. The experimental results are as follows: Figure 3 and Figure 4 As shown.
[0319] The experimental results above show that, through the collaborative operation of dual-coordinate indexing, node semantic profiling, semantic collaborative addressing, hop-by-hop retrieval accumulation, and semantic anchor tracing, this invention improves the recall capability of semantically related resources and the quality of result ranking compared to DHT schemes that rely solely on identifier addressing and ordinary semantic routing schemes; compared to flooding broadcast schemes, it reduces the communication overhead between nodes, thereby achieving a balance between resource discovery effectiveness and communication cost.
Claims
1. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing, characterized in that, It operates within a closed logical addressing space composed of multiple distributed nodes, including a dual-coordinate index construction module, a semantic anchor registration module, a node semantic profiling module, a query intent awareness module, a semantic collaborative addressing module, a hop-by-hop retrieval accumulation module, a semantic anchor tracing module, and a result fusion output module; each module operates collaboratively according to the following process: (1) The dual-coordinate index construction module obtains the resource identifier and resource semantic representation of the resource entity, performs deterministic mapping on the resource identifier, generates identifier coordinates for precise addressing, and maps the resource semantic representation through multiple sets of semantic projection functions to generate multiple semantic coordinates for semantic approximation discovery, and maps the identifier coordinates and multiple semantic coordinates to the same closed logical addressing space. (2) The semantic anchor registration module determines the corresponding semantic anchor node based on multiple semantic coordinates, and registers a semantic pointer containing the resource semantic coordinates, resource identifier coordinates and the address of the node to which the resource belongs to the semantic anchor node for subsequent semantic anchor tracing. (3) The node semantic profiling module generates one or more semantic centroids and corresponding semantic centroid coordinates of a node based on the semantic representation of the local hosted resources of each distributed node. It establishes a set of semantic neighbor candidate nodes based on the semantic centroid coordinates and centroid weights, and provides the set of semantic neighbor candidate nodes to the semantic collaborative addressing module. (4) The query intent perception module parses the precise identifier field and unstructured semantic field in the query payload, determines the precise query mode, semantic query mode or mixed query mode based on the effectiveness of the precise identifier field and the semantic focus of the unstructured semantic field, and starts the corresponding query path. (5) In precise query mode, the identification responsibility node is located and precise addressing results are obtained based on the precise identification field; in semantic query mode, the semantic collaborative addressing module generates query semantic coordinates based on the unstructured semantic field, merges the topology routing candidate nodes of the current node with the semantic neighbor candidate nodes, and selects the next hop node based on the proximity of the candidate node to the query semantic coordinates. The hop-by-hop retrieval accumulation module performs local resource retrieval and result deduplication accumulation at the nodes traversed by the query packet; in mixed query mode, precise query and semantic query are executed simultaneously; when the accumulated result meets the preset result requirements, semantic collaborative addressing stops; when semantic collaborative addressing terminates and the accumulated result does not meet the preset result requirements, the semantic anchor tracing module locates the query semantic anchor node based on the query semantic coordinates, obtains the semantic pointer, and initiates a tracing request to the node to which the corresponding resource belongs. (6) The result fusion output module deduplicatizes the precise addressing results, hop-by-hop cumulative results and tracing results to generate the final resource discovery results.
2. The decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing according to claim 1, characterized in that, The dual-coordinate index construction module includes: First, the system obtains the resource entity. resource content And generate resource semantic representation vectors through semantic mapping. ; Next, the system obtains the resource entity. Resource Identifier Primitives Resource identification primitives consist of unique identifiers for each resource. Resource types After fixed field order, unified encoding, field cleaning, null value placeholders, delimiter concatenation, and deterministic serialization, the data is then concatenated together. = ; Based on the above resource identifier primitives The identifier coordinates are generated in the closed logical address space through a hash mapping function. Then, based on semantic resource representation Resource entities are generated through semantic projection functions. The L semantic coordinates are represented by the set of projection vectors as follows: in, , Indicates the first The first group semantic projection function One projection vector; The number of bits in the binary code corresponding to each semantic coordinate; For resource entities In the The i-th semantic code under the group projection function is generated as follows: in, Represents resource entities semantic vectors In the The binarization result of the i-th projection direction; the binarization result is used to characterize the semantic distinction state of the resource semantic representation vector in the corresponding projection direction; then, the i-th... The group projection function obtained Bitwise codes are concatenated into integer coordinates to obtain the resource entity. The Semantic coordinates: Multiple semantic coordinates are combined to obtain a semantic coordinate set: Each semantic coordinate All fall into the closed logical address space; Identify coordinates Used to locate resource entities Identifying the responsible node, a set of semantic coordinates Used to locate resource entities Multiple semantic anchor nodes, and the dual-coordinate index data does not contain the original resource data; In addition, the L-group projection functions are also configured to allow the query intent awareness module to generate query semantic coordinates. And the node semantic centroid coordinates generated by the node semantic profiling module.
3. The decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing according to claim 1, characterized in that, The semantic anchor registration module includes: based on resource entities L semantic coordinates In the closed logical address space, determine L corresponding semantic anchor nodes respectively; and send the corresponding semantic pointers to the semantic anchor nodes respectively. Closed logic address space is modulo In a ring-shaped space, the clockwise distance between any two coordinates a and b on the ring is defined as: For any target coordinate Its corresponding successor responsibility node Defined as node identifiers on the ring not less than the target coordinates. Or, the first node reached after crossing zero, its mathematical expression is: in, Represents the set of nodes in a closed logical address space; Identifier representing a node; Indicates from the target coordinates Start by moving clockwise through the closed logical address space to reach the node. The distance on the ring; This means returning the ID of the specific node that minimizes the distance, which is... If the distance to the previous node and the distance to the next node are equal, then the next node is chosen. For resource entities The semantic coordinates The corresponding semantic anchor nodes are determined as follows: in, Represents resource entities In the The semantic anchor node corresponding to each semantic coordinate; then the semantic anchor node receives and saves the first system push. Semantic pointers corresponding to each semantic coordinate: in, Represents resource entities The A semantic coordinate, Represents resource entities The coordinates of the identifier; Indicates the address of the publishing node; Semantic pointers are used in subsequent semantic queries when the accumulated results formed by hop-by-hop semantic addressing are insufficient. They allow the query initiating node to locate the corresponding query semantic anchor node based on the query semantic coordinates and obtain candidate resource pointers through the query semantic anchor node to trace the node to which the resource belongs.
4. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing as described in claim 1, characterized in that, The node semantic profiling module includes: calculating the semantic centroid coordinates of the node based on the semantic vector of the local hosted resources of the current node; and constructing a corresponding set of semantically neighboring candidate nodes based on the semantic centroid coordinates of the node. Set the current node Locally hosted k resource entities The set of semantic vectors of k resource entities is = ; The node semantic profiling module calculates the current node Single semantic centroid vector: When node number of semantic centroids =1 indicates a node Using single-semantic centroid profiling, when >1 indicates a node Employing multi-semantic centroid profiling; When node When using multi-semantic centroid profiling, the semantic vector set Perform semantic clustering to obtain Semantic clusters: Then, the semantic centroid vector of each semantic cluster is calculated separately: Subsequently, using the same L sets of local sensitive projection functions as in the resource semantic representation vector coordinate generation stage, the semantic centroid vector is mapped to L semantic centroid coordinates: Furthermore, a centroid weight is assigned to each semantic centroid. The semantic centroid weight is used to represent the proportion of the resource set represented by the semantic centroid in the node's locally hosted resources; it is determined in the following manner: in, This represents the resource entities contained in the t-th semantic cluster. Quantity, k represents the number of nodes Total number of locally hosted resource entities; This indicates that the t-th semantic centroid is located at node . The weights in the semantic profile, and satisfying: Each candidate node entry in the semantic neighbor candidate node set includes at least: in, Indicates neighboring candidate nodes Node identifier, Indicates neighboring candidate nodes The access address, Indicates neighboring candidate nodes The number of semantic centroids; The set of semantic centroid coordinates is as follows: The set of centroid weights is: Among them, when When, candidate node entries degenerate into single centroid entries containing L semantic centroid coordinates; when At that time, the candidate node entries contain the semantic centroid coordinates and their weights corresponding to multiple semantic centroids; The node semantic profiling module provides a set of semantically neighboring candidate nodes to the semantic collaborative addressing module, enabling the semantic collaborative addressing module to determine the semantic proximity of the candidate node to the query intent based on the on-ring distance between the semantic centroid coordinates of the candidate node and the query semantic coordinates when selecting the next hop.
5. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing as described in claim 1, characterized in that, The query intent awareness module employs an adaptive mode switching mechanism, specifically including the following steps: ① The query intent awareness module receives the query payload. Then parse the identifier field. The system extracts structured fields and generates standardized identifier key-value pairs; simultaneously, the system also extracts semantic domains. Unstructured text content in Query semantic feature vectors mapped to a high-dimensional continuous space ; ② Module calculates the discrete convergence degree. With semantic energy diffusion entropy This is used to quantitatively evaluate the intent focus and traffic splitting benchmark of query load Q; Discrete convergence The deterministic calculation process is as follows: Completeness retrieval is performed on standardized identifier key-value pairs to construct an existence feature state vector containing N preset metadata baseline dimensions. ; In establishing the state vector If the i-th metadata baseline dimension exists in the input and passes the format validity check, then the state value of that bit is assigned a value. Otherwise, assign a value ; Further call to the field convergence weight vector and first-order hard convergence control operator The discrete convergence degree of the identifier is quantified and calculated according to the following formula. : in, If and only if the state vector When a globally unique identifier or responsibility node identifier field exists that can uniquely lock the one-dimensional coordinates of the closed address space, set =1; If all key positioning elements used to generate precise one-dimensional coordinates are missing, a gated deadlock setting is triggered. =0, at this time Forced convergence to 0; Semantic energy diffusion entropy The probabilistic statistical process is as follows: The query semantic feature vector... Input L sets of hash space projection operators, calculate The set of marginal probability distributions under each projection axis Then, the projection space divergence on each projection axis is quantitatively calculated to obtain the semantic energy dispersion entropy. : The lower the value, the more it tends to converge to the local routing area in the closed address space, which means that the retrieval intent of unstructured text is more explicit; The query intent awareness module calculates the discrete convergence of the identifier. and the inverse measure of semantic energy diffusion entropy (1- ), respectively with the preset discrete convergence hard threshold and elastic semantic focus threshold Perform interval matching and dynamically distribute the current query load Q to the corresponding routing channel based on different combinations of conditions: Condition combination one: When the condition is met When both the identifier domain and semantic domain features are valid, the hybrid cooperative addressing mode is activated. Combination of conditions 2: When conditions are met When only semantic domain features are valid, the single semantic approximate routing mode is activated. Combination of conditions 3: When conditions are met When the determination is made that only the identifier field feature is valid, the single-label precise addressing mode is activated; Combination of conditions 4: When conditions are met If the query characteristics fail to meet the preset routing baseline, the module terminates the initialization routing process and returns an invalid query message to the client.
6. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing as described in claim 1, characterized in that, During the semantic query process, the semantic collaborative addressing module uses a routing scoring model based on centroid set fusion of semantic distance and path duplication penalty to select the next hop node. The node where the current query package is located is denoted as The semantic cooperative addressing module merges the current node's topological routing candidate node set with the semantic neighbor candidate node set to form a candidate node set. ; When the query packet reaches the current node At this time, the semantic cooperative addressing module uses the current node's topological routing candidate nodes and semantic neighbor candidate nodes as candidate next-hop nodes; for any candidate next-hop node The semantic collaborative addressing module reads the L semantic centroid coordinates of the candidate node: When the query initiating node receives the unstructured query payload, it maps the query payload to a query semantic vector. The same function as the resource semantic coordinate generation stage is used to generate L semantic coordinates for the query, where L is the preset number of semantic projection groups. Each semantic projection function corresponds to a semantic projection channel, which is used to generate a query semantic coordinate in the closed logical address space. For the semantic coordinates generated by the l-th semantic projection function, the semantic cooperative addressing module calculates the candidate next-hop node. The on-ring distance between the semantic centroid coordinates of query Q and the query semantic coordinates of query Q: in, Indicates the candidate next hop node The ring distance between the t-th semantic centroid and the query semantic coordinates generated by query Q in the same semantic projection channel under the L-th semantic projection channel; Furthermore, the system employs a temperature coefficient By fusing the ring distances under each semantic projection channel, candidate next-hop nodes are obtained. The fused semantic distance relative to query Q: in, Indicates the candidate next hop node The fused semantic distance between the query Q and the query Q This is a temperature coefficient used to control the degree of bias in the fused semantic distance towards the minimum distance semantic hash table; The semantic cooperative addressing module then calculates the current node. relative to the semantic distance of the L table fusion for query Q And based on candidate next-hop nodes With the current node Based on the fusion of semantic distance differences, candidate next-hop nodes are calculated. Relative semantic progress rate: in, For smoothing parameters, Indicates starting from the current node Jump to candidate next hop node The relative semantic distance reduction ratio obtained afterwards; when When, it indicates a candidate next-hop node. Relative to the current node Closer to the semantic target of the query; The semantic cooperative addressing module determines candidate next-hop nodes based on the visited node records carried in the query packet. Path duplication penalty: in, This represents the set of visited nodes carried in the query packet; The semantic cooperative addressing module calculates candidate next-hop nodes according to the following routing scoring function. Route rating: in, These are non-negative weighting coefficients; Indicates the candidate next hop node Semantic similarity to query Q; Indicates the candidate next hop node Contribution to effective semantic progress; Indicates a penalty for duplicate paths; The semantic cooperative addressing module selects the node with the highest routing score as the next hop node from the candidate next hop nodes; when all candidate next hop nodes satisfy: The semantic cooperative addressing module determines the current node. It is located in a locally semantically optimal region or a low-yield routing region, triggering an early termination decision or semantic anchor point tracing process for semantic cooperative addressing; among which, This is the minimum effective progress threshold.
7. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing as described in claim 1, characterized in that, The hop-by-hop retrieval accumulation module adopts a result accumulation mechanism based on cross-node deduplication and incremental benefit evaluation to update the accumulated result set and routing state quantity carried by the query packet during the semantic collaborative addressing process. Specifically, when the query mode is a semantic query or a hybrid query, the query packet carries at least a query semantic vector during the semantic cooperative addressing process. Query semantic coordinate set Cumulative result set The set of visited nodes Current jump count h and consecutive low-reward jump count ; When the query packet reaches the h-th hop node At that time, based on the query semantic vector and query semantic coordinate set ), at node Obtain the local candidate result set for this hop from the local resource index. ; For any candidate result r, the coordinates are identified. Unique identifier for unified resources Determine the unique deduplication key; Then, for the local candidate result set of this hop Cumulative result set with the previous hop Perform cross-node deduplication and merging to obtain a candidate merge result set. ; The hop-by-hop retrieval cumulative module further calculates the relevance score between the resource result r and the query Q. For the candidate merge result set Sort the results and retain the top K' candidate results to obtain the updated cumulative result set. ; The hop-by-hop retrieval accumulation module calculates the incremental revenue for the current hop based on the newly added valid result set; whereby the newly added valid result set for the current hop is defined as: in, The threshold for the relevance of valid results; The incremental profit for this jump is defined as: in, This indicates the effective contribution of the h-th hop node to the cumulative result set; The hop-by-hop retrieval and accumulation module updates the number of consecutive low-yield hops based on the incremental profit of the current hop; the low-yield hop count is incremented by 1 only when the incremental profit of the current hop is lower than the preset incremental profit threshold. Then the current node Write to the set of visited nodes And update the current jump count; Updated cumulative result set Continuous low-yield jumps The set of visited nodes The current hop count h is written back to the query packet for subsequent semantic collaborative addressing, early termination determination, and semantic anchor tracing.
8. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing according to claim 7, characterized in that, The hop-by-hop retrieval accumulation module also includes a semantic cooperative addressing termination determination mechanism. The semantic cooperative addressing termination determination mechanism is used to determine whether to terminate the current semantic cooperative addressing process based on the accumulated result set carried by the query packet, the number of consecutive low-yield hops, the visited node records and the current route hops, and to determine the processing path after termination. After the query packet completes the local resource retrieval and cumulative result update at the h-th hop node, the hop-by-hop retrieval and accumulation module reads the addressing status variables carried in the query packet: in, This represents the set of visited nodes; Indicates the current hop count; This indicates the number of consecutive low-yield jumps; This represents the incremental gain of the h-th hop node on the cumulative result set for this hop; Then calculate the cumulative satisfaction level after the h-th jump: in, This represents the cumulative satisfaction level after the h-th hop; Indicates the preset target number of output results. Represents the cumulative result set The number of resource results in the middle. Represents the cumulative result set Compared to the overall relevance quality of query Q, and The weight coefficients are non-negative and satisfy the following conditions: Among them, overall relevance quality Determine as follows: in, Indicates from the cumulative result set The top-ranked candidates were selected based on their relevance scores. A collection of resource results; when hour, = ; This represents the relevance score between the resource result r and the query Q; when When empty, set ; The hop-by-hop retrieval accumulation module further calculates the semantic route decay after the h-th hop based on the number of consecutive low-yield hops, visited node records, and the current route hop count. in, This represents the semantic route decay after the h-th hop; This indicates the maximum number of consecutive low-return jumps that can be preset. This indicates the preset maximum number of route hops. Indicates the path repetition intensity. , , The weight coefficients are non-negative and satisfy the following conditions: Among them, path repetition intensity Based on the set of visited nodes Determining the degree of node repetition: in, This indicates that the records are from visited nodes. The set of nodes obtained after deduplication This indicates the number of visited nodes after deduplication; Semantic route decay This is used to characterize the degree of benefit decay when continuing to perform semantic cooperative addressing; among them, the higher the number of consecutive low-benefit hops, the higher the path repetition intensity, and the closer the current route hop count is to the preset maximum route hop count, the higher the semantic route decay degree. The higher; The hop-by-hop retrieval accumulation module further sets hard termination boundaries; when the number of consecutive low-yield hops reaches the preset upper limit, or the current route hop count reaches the preset maximum route hop count, or the node where the current query packet is located already exists in the record of the visited nodes before it reaches the current node, it is determined that the current semantic cooperative addressing triggers a hard termination boundary. Then, based on the cumulative result satisfaction, semantic route decay, and hard termination boundary, the subsequent processing path of semantic cooperative addressing is determined: When the cumulative result satisfaction reaches the preset cumulative result satisfaction threshold, the current semantic collaborative addressing process is stopped, and the cumulative result set is input into the result fusion output module, which then organizes the results to generate semantic query results. When the cumulative result satisfaction does not reach the preset cumulative result satisfaction threshold, and the semantic route decay reaches the preset semantic route decay threshold, or a hard termination boundary is triggered, the current semantic cooperative addressing is determined to enter the route decay termination state, the current semantic cooperative addressing process is stopped, and the query packet state is switched to the semantic anchor tracing state. When the cumulative result satisfaction does not reach the preset cumulative result satisfaction threshold, and the semantic route decay does not reach the preset semantic route decay threshold, and no hard termination boundary is triggered, it is determined that the current semantic cooperative addressing has not yet reached the termination condition. The updated query packet state is written back to the query packet, and the semantic cooperative addressing module is called again to select the next hop node.
9. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing according to claim 1, characterized in that, The semantic anchor tracing module executes a tracing mechanism based on result gaps and tracing budget constraints, which is used to supplement the discovery of candidate resources when semantic cooperative addressing terminates and the accumulated result set does not meet the preset output requirements. The semantic anchor tracing module receives the query packet transmitted by the hop-by-hop retrieval accumulation module. The query packet carries at least the query semantic coordinate set, the accumulated result set, the visited node record, and the current route hop count. The semantic anchor tracing module reads the cumulative result set formed when semantic cooperative addressing terminates. And calculate the number of result gaps, i.e. the maximum number of semantic pointers, based on the preset target result number K: The semantic anchor tracing module locates the corresponding query semantic anchor node in the closed logical address space based on the query semantic coordinate set, and sends a backup tracing request to the query semantic anchor node. After receiving the backup tracing request, the query semantic anchor node reads candidate semantic pointers from the locally stored semantic pointer directory, and calculates the tracing score of the candidate semantic pointers based on the on-ring distance between the resource semantic coordinates and the query semantic coordinates in the candidate semantic pointers. in, Represents a candidate semantic pointer. This represents the semantic coordinates of the resource corresponding to the candidate semantic pointer. Indicates the semantic coordinates of the query. Represents the distance on the ring within the closed logical address space. Represents the coordinate scale of a closed logical address space; The query semantic anchor node sorts the candidate semantic pointers according to the traceability score and includes the traceability budget. Within the range, a set of candidate semantic pointers is returned. When the semantic anchor tracing module receives the set of candidate semantic pointers, it generates a deduplication identifier based on the resource identifier coordinates, resource unique identifier, or resource node access location information in the candidate semantic pointers, and matches the candidate semantic pointers with the existing resource results in the current accumulated result set to remove duplicate semantic pointers. The semantic anchor tracing module initiates a secondary resource discovery request to the corresponding resource node only for the candidate semantic pointers retained after deduplication, obtains supplementary resource results, and then merges the supplementary resource results with the current cumulative result set before inputting them into the result fusion output module.
10. A decentralized resource discovery and retrieval system based on dual-coordinate system indexing and semantically aware routing according to claim 1, characterized in that, The result fusion output module performs unified fusion output on the identifier addressing result, the cumulative result set formed by semantic collaborative addressing, and the traceability result set formed by the semantic anchor tracing module; The result fusion output module generates a result deduplication key based on the resource identifier coordinates and resource unique identifier information, and merges and deduplicates resource results from different sources according to the result deduplication key; The result fusion output module merges and organizes the deduplicated resource results to generate the final resource discovery results.