Distance perception-based graph division retrieval method and system and readable storage medium
By considering the distance information between nodes during the graph division process and adopting a partitioning method that minimizes the sum of the cut edge weights, the high construction cost and graph division defects of the existing graph division algorithm on large data sets are solved, and more efficient and accurate graph retrieval is achieved.
Patent Information
- Application Number
- CN202510472804.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When building graph structures of large data sets, the existing graph division algorithm has high construction costs and graph division defects, resulting in inaccurate query results and inefficient efficiency.
The graph division search method based on distance perception is adopted, and the graph index structure is optimized by selecting distance metrics to assign weights to each edge, combining nearest neighbor pairs, segmenting the graph structure at a hierarchical level, and optimizing the graph index structure by minimizing the sum of the cut edge weights.
It effectively reduces the number of cross-partition queries, improves query efficiency and accuracy, and maintains the spatial proximity of data.
Smart Images

Figure CN119988694A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vector databases, and in particular to a distance-aware graph partition retrieval method, system and readable storage medium. Background Art
[0002] In the field of computer vision, image matching and retrieval are core tasks, which are widely used in face recognition, autonomous driving, medical image analysis and other fields. With the explosive growth of image data (such as social media, medical image libraries, etc.), traditional retrieval methods based on text annotation or low-level visual features (color, texture) have gradually exposed their limitations: 1) Strong reliance on manual annotation: Traditional methods require manual annotation of image content, which is costly and difficult to cover complex scenes; 2) High feature dimension: Image features (such as SIFT and ORB) are usually high-dimensional vectors. The computational complexity of direct matching is as high as O(N²), which makes it difficult to handle large-scale data. Semantic gap: Low-level features are difficult to capture image semantic information, resulting in limited cross-modal retrieval effects.
[0003] To solve the above problems, the approximate nearest neighbor (ANN) method has become a research hotspot. Its core goal is to significantly improve the retrieval efficiency while ensuring a certain accuracy through dimensionality reduction or structured indexing technology.
[0004] The current mainstream ANN methods can be divided into four categories: 1. Tree-based methods (such as KD-Tree, Ball-Tree): accelerate nearest neighbor search through space partitioning, but have limited effect on high-dimensional data; 2. Quantization-based methods (such as LSH and PQ): feature dimensions are compressed by quantization, but quantization errors may affect matching accuracy; 3. Hash-based methods (such as CNNH and DPSH): fast matching is achieved through hash coding, but they are not robust enough to complex geometric deformations and lighting changes.
[0005] 4. Graph-based methods (such as Graph Neural Networks): Efficient retrieval is achieved by building an image similarity graph structure, with optimal performance, but there are the following bottlenecks: High construction cost: Graph construction of large data sets (such as millions of images) requires a lot of computing resources and storage space; Graph partitioning defects: Traditional graph partitioning algorithms (such as METIS) aim to minimize the number of edge cuts, which may split high-similarity nodes into different subgraphs, reducing indexing performance; Therefore, it is necessary to study a new graph partitioning algorithm to maintain the spatial proximity of data and make the query results more accurate. Summary of the invention
[0006] The purpose of the present invention is to provide a distance-aware graph partition retrieval method, system and readable storage medium to solve the technical problems existing in existing graph partition algorithms, reduce the number of cross-partition queries, and improve query efficiency and the accuracy of query results.
[0007] In order to achieve the above object, the present invention provides a distance-aware graph partition retrieval method, comprising the following steps: Construct the original data into an initial graph structure; Selecting a distance metric, assigning a weight to each edge, and merging nearest neighbor node pairs based on the weight to gradually simplify the initial graph structure to obtain a simplified coarsened graph structure; The coarsened graph structure is hierarchically segmented, and based on the segmentation situation, it is iteratively back-projected onto the initial graph structure, and finally a plurality of initial cut edge sets are obtained; Find the corresponding node according to each initial cut edge in the initial cut edge set, randomly put the node into different partitions, and optimize each initial cut edge set one by one by minimizing the sum of the cut edge weights as the objective function to obtain the final graph index structure; A query is input, the graph index structure is called, and the distances between the query and the central nodes of each partition of the graph index structure are calculated respectively, and the partition with the closest distance is selected for retrieval.
[0008] Optionally, the distance metric is used to measure the proximity between adjacent nodes, and the distance metric is Euclidean distance, cosine similarity or Manhattan distance.
[0009] Optionally, the weight is expressed as follows:
[0010] Among them, w uv represents the weight, d uv Represents the distance between adjacent nodes u and v, and ε is a constant.
[0011] Based on the same inventive concept, the present invention also provides a distance-aware graph partition retrieval system, comprising: A graph construction module, used to construct the original data into an initial graph structure; A coarsening module, configured to select a distance metric, assign a weight to each edge, and merge nearest neighbor node pairs based on the weight to gradually simplify the initial graph structure to obtain a simplified coarsening graph structure; A graph segmentation module, used for hierarchically segmenting the coarsened graph structure, and iteratively back-projecting it onto the initial graph structure based on the segmentation situation, and finally obtaining a plurality of initial cut edge sets; The edge cutting optimization module is used to find the corresponding node according to each initial edge cutting in the initial edge cutting set, randomly put the nodes into different partitions, and optimize each initial edge cutting set one by one by minimizing the sum of edge cutting weights as the objective function to obtain the final graph index structure; The retrieval module is used to input a query, call the graph index structure and respectively calculate the distance between the query and the central node of each partition of the graph index structure, and select the partition with the closest distance for retrieval.
[0012] Optionally, the distance metric is used to measure the proximity between adjacent nodes, and the distance metric is Euclidean distance, cosine similarity or Manhattan distance.
[0013] Optionally, the weight is expressed as follows:
[0014] Among them, w uv represents the weight, d uv Represents the distance between adjacent nodes u and v, and ε is a constant.
[0015] Based on the same inventive concept, the present invention also provides a readable storage medium on which a computer program is stored. When the computer program is executed, it can implement the distance-aware graph partition retrieval method as described above.
[0016] In the distance-aware graph partitioning retrieval method, system and readable storage medium provided by the present invention, by considering the distance information between nodes during the partitioning process, nodes with close distances are effectively divided into the same partition, thereby reducing the number of cross-partition queries and improving query efficiency. At the same time, by adopting a node partitioning method that minimizes the sum of the cut edge weights, the spatial proximity of the data can be maintained, making the query results more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Those skilled in the art should understand that the drawings are provided for a better understanding of the present invention and do not constitute any limitation on the scope of the present invention. Figure 1 A flowchart of a distance-aware graph partitioning retrieval method provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0018] As described in the background technology, traditional graph partitioning methods mainly focus on minimizing the number of cut edges, and often ignore the distance between adjacent nodes being partitioned. This oversight may cause closer nodes to be assigned to different partitions. In order to better maintain the local neighborhood structure, the present invention proposes a distance-aware graph partitioning retrieval method, which effectively divides nodes with closer distances into the same partition by considering the distance information between nodes during the partitioning process, thereby reducing the number of cross-partition queries and improving query efficiency. At the same time, by adopting a node partitioning method that minimizes the sum of the cut edge weights, the spatial proximity of the data can be maintained, making the query results more accurate.
[0019] In order to make the purpose, advantages and features of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the accompanying drawings are in a very simplified form and use non-precise proportions, which are only used to conveniently and clearly assist in explaining the purpose of the embodiments of the present invention. In order to make the purpose, features and advantages of the present invention more obvious and easy to understand, please refer to the accompanying drawings. It should be noted that the structure, proportion, size, etc. illustrated in the drawings of this specification are only used to match the content disclosed in the specification for people familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Any modification of the structure, change in the proportional relationship or adjustment of the size, under the same or similar conditions as the effects that can be produced by the present invention and the purposes that can be achieved, should still fall within the scope of the technical content disclosed by the present invention.
[0020] As used in the present invention, the singular forms "a", "an", and "the" include plural referents unless the context clearly indicates otherwise. As used in the present invention, the term "or" is generally used in a sense including "and / or" unless the context clearly indicates otherwise.
[0021] Please refer to Figure 1 , this embodiment provides a distance-aware graph partition retrieval method, comprising the following steps: S1, construct the original data into an initial graph structure; S2. Select a distance metric, assign a weight to each edge, and merge the nearest neighbor node pairs based on the weight to gradually simplify the initial graph structure to obtain a simplified coarsened graph structure; S3, hierarchically segmenting the coarsened graph structure, and iteratively back-projecting it onto the initial graph structure based on the segmentation conditions, and finally obtaining a plurality of initial cut edge sets; S4, finding corresponding nodes according to each initial cut edge in the initial cut edge set, randomly placing the nodes into different partitions, taking minimizing the sum of the cut edge weights as the objective function, optimizing each initial cut edge set one by one, and obtaining the final graph index structure; S5. Input a query, call the graph index structure, calculate the distance between the query and the central node of each partition of the graph index structure, and select the partition with the closest distance for retrieval.
[0022] First, execute S1 to construct the original data into an initial graph structure. In this embodiment, a traditional graph construction algorithm can be used to construct the original data into an initial graph structure, and the present invention does not limit this.
[0023] Then, S2 is executed to select a distance metric, assign a weight to each edge, and merge the nearest neighbor node pairs based on the weight to gradually simplify the initial graph structure to obtain a simplified coarsened graph structure.
[0024] In this embodiment, the distance metric is used to measure the proximity between adjacent nodes. Depending on the nature of the data set, common choices include Euclidean distance, cosine similarity, and Manhattan distance. Depending on the distance metric selected, each edge in the graph can be assigned a weight, which helps to prioritize connections between adjacent nodes during the partitioning process.
[0025] After choosing the distance metric, each edge is assigned a weight that is inversely proportional to the distance of the connected nodes. This ensures that edges connecting closely related nodes receive higher weights. The weights are expressed as follows:
[0026] Among them, w uv represents the weight, d uv Represents the distance between adjacent nodes u and v, and ε is a constant to avoid division by zero.
[0027] This approach ensures that edges between closer nodes are prioritized during the partitioning process, thereby maintaining local connectivity more effectively.
[0028] The S2 actually simplifies the graph step by step by merging node pairs. In the specific implementation, the node pairs with weights can be merged preferentially, and the weights indicate that the node pairs are closer. By merging nodes with their nearest neighbors, this stage retains the key neighborhood relationships and maintains the local structure of the graph, which is very important for accurate segmentation of the graph. This weighted coarsening method allows the graph to maintain its proximity-based structure while reducing its size.
[0029] It should be noted that the simplified graph mentioned here does not only refer to simplification once, but also includes multiple simplifications. For example, the initial graph structure includes 100 nodes, and after the first merging of node pairs, there are 50 nodes left, and then after the second merging of node pairs, there are 25 nodes left, and so on, until the smallest and simplest version of the coarsened graph structure is obtained.
[0030] Then, S3 is executed to hierarchically segment the coarsened graph structure, and iteratively back-project it onto the initial graph structure based on the segmentation situation, and finally obtain a plurality of initial cut edge sets.
[0031] In this embodiment, partitions are divided according to the number of data points on the final coarsened graph structure, and the number of partitions is the same as the number of points.
[0032] Then, S4 is executed to find the corresponding node according to each initial cut edge in the initial cut edge set, randomly put the node into different partitions, and optimize each initial cut edge set one by one with minimizing the sum of the cut edge weights as the objective function to obtain the final graph index structure.
[0033] It should be emphasized that the traditional graph partitioning algorithm minimizes the total number of cut edges by not considering the strength or importance of each edge, ignoring the distance information between nodes, which easily leads to adjacent nodes being divided into different partitions, thus affecting the query efficiency and recall rate. In this technology, the present invention introduces distance-based weights and adjusts the optimization goal to minimize the sum of cut edge weights rather than just the number of edges. Such a design can effectively reduce the possibility of separating closely connected nodes, thereby improving the quality of segmentation in maintaining local structure, while also maintaining the spatial proximity of data, making query results more accurate.
[0034] In this embodiment, the modified objective function becomes:
[0035] Among them, E cut represents the initial cut edge set, w uv Represents the weight of each cut edge, reflecting the proximity between connected nodes.
[0036] In this embodiment, each initial edge cutting set is optimized in turn. The optimization process is further described below using one of the edge cutting sets as an example.
[0037] 1) Find the corresponding graph nodes according to the initial edge cutting set, randomly put the nodes into different partitions, update the initial edge cutting set, and then calculate which of the updated edge cutting sets has the smallest corresponding objective function value. After finding the minimum value, end the optimization of this initial edge cutting set.
[0038] 2) Then start the optimization of the second initial cut edge set, the method is the same as the above process.
[0039] 3) The entire optimization process ends when all initial cut edge sets are optimized.
[0040] It should be noted that the most accurate way to update the initial cut edge set is to traverse all possibilities. However, considering that there are many points involved (for example, the number of nodes involved is m), there will be many corresponding solutions for moving partitions (the solutions involved are 2 m ), so we can set a threshold x and select the minimum value of the objective function after calculating n partitioning schemes. The specific value of x needs to be determined according to the number of nodes involved, for example, x = 1 / 3 *2 m .
[0041] Finally, S5 is executed to input a query, call the graph index structure, calculate the distance between the query and the central node of each partition of the graph index structure, and select the partition with the closest distance for retrieval.
[0042] For example, after obtaining a query (data), the query is first calculated with the central node of each partition of the graph index structure obtained by the distance-aware graph partitioning retrieval method, so as to select which partition to enter for calculation, and finally return the top-k values as the final retrieval results. This can reduce the calculation steps and improve indexing efficiency.
[0043] The following compares the recall rate and the number of queries per second of the distance-aware graph partitioning retrieval method provided by the present invention with the existing methods through specific experimental data.
[0044] Experimental environment: Operating system: Ubuntu 20.04.5 LTS CPU: Intel(R) Xeon(R) Platinum 8362 CPU @ 2.80GHz Memory: 256G Datasets: SIFT1M, GIST The recall rates of each method for SIFT128, 1M when calculating 12,000 distance calculations are shown in Table 1: Table 1
[0045] The number of queries per second for each method when querying 1000 queries for SIFT128, 1M is shown in Table 2: Table 2
[0046] GIST960, 1M The corresponding recall rates of each method when calculating 12,000 distance calculations are shown in Table 3: Table 3
[0047] The number of queries per second for each method when GIST960, 1M queries 1000 queries is shown in Table 4: Table 4
[0048] According to Tables 1 to 4, when the initial degrees of the same edges are consistent, the method proposed in the present invention is superior to the existing methods in terms of recall rate (Recall) and number of queries per second (Query Per Second, QPS).
[0049] Based on the same inventive concept, an embodiment of the present invention further proposes a graph partitioning retrieval system based on distance perception, comprising: A graph construction module, used to construct the original data into an initial graph structure; A coarsening module, configured to select a distance metric, assign a weight to each edge, and merge nearest neighbor node pairs based on the weight to gradually simplify the initial graph structure to obtain a simplified coarsening graph structure; A graph segmentation module, used for hierarchically segmenting the coarsened graph structure, and iteratively back-projecting it onto the initial graph structure based on the segmentation situation, and finally obtaining a plurality of initial cut edge sets; The edge cutting optimization module is used to find the corresponding node according to each initial edge cutting in the initial edge cutting set, randomly put the nodes into different partitions, and optimize each initial edge cutting set one by one by minimizing the sum of edge cutting weights as the objective function to obtain the final graph index structure; The retrieval module is used to input a query, call the graph index structure and respectively calculate the distance between the query and the central node of each partition of the graph index structure, and select the partition with the closest distance for retrieval.
[0050] In this embodiment, the distance metric is used to measure the proximity between adjacent nodes, and the distance metric is Euclidean distance, cosine similarity or Manhattan distance.
[0051] Preferably, the weight is expressed as follows:
[0052] Among them, w uv represents the weight, d uv Represents the distance between adjacent nodes u and v, and ε is a constant to avoid division by zero.
[0053] Based on the same inventive concept, an embodiment of the present invention further proposes a readable storage medium on which a computer program is stored. When the computer program is executed, it can implement the distance-aware graph partition retrieval method as described above.
[0054] The readable storage medium can be a tangible device that can keep and store the instructions used by the instruction execution device, such as but not limited to an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device or any suitable combination thereof. The more specific example (non-exhaustive list) of the readable storage medium includes: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination thereof. The computer program described herein can be downloaded from the readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network and / or a wireless network. The network can include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer program from the network and forwards the computer program for storage in the readable storage medium in each computing / processing device. The computer program for performing the operation of the present invention can be an assembly instruction, an instruction set architecture (ISA) instruction, a machine instruction, a machine-related instruction, a microcode, a firmware instruction, a state setting data, or a source code or object code written in any combination of one or more programming languages, including object-oriented programming languages-such as Smalltalk, C++, etc., and conventional procedural programming languages-such as "C" language or similar programming languages. The computer program can be executed completely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on the remote computer, or completely on the remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, using an Internet service provider to connect through the Internet). In some embodiments, by utilizing the state information of a computer program to personalize an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute computer-readable program instructions to implement various aspects of the present invention.
[0055] Here, various aspects of the present invention are described with reference to the flowchart and / or block diagram of the method, system and computer program product according to the embodiment of the present invention. It should be understood that each square frame of the flowchart and / or block diagram and the combination of the square frames in the flowchart and / or block diagram can be realized by a computer program. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device, so as to produce a machine, so that these programs, when executed by the processor of a computer or other programmable data processing device, produce a device for realizing the function / action specified in one or more square frames in the flowchart and / or block diagram. These computer programs can also be stored in a readable storage medium, and these computer programs make the computer, programmable data processing device and / or other equipment work in a specific way, so that the readable storage medium storing the computer program includes a manufactured product, which includes instructions for realizing various aspects of the function / action specified in one or more square frames in the flowchart and / or block diagram.
[0056] The computer program may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are executed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the computer program executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0057] In summary, the embodiments of the present invention provide a distance-aware graph partitioning retrieval method, system, and readable storage medium. By considering the distance information between nodes during the partitioning process, nodes with close distances are effectively divided into the same partition, thereby reducing the number of cross-partition queries and improving query efficiency. At the same time, by adopting a node partitioning method that minimizes the sum of the cut edge weights, the spatial proximity of the data can be maintained, making the query results more accurate.
[0058] In addition, it should be recognized that although the present invention has been disclosed as a preferred embodiment, the above embodiment is not intended to limit the present invention. For any technician familiar with the art, without departing from the scope of the technical solution of the present invention, the technical content disclosed above can be used to make many possible changes and modifications to the technical solution of the present invention, or modified into equivalent embodiments of equivalent changes. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still belongs to the scope of protection of the technical solution of the present invention.
Claims
1. A distance-aware graph partitioning retrieval method, characterized in that: The following steps are involved: Construct the original data into an initial graph structure; Selecting a distance metric, assigning a weight to each edge, and merging nearest neighbor node pairs based on the weight to gradually simplify the initial graph structure to obtain a simplified coarsened graph structure; The coarsened graph structure is hierarchically segmented, and based on the segmentation situation, it is iteratively back-projected onto the initial graph structure, and finally a plurality of initial cut edge sets are obtained; Find the corresponding node according to each initial cut edge in the initial cut edge set, randomly put the node into different partitions, and optimize each initial cut edge set one by one by minimizing the sum of the cut edge weights as the objective function to obtain the final graph index structure; A query is input, the graph index structure is called, and the distances between the query and the central nodes of each partition of the graph index structure are calculated respectively, and the partition with the closest distance is selected for retrieval.
2. The distance-aware graph partitioning retrieval method according to claim 1, characterized in that: The distance metric is used to measure the proximity between adjacent nodes, and the distance metric is Euclidean distance, cosine similarity or Manhattan distance.
3. The distance-aware graph partitioning retrieval method according to claim 1 is characterized in that: The weights are expressed as follows: Among them, w uv represents the weight, d uv Represents the distance between adjacent nodes u and v, and ε is a constant.
4. A distance-aware graph partitioning retrieval system, characterized in that: include: A graph construction module, used to construct the original data into an initial graph structure; A coarsening module, configured to select a distance metric, assign a weight to each edge, and merge nearest neighbor node pairs based on the weight to gradually simplify the initial graph structure to obtain a simplified coarsening graph structure; A graph segmentation module, used for hierarchically segmenting the coarsened graph structure, and iteratively back-projecting it onto the initial graph structure based on the segmentation situation, and finally obtaining a plurality of initial cut edge sets; The edge cutting optimization module is used to find the corresponding node according to each initial edge cutting in the initial edge cutting set, randomly put the nodes into different partitions, and optimize each initial edge cutting set one by one by minimizing the sum of edge cutting weights as the objective function to obtain the final graph index structure; The retrieval module is used to input a query, call the graph index structure and respectively calculate the distance between the query and the central node of each partition of the graph index structure, and select the partition with the closest distance for retrieval.
5. The distance-aware graph partitioning retrieval system according to claim 4 is characterized in that: The distance metric is used to measure the proximity between adjacent nodes, and the distance metric is Euclidean distance, cosine similarity or Manhattan distance.
6. The distance-aware graph partitioning retrieval system according to claim 4, characterized in that: The weights are expressed as follows: Among them, w uv represents the weight, d uv Represents the distance between adjacent nodes u and v, and ε is a constant.
7. A readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed, it can implement the distance-aware graph partitioning retrieval method according to any one of claims 1 to 3.
Citation Information
Patent Citations
A hierarchical image segmentation method based on multi-scale edge cues
CN109272467A
Hypergraph segmentation method and system based on hyperedge clustering coarsening and medium
CN114626331A
Mass high-dimensional data learning index construction method based on partition hierarchy diagram
CN116992091A
Node classification method, system, equipment and product based on improved graph Transform model
CN118916757A
Image Searching By Approximate k-NN Graph
US20130230255A1
Cited By
Distributed graph index nearest neighbor search method oriented to high-dimensional space
CN121070937A
Image retrieval method and system based on large model
CN121479007A