Data retrieval method and system based on adaptive space division and hash coding

By using an adaptive space partitioning and hash encoding method, the problem of high computational complexity of existing learning hashing methods on large-scale datasets is solved, enabling fast and low-cost data retrieval, which is suitable for similarity search on large-scale datasets.

CN120873244APending Publication Date: 2025-10-31NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511174899.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing learning hash (L2H) methods are computationally complex and costly on large-scale datasets, failing to meet real-time requirements. Furthermore, traditional methods rely on complex learning and optimization processes, resulting in long training times.

Method used

An adaptive spatial partitioning and hash coding method is adopted. By extracting the feature vectors of the original data in the database, the feature space sub-units are generated using the adaptive spatial partitioning method, and binary hash codes are generated based on the hash coding mechanism to achieve fast similarity retrieval.

Benefits of technology

It reduces computational complexity, improves retrieval efficiency, and reduces memory usage, making it suitable for fast similarity retrieval of large-scale datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873244A_ABST
    Figure CN120873244A_ABST
Patent Text Reader

Abstract

The invention relates to a data retrieval method based on adaptive space division and hash coding, and relates to the technical field of data retrieval. The method comprises the following steps: extracting feature vectors of original data and to-be-queried data; dividing a feature space Rd formed by the feature vectors based on a preset adaptive space division method, and repeating T times to generate T groups of feature space subunits of the Rd; distributing the data points of the Rd to the feature space subunits of the Rd; coding the T psi-bit vector sets H (x) based on a preset Hash coding mechanism, and connecting to generate an L-bit binary Hash code of each data point; and calculating the distance between the original data and the L-bit binary hash code of the to-be-queried data, and taking the original data corresponding to the minimum distance as a target retrieval result. The compact binary code is generated, and rapid similarity retrieval is realized in the binary code space, so that the problems of complex data coding, high calculation cost, low retrieval efficiency and high memory occupation in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data retrieval, and in particular to a data retrieval method and system based on adaptive spatial partitioning and hash coding. Background Technology

[0002] With the advent of the big data era, fast and accurate similarity search and data retrieval from massive amounts of data has become a key challenge in the field of information technology. Traditional precise search methods, such as linear scanning, are inefficient when dealing with huge datasets and cannot meet real-time requirements. Learning-to-Hash (L2H) technology significantly improves retrieval speed by mapping high-dimensional data to compact binary hash codes and performing efficient approximate nearest neighbor search in Hamming space. However, existing L2H methods typically rely on complex learning and optimization processes to ensure the quality of hash codes, which not only increases computational overhead but also has inherent limitations in the types of hash functions they rely on (such as thresholding, hypersphere, and hyperplane methods).

[0003] The complexity of existing Learning to Hash (L2H) methods stems from the nature of their "learning" process. Instead of using fixed, data-independent hash functions, these methods aim to actively learn a set of hash functions that preserve the similarity relationships between original data points in the target Hamming space through a carefully designed optimization process. This "learning" process typically involves multiple computationally intensive stages, which is the main reason for their long training time and high computational cost.

[0004] Specifically, its complexity is mainly reflected in the following aspects:

[0005] 1. Full Modeling of Data Structure: The first step in many mainstream L2H methods is to model the internal structure of the entire training dataset. This typically involves constructing a large-scale N×N similarity matrix W (where N is the number of training samples), where each element Wij represents the similarity between data points i and j. Simply constructing this matrix requires O(N²) level pairwise comparison calculations, and when N reaches millions, the computational and storage overhead becomes extremely large.

[0006] 2. NP-hard optimization objectives: These methods define an objective function that mathematically describes the conditions that a "good hash code" should satisfy. A typical objective is to minimize the sum of weighted Hamming distances between the hash codes of all similar pairs of points. Since hash codes are discrete binary values, such objective functions usually constitute a combinatorial optimization problem, which is often NP-hard to solve directly and cannot find the optimal solution in polynomial time.

[0007] 3. High cost of relaxation process: To solve the aforementioned NP-hard problem, existing methods generally employ a "relaxation" strategy, which transforms the discrete binary optimization problem into a problem in a continuous domain, and then obtains the final binary code through quantization or thresholding. This relaxation process is the core of the computational cost.

[0008] Taking the classic Spectral Hashing (SH) method as an example, its process illustrates the aforementioned complexity: First, a complete N×N similarity matrix W needs to be calculated based on a Gaussian kernel function. Next, the objective of minimizing hash code distance is relaxed to solving the eigenvector problem of a graph Laplacian matrix. This step requires eigenvalue decomposition, which typically has a computational complexity of O(N³) for a dense N×N matrix, resulting in extremely high computational costs. To ensure the quality of the hash code, additional constraints of "maximizing entropy" (ensuring a balanced distribution of 0s and 1s) and "bit independence" (ensuring no redundancy in different hash bits) must be imposed, further increasing the complexity of the optimization problem. Finally, the obtained real-valued eigenvectors need to undergo thresholding before the final binary code is generated.

[0009] Another type of method, iterative optimization learning, such as Iterative Quantization (ITQ), avoids constructing an N×N matrix, but its learning process is equally complex. ITQ first requires dimensionality reduction and decorrelation of the data using Principal Component Analysis (PCA). Then, it iteratively optimizes multiple times to learn an optimal rotation matrix that minimizes the quantization error between the data point and the nearest hypercube vertex after rotation. The convergence speed and computational cost of this iterative process are also considerable.

[0010] In summary, the "learning" process of existing L2H methods is essentially a comprehensive process involving full data structure modeling, complex objective function definition, NP-hard problem relaxation, large-scale matrix operations (such as eigenvalue decomposition), and multi-step iterative optimization. These intensive and complex computational steps collectively contribute to their high training time and computational cost, thus limiting their application potential in scenarios with large-scale datasets and high real-time requirements. Summary of the Invention

[0011] Therefore, it is necessary to provide a data retrieval method and system based on adaptive spatial partitioning and hash encoding to address the above-mentioned technical problems. A data-dependent binary method framework that does not require learning is disclosed. By converting raw data of any type into a unified vectorized representation, and then using an adaptive spatial partitioning method and encoding mechanism to generate compact binary codes, fast similarity retrieval is finally achieved in the binary code space. This solves the problems of complex encoding, high computational cost, low retrieval efficiency and high memory consumption in existing data.

[0012] Firstly, this application provides a data retrieval method based on adaptive spatial partitioning and hash encoding, including:

[0013] Extract feature vectors from the original data in the database to generate a high-dimensional feature vector set X∈R. d×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension of the feature vectors of the original data.

[0014] R is partitioned based on a preset adaptive space partitioning method. d Repeat this process T times to generate T sets of R. d The characteristic space subunit;

[0015] For each group R d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-vectors for each data point, H(x) = [h1(x),...,h...]. i (x),...,h ψ (x)];where h i (x)∈{0,1} is the i-th hash function for data point x if and only if x belongs to the i-th cell. i (x) = 1, and the rest are 0; ψ is the number of a set of feature space sub-units;

[0016] Based on a preset hash encoding mechanism, the set of T ψ-bit vectors H(x) is encoded to generate T ω-bit encoding blocks for each data point. The T ω-bit encoding blocks for each data point are then connected to generate the L-bit binary hash code B(x) for each data point; where T = L / ω.

[0017] Receive the data to be queried and convert it into an L-bit binary hash code B(q). Calculate the distance between B(q) and B(x) of each data point, and take the original data with the smallest distance as the target retrieval result.

[0018] In one embodiment, R is partitioned based on a preset adaptive spatial partitioning method. d ,include:

[0019] From R d ψ data points are randomly sampled from the corresponding high-dimensional feature vector set to generate a data point set D;

[0020] Based on a preset spatial partitioning algorithm, R is divided according to ψ data points in D. d The space is divided into ψ feature space sub-units; among them, R is divided using a spatial partitioning algorithm. d Covering the entire R d In D, the data points are located at the center of their corresponding feature space sub-units.

[0021] In one embodiment, the spatial partitioning algorithm includes the Voronoi diagram method, the KD-tree based method, and a variant partitioning method based on a ball tree.

[0022] In one embodiment, R d Data points assigned to R d The feature space subunits include:

[0023] Calculate R d The distance between the data points and ψ data points in D, R d The data points are assigned to the feature space sub-units corresponding to the data points with the smallest distance in D.

[0024] In one embodiment, R d Data points assigned to R d The feature space subunit also includes:

[0025] If R d When a data point is located at the boundary of a feature space sub-unit, the data point is assigned to any feature space sub-unit that shares that boundary.

[0026] In one embodiment, encoding a set of T ψ bit vectors H(x) based on a preset hash encoding mechanism includes:

[0027] Based on a pre-defined hash coding mechanism, the set of T ψ-bit vectors H(x) is converted into corresponding T ω-bit coded blocks b(x) = [e1(x), e2(x), ..., e ω (x)];where e j The formula for calculating (x) is as follows:

[0028]

[0029] Where j = 1, 2, ..., ω; ψ ≤ 2 ω .

[0030] In one embodiment, the formula for calculating the distance between the binary hash code of the data to be queried and the binary hash code of each data item in the original data is as follows:

[0031]

[0032] Where; B(x) i Let x be the i-th data point. i The L-bit binary hash code B(x); b k (x i Let x be the i-th data point. i The k-th coded block; I(·) is the indicator function, when b k (x) and b k (q) Take 1 if they are not equal, otherwise take 0.

[0033] Secondly, this application also provides a data retrieval system based on adaptive spatial partitioning and hash encoding, including:

[0034] The feature extraction module is used to extract feature vectors from the original data in the database and generate a high-dimensional feature vector set X∈R of the original data. d×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension of the feature vectors of the original data.

[0035] The feature vector space partitioning module is used to partition R based on a preset adaptive space partitioning method. d Repeat this process T times to generate T sets of R. d The characteristic space subunit;

[0036] The data point allocation module is used for each group of R d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-vectors for each data point, H(x) = [h1(x),...,h...]. i (x),...,h ψ (x)];where h i (x)∈{0,1} is the i-th hash function for data point x if and only if x belongs to the i-th cell. i (x) = 1, and the rest are 0; ψ is the number of a set of feature space sub-units;

[0037] The encoding module is used to encode the set of T ψ-bit vectors H(x) based on a preset hash encoding mechanism, generate T ω-bit encoding blocks for each data point, and connect the T ω-bit encoding blocks of each data point to generate the L-bit binary hash code B(x) for each data point; where T = L / ω;

[0038] The distance calculation module is used to receive the data to be queried and convert it into an L-bit binary hash code B(q). It calculates the distance between B(q) and B(x) of each data point and takes the original data with the smallest distance as the target retrieval result.

[0039] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:

[0040] Extract feature vectors from the original data in the database to generate a high-dimensional feature vector set X∈R. d×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension of the feature vectors of the original data.

[0041] R is partitioned based on a preset adaptive space partitioning method. d Repeat this process T times to generate T sets of R. d The characteristic space subunit;

[0042] For each group R d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-bit hash functions H(x) = [h1(x),...,h2] for each data point. i (x),...,h ψ (x)];where h i (x)∈{0,1} is the i-th hash function for data point x if and only if x belongs to the i-th cell. i (x) = 1, and the rest are 0; ψ is the number of a set of feature space sub-units;

[0043] Based on a preset hash encoding mechanism, the set of T ψ-bit vectors H(x) is encoded to generate T ω-bit encoding blocks for each data point. The T ω-bit encoding blocks for each data point are then connected to generate the L-bit binary hash code B(x) for each data point; where T = L / ω.

[0044] Receive the data to be queried and convert it into an L-bit binary hash code B(q). Calculate the distance between B(q) and B(x) of each data point, and take the original data with the smallest distance as the target retrieval result.

[0045] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:

[0046] Extract feature vectors from the original data in the database to generate a high-dimensional feature vector set X∈R. d×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension of the feature vectors of the original data.

[0047] R is partitioned based on a preset adaptive space partitioning method. d Repeat this process T times to generate T sets of R. d The characteristic space subunit;

[0048] For each group R d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-bit hash functions H(x) = [h1(x),...,h2] for each data point. i (x),...,h ψ (x)];where h i (x)∈{0,1} is the i-th hash function for data point x if and only if x belongs to the i-th cell. i (x) = 1, and the rest are 0; ψ is the number of a set of feature space sub-units;

[0049] Based on a preset hash encoding mechanism, the set of T ψ-bit vectors H(x) is encoded to generate T ω-bit encoding blocks for each data point. The T ω-bit encoding blocks for each data point are then connected to generate the L-bit binary hash code B(x) for each data point; where T = L / ω.

[0050] Receive the data to be queried and convert it into an L-bit binary hash code B(q). Calculate the distance between B(q) and B(x) of each data point, and take the original data with the smallest distance as the target retrieval result.

[0051] This application employs the aforementioned data retrieval method and system based on adaptive spatial partitioning and hash encoding, which has the following beneficial effects:

[0052] 1. The method proposed in this invention only includes random sampling of data points and partitioning based on nearest neighbor rules in its "training" phase. The computational complexity is linearly related to the size and dimension of the dataset. It does not require the complex learning or optimization process required by the traditional L2H method, thereby reducing computational complexity, helping to reduce the amount of computation and improve coding efficiency.

[0053] 2. During the retrieval phase, calculating the distance between binary codes can accelerate the query process, improving retrieval efficiency compared to calculating complex metrics such as Euclidean distance in the original high-dimensional space.

[0054] 3. The method proposed in this application converts high-dimensional floating-point feature vectors into low-dimensional binary hash codes, which can greatly compress data storage space and reduce memory and storage costs. Attached Figure Description

[0055] Figure 1 This is a flowchart of a data retrieval method based on adaptive space partitioning and hash encoding in one embodiment;

[0056] Figure 2 This is a schematic diagram of the mAP scores of VDeH and the competitive hashing method on a benchmark dataset in one embodiment.

[0057] Figure 3 This is a schematic diagram of the PR curves of VDeH and the contention hash method when the hash bit length is 512 in one embodiment;

[0058] Figure 4 This is a schematic diagram illustrating the training time of VDeH and the competition hashing method in another embodiment. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0060] The encoding method used in this application needs to satisfy full space coverage, entropy maximization, and bit independence. Full space coverage is defined as the partitioning strategy being able to cover the entire feature space R. d This ensures that any data point can be assigned to a partition, thereby obtaining an effective hash code; the entropy maximization is defined as that for a given dataset X, each partition should contain approximately the same number of data points, so that every bit of the hash code is fully utilized, thereby maximizing the information entropy of the hash code; bit independence means that every bit of the final generated binary hash code should be statistically independent of each other, avoiding information redundancy and maximizing the discriminative ability of the hash code.

[0061] It should be noted that any encoding method that satisfies the above three properties can implement the solution of this invention.

[0062] This application provides a general definition of an adaptive spatial partitioning method, namely, an adaptive spatial partitioning hash cluster H(X) is a family of hash functions derived from a dataset X and associated with a distribution P on this family of functions. H Its core characteristics must satisfy the following conditions: for the feature space R d The joint probability that any two data points x and y in the space are mapped to 1 by the same hash function h randomly drawn from this family of functions is equal to a data-dependent similarity function K(x,y|H(X)) defined by the space partitioning strategy itself. Its mathematical form is:

[0063]

[0064] In this application, adaptive space partitioning methods include Voronoi graph methods, KD-tree based methods, and variant partitioning methods based on ball trees. The following section uses Voronoi graph-based encoded hashing (VDeH) as an example for illustration.

[0065] Table 1: Key symbols used.

[0066]

[0067] Specifically, referring to Table 1, the following findings were made based on the Voronoi diagram:

[0068] A kernel function based on Voronoi diagrams, called the isolation kernel, is defined as follows:

[0069]

[0070] Each Voronoi diagram H has a cell θ, and 1(·) is an indicator function. By rewriting each cell as a hash function h, it can be reformulated as:

[0071]

[0072] Based on the above findings and the fact that Voronoi diagrams naturally satisfy the three properties mentioned above, this application proposes a hash definition. For a given dataset XCR... d×N The proposed VDeH scheme is defined as follows:

[0073] The VDeH scheme is a family of hash functions H(X) created from Voronoi diagrams, which corresponds to a distribution P on H(X) generated from a dataset X. H The correlation makes it satisfy the following formula:

[0074]

[0075] in, H is derived from the Voronoi graph generated by X, and H is the set of all hash functions h derived from a Voronoi graph.

[0076] The above definition simplifies the implementation of VDeH, as the hash function can be obtained by simply adding an extra encoding to the Voronoi diagram already used in the isolated core. VDeH does not require special design or learning of hash functions to achieve full space coverage, entropy maximization, and bit independence.

[0077] Firstly, based on the above content, referring to Figure 1 This application provides a data retrieval method based on adaptive spatial partitioning and hash encoding, including:

[0078] S100, Extract the feature vectors from the original data in the database to generate a high-dimensional feature vector set X∈R of the original data. d ×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension corresponding to the feature vectors of the original data.

[0079] In one embodiment, the raw data in the database includes image data, text data, graph data, trajectory data, and audio data.

[0080] For image data, the outputs of fully connected or pooling layers are extracted as feature vectors using computer vision feature extractors or pre-trained deep convolutional neural networks (CNNs). Computer vision feature extractors include SIFT, SURF, and ORB; CNNs include ResNet, VGG, and EfficientNet.

[0081] For text data such as articles, web pages, and user comments, traditional methods such as the bag-of-words (BoW) model and TF-IDF can be used, or pre-trained language models such as BERT, GPT, or various Transformer variants can be employed to encode the text into fixed-length semantic vectors.

[0082] For graph-structured data such as social networks and molecular structures, algorithms such as Graph Neural Networks (GCN) and GraphSAGE can be used to learn the node embeddings in the graph and generate feature vectors for the graph structure.

[0083] For trajectory data such as GPS trajectories and user behavior sequences, feature vectors can be extracted by segmentation and aggregation, or by using sequence models such as recurrent neural networks (RNN) and LSTM.

[0084] For audio data, features such as Mel-frequency cepstral coefficients (MFCCs) and spectrograms can be extracted, or audio processing deep models (such as VGGish) can be used to vectorize the audio data to generate feature vectors.

[0085] Based on the feature extraction method described above, any type of raw data in the database can be converted into a high-dimensional feature vector set X = [x1,...,x...]. N ]∈R d×N .

[0086] S200, based on a preset adaptive spatial partitioning method, partitions R. d Repeat this process T times to generate T sets of R. d The characteristic space sub-unit.

[0087] In one embodiment, R is partitioned based on a preset adaptive spatial partitioning method. d Including: from R d Randomly sample ψ data points from the corresponding high-dimensional feature vector set to generate a data point set D; based on a preset spatial partitioning algorithm, divide R according to the ψ data points in D. d The space is divided into ψ feature space sub-units; among them, R is divided using a spatial partitioning algorithm. d Covering the entire R d This ensures that any data point can be assigned to a feature space sub-unit, thereby obtaining an effective hash code; the data points in D are located at the center of their corresponding feature space sub-units.

[0088] In one embodiment, the spatial partitioning algorithm includes the Voronoi diagram method, the KD-tree based method, and a variant partitioning method based on a ball tree.

[0089] Randomly sample ψ data points from dataset X to form a data point set D = {s1,...,s...} ψ In this embodiment, to facilitate subsequent binary encoding, ψ is set to 2. ω Where ω is an integer. Taking VDeH as an example, a Voronoi diagram is constructed using the data point set D, which represents the entire feature space R. d It is divided into ψ characteristic space sub-units, namely Voronoi units. (Refer to...) Figure 2 , Figure 2 The process of encoding using VDeH is given when N=5, d=5, L=4 and w=2.

[0090] Specifically, for a given set of data points D = {s1,...,s...} i ,...,s ψ}CX, containing ψ randomly selected data points. A Voronoi diagram H created by D will be R dThe space is divided into ψ Voronoi units, where s i Located at the center of Voronoi unit i. Let the set of hash functions created by D be H(x) = [h1(x),...,h...]. i (x),...,h ψ (x)]. For any data point x∈R d h i The definition of (x) is:

[0091]

[0092] Where, dist(x, s) i )=||xs i ||For x and s i The L2 distance between them; dist(x, D) = min y∈D / {X} ||xy||.

[0093] S300, for each group R d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-bit hash functions H(x) = [h1(x),...,h2] for each data point. i (x),...,h ψ (x)];where h i (x)∈{0,1} is the i-th hash function for data point x if and only if x belongs to the i-th cell. i (x) = 1, and the rest are 0; ψ is the number of a set of characteristic space sub-units.

[0094] In one embodiment, R d Data points assigned to R d The feature space subunits include:

[0095] Calculate R d The distance between the data points and ψ data points in D, R d The data points are assigned to the feature space sub-units corresponding to the data points with the smallest distance in D.

[0096] In one embodiment, R d Data points assigned to R d The feature space subunit also includes:

[0097] If R d When a data point is located at the boundary of a feature space sub-unit, the data point is randomly assigned to any feature space sub-unit that shares that boundary.

[0098] S400: Based on a preset hash encoding mechanism, T sets of ψ-bit vectors H(x) are encoded to generate T ω-bit encoding blocks for each data point. The T ω-bit encoding blocks for each data point are then connected to generate an L-bit binary hash code B(x) for each data point; where T = L / ω.

[0099] In one embodiment, encoding a set of T ψ bit vectors H(x) based on a preset hash encoding mechanism includes:

[0100] Based on a pre-defined hash coding mechanism, the set of T ψ-bit vectors H(x) is converted into corresponding T ω-bit coded blocks b(x) = [e1(x), e2(x), ..., e ω (x)];where e j The formula for calculating (x) is as follows:

[0101]

[0102] Where j = 1, 2, ..., ω; ψ ≤ 2 ω .

[0103] S500 receives the data to be queried and converts it into an L-bit binary hash code B(q). It calculates the distance between B(q) and B(x) of each data point and takes the original data with the smallest distance as the target retrieval result.

[0104] It is understandable that B(x) = [b 1 (x),...,b k (x),...,b L / ω (x)]; where, when data point x is the i-th data point x i When, that is, x = x i When, B(x) i )=[b 1 (x i ),...,b k (x i ),...,b L / ω(x i )).

[0105] It is understandable that B(q) = [b 1 (q),...,b k (q),...,b L / ω (q)].

[0106] In one embodiment, the formula for calculating the distance between the binary hash code of the data to be queried and the binary hash code of each data item in the original data is as follows:

[0107]

[0108] Where; B(x) i Let x be the i-th data point. i The L-bit binary hash code B(x); b k (x i Let x be the i-th data point. i The k-th coded block; I(·) is the indicator function, when b k (x) and b k (q) is set to 1 if they are not equal, otherwise it is set to 0. This distance calculation only involves XOR and counting operations, which helps to improve the calculation speed.

[0109] The system sorts data based on calculated distances and selects one or more data items with the smallest distance as the most similar results. This application does not limit the number of most similar results; it can be set according to actual needs.

[0110] The encoding method employed in this application converts high-dimensional floating-point feature vectors into low-dimensional binary hash codes, which can significantly compress data storage space and reduce memory and storage costs. This is crucial for deploying large-scale search engines on memory-constrained devices (such as mobile terminals and embedded systems) or for processing ultra-large-scale datasets that are already in the terabyte or even petabyte range.

[0111] For example, suppose a database contains N = 1,000,000 data points, each represented by a d = 512-dimensional feature vector, and each dimension of the feature vector is stored using half-precision floating-point numbers (float16, i.e., 2 bytes). The original storage space would be:

[0112] N×d×sizeof(float16)=1000000×512×2bytes=1024000000bytes≈1024GB

[0113] The storage space after using the encoding method of this invention (mapped to a 256-bit hash code) is:

[0114] N×L÷8=1000000×256bits÷8=32000000bytes=32MB

[0115] To evaluate VDeH, experiments were conducted using four public datasets: two image datasets and two text datasets. Notably, existing research primarily evaluates performance on image datasets; this application extends this to text datasets to leverage their inherent complexity for a comprehensive analysis of the aforementioned methods.

[0116] The four public datasets differ in size and dimensionality, specifically: CIFAR-10, GIST, Nytimes, and Kosarak. CIFAR-10 contains 60,000 images (512 dimensions); GIST contains one million images represented by 960-dimensional descriptors; Nytimes contains 290,000 articles (256 dimensions); and Kosarak contains 74,962 clickstream news items (27,983 dimensions).

[0117] This application randomly samples 10,000 instances from the aforementioned data as the training set, and then randomly samples 500 instances different from the training set as query instances. Retrieval performance is evaluated using two commonly used metrics: mean precision (mAP) and precision-recall curve (PR curve). The performance of VDeH is compared in this application with three types of state-of-the-art methods: (1) threshold-based methods: spectral hashing (SH), cyclic binary embedding (CBE), iterative quantization (ITQ), bilinear projection (BP), and sparse projection (SP); (2) hypersphere-based methods: spherical hashing (SpH); and (3) hyperplane-based methods: density-sensitive hashing (DSH), sparse embedding and minimum variance coding (SELVE), and refined code for locality-sensitive hashing (rcLSH). Parameters in each hashing method are set to the authors' default or suggested values. For VDeH, parameters are searched in t222328u. All experiments were performed on a Linux CPU machine: an AMD 128-core CPU, each core running at 2GHz, and 1TB of memory.

[0118] Figure 2 The mAP scores of VDeH and competing hashing methods on benchmark datasets are presented. VDeH achieves state-of-the-art results compared to existing methods across all hash bit lengths tested (from 128 bits to 2048 bits). VDeH's mAP consistently increases with increasing hash bit length. Notably, besides VDeH, SH, and DSH, the performance of many existing methods on text datasets significantly decreases with increasing hash bit length. This is likely due to the complex density variations in text data. DSH and SH adapt by using more hash functions in dense distributions and fewer in sparse distributions; while VDeH adapts naturally because it creates smaller Voronoi units in dense distributions and larger units in sparse distributions. Furthermore, Figure 3 The PR curves for VDeH and the competing hash method are provided when the hash bit length is 512. We can observe that the VDeH curve (red) consistently dominates the curves of other methods on all datasets, demonstrating its superior precision and recall scores. Similar results can be observed for different hash bit lengths.

[0119] Unlike existing hashing methods that rely on optimizing the learning hash function, VDeH's training process is simple and straightforward. It only requires randomly sampling a given number of points from the training data, with each point corresponding to a nearest-neighbor-based hash function. Note that all methods require converting the training data into binary code using the hash function. In this experiment, the training time of multiple methods was tested on the GIST dataset, such as... Figure 4 As shown, VDeH has the shortest execution time for each training data size (from 10k to 1M), demonstrating its efficiency advantage over other methods. This is because VDeH does not require learning, while all other existing methods must.

[0120] Secondly, this application also provides a data retrieval system based on adaptive spatial partitioning and hash encoding, including:

[0121] The feature extraction module is used to extract feature vectors from the original data in the database and generate a high-dimensional feature vector set X∈R of the original data. d×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension of the feature vectors of the original data.

[0122] The feature vector space partitioning module is used to partition R based on a preset adaptive space partitioning method. d Repeat this process T times to generate T sets of R. d The characteristic space subunit;

[0123] The data point allocation module is used for each group of R d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-vectors for each data point, H(x) = [h1(x),...,h...]. i (x),...,h ψ [x]; where h is true if and only if x belongs to the i-th unit. i (x) = 1, and the rest are 0; ψ is the number of a set of feature space sub-units;

[0124] The encoding module is used to encode the set of T ψ-bit vectors H(x) based on a preset hash encoding mechanism, generate T ω-bit encoding blocks for each data point, and connect the T ω-bit encoding blocks of each data point to generate the L-bit binary hash code B(x) for each data point; where T = L / ω;

[0125] The distance calculation module is used to receive the data to be queried and convert it into an L-bit binary hash code B(q). It calculates the distance between B(q) and B(x) of each data point and takes the original data with the smallest distance as the target retrieval result.

[0126] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0127] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0128] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0129] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0130] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0131] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A data retrieval method based on adaptive spatial partitioning and hash coding, characterized in that, include: Extract feature vectors from the original data in the database to generate a high-dimensional feature vector set X∈R. d×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension of the feature vectors of the original data. R is partitioned based on a preset adaptive space partitioning method. d Repeat this process T times to generate T sets of R. d The characteristic space subunit; For each group R d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-position vectors H(x) = [h1(x),...,h2] for each data point. i (x),...,h ψ [x]; where h is true if and only if x belongs to the i-th unit. i (x) = 1, and the rest are 0; ψ is the number of a set of feature space sub-units; Based on a preset hash encoding mechanism, the set of T ψ-bit vectors H(x) is encoded to generate T ω-bit encoding blocks for each data point. The T ω-bit encoding blocks for each data point are then connected to generate the L-bit binary hash code B(x) for each data point; where T = L / ω. Receive the data to be queried and convert it into an L-bit binary hash code B(q). Calculate the distance between B(q) and B(x) of each data point, and take the original data with the smallest distance as the target retrieval result.

2. The method according to claim 1, characterized in that, R is partitioned based on a preset adaptive space partitioning method. d ,include: From R d ψ data points are randomly sampled from the corresponding high-dimensional feature vector set to generate a data point set D; Based on a preset spatial partitioning algorithm, R is divided according to ψ data points in D. d The space is divided into ψ feature space sub-units; among them, R is divided using a spatial partitioning algorithm. d Covering the entire R d In D, the data points are located at the center of their corresponding feature space sub-units.

3. The method according to claim 2, characterized in that, Spatial partitioning algorithms include Voronoi diagram method, KD tree-based method and variant partitioning method of ball tree.

4. The method according to claim 2, characterized in that, R d Data points assigned to R d The feature space subunits include: Calculate R d The distance between the data points and ψ data points in D, R d The data points are assigned to the feature space sub-units corresponding to the data points with the smallest distance in D.

5. The method according to claim 4, characterized in that, R d Data points assigned to R d The feature space subunit also includes: If R d When a data point is located at the boundary of a feature space sub-unit, the data point is assigned to any feature space sub-unit that shares that boundary.

6. The method according to claim 1, characterized in that, The set of T ψ bit vectors H(x) is encoded based on a preset hash encoding mechanism, including: Based on a pre-defined hash coding mechanism, the set of T ψ-bit vectors H(x) is converted into corresponding T ω-bit coded blocks b(x) = [e1(x), e2(x), ..., e ω (x)];where e j The formula for calculating (x) is as follows: Where j = 1, 2, ..., ω; ψ ≤ 2 ω .

7. The method according to claim 2, characterized in that, The formula for calculating the distance between the binary hash code of the data to be queried and the binary hash code of each data in the original data is as follows: in; B(x i Let x be the i-th data point. i The L-bit binary hash code B(x); b k (x i Let x be the i-th data point. i The k-th coded block; I(·) is the indicator function, when b k (x) and b k (q) Take 1 if they are not equal, otherwise take 0.

8. A data retrieval system based on adaptive spatial partitioning and hash coding, characterized in that, include: The feature extraction module is used to extract feature vectors from the original data in the database and generate a high-dimensional feature vector set X∈R of the original data. d×N Among them, R d N represents the feature vector space corresponding to the feature vectors of the original data, where one feature vector represents one data point; N is the dimension of the feature vectors of the original data. The feature vector space partitioning module is used to partition R based on a preset adaptive space partitioning method. d Repeat this process T times to generate T sets of R. d The characteristic space subunit; The data point allocation module is used to assign data points to each group of R. d The feature space subunit, R d Data points assigned to R d The feature space sub-unit generates a set of T ψ-position vectors H(x) = [h1(x),...,h2] for each data point. i (x),...,h ψ (x)];where h i (x)∈{0,1} is the i-th hash function for data point x if and only if x belongs to the i-th cell. i (x) = 1, and the rest are 0; ψ is the number of a set of feature space sub-units; The encoding module is used to encode the set of T ψ-bit vectors H(x) based on a preset hash encoding mechanism, generate T ω-bit encoding blocks for each data point, and connect the T ω-bit encoding blocks of each data point to generate the L-bit binary hash code B(x) for each data point; where T = L / ω; The distance calculation module is used to receive the data to be queried and convert it into an L-bit binary hash code B(q). It calculates the distance between B(q) and B(x) of each data point and takes the original data with the smallest distance as the target retrieval result.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.