Data search method and device, computer equipment and medium
By generating initial clusters and dividing them into sub-clusters, building a hierarchical index, and using the clustering types of non-leaf nodes and leaf nodes for search, the problem of low search efficiency in high-dimensional big data environments is solved, and efficient and accurate nearest neighbor search is achieved.
Patent Information
- Application Number
- CN202510724190.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing high-dimensional nearest neighbor search methods are inefficient in high-dimensional big data environments, especially when the data volume is huge and the feature vector dimension is high. The search efficiency of traditional methods is low, and the KMeans algorithm cannot automatically determine the number of clusters. The iterative calculation cost is high and the adaptability is poor.
By generating initial clusters and dividing them into sub-clusters, we ensure that the farthest data points of data points in the same sub-cluster are the same, build a hierarchical target cluster index, and use the clustering types of non-leaf nodes and leaf nodes for searching to avoid excessive search costs caused by too many sub-clusters.
It improves the accuracy and efficiency of search, adapts to data distribution patterns, reduces computational overhead, and improves the performance of high-dimensional nearest neighbor search.
Smart Images

Figure CN120653812A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data search technology, and in particular to a data search method, apparatus, computer equipment, and medium. Background Art
[0002] High-dimensional nearest neighbor search is a core problem in the field of unstructured data retrieval. Unstructured data refers to data that does not conform to a fixed format or schema. Unlike structured data (such as tables in a relational database), it cannot be organized and stored using predefined data models. Unstructured data comes in a variety of forms, including text, images, audio, video, and log files, making it difficult to process using traditional database management systems.
[0003] In the related technologies, a variety of technical paths have been proposed in the field of high-dimensional indexing, including methods based on tree-based space division and new scanning methods. However, these methods require multiple rounds of iterations, or are easily limited by data volume, resulting in low search efficiency. Summary of the Invention
[0004] The present application provides a data search method, apparatus, computer device, and medium to at least solve the problem of low search efficiency.
[0005] This application provides a data search method, including:
[0006] Acquire a target data set, and generate initial clusters based on the target data set;
[0007] Dividing the initial cluster into sub-clusters to obtain a target cluster index; wherein the farthest data points of the data points in the same sub-cluster are the same;
[0008] Based on the cluster type of the sub-cluster in the target cluster index, a target data point closest to the target query point in the target cluster index is searched and obtained, where the cluster type includes non-leaf nodes and leaf nodes.
[0009] The present application also provides a data search device, comprising:
[0010] An initialization module, configured to obtain a target data set and generate initial clusters based on the target data set;
[0011] A sub-clustering division module is used to divide the initial cluster into sub-clusters to obtain a target cluster index; wherein the farthest data point of the data points in the same sub-cluster is the same;
[0012] The search module is configured to search for a target data point closest to a target query point in the target cluster index based on a cluster type of a sub-cluster in the target cluster index, wherein the cluster type includes a non-leaf node and a leaf node.
[0013] The present application also provides a computer device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data search methods when executing the computer program.
[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data search methods are implemented.
[0015] The data search method provided in this embodiment includes obtaining a target data set and generating an initial cluster based on the target data set; dividing the initial cluster into sub-clusters to obtain a target cluster index; wherein the farthest data point of the data points in the same sub-cluster is the same; based on the cluster type of the sub-cluster in the target cluster index, searching for the target data point closest to the target query point in the target cluster index, the cluster type includes non-leaf nodes and leaf nodes. In order to avoid large differences in the scales of different clusters, this method further divides the larger cluster into sub-clusters until all clusters are of similar scale, thereby obtaining the target cluster index and ensuring the hierarchical nature of the index. Searching based on the type of sub-cluster can avoid excessive search costs caused by too many lower-level sub-clusters of the factor cluster, thereby improving the accuracy and efficiency of the search. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 A flowchart of a data search method provided in an embodiment of the present application;
[0018] Figure 2 A schematic diagram of data search provided in an embodiment of the present application;
[0019] Figure 3 A schematic diagram of a data search device provided in an embodiment of the present application;
[0020] Figure 4 A schematic diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0022] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0023] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0024] Nearest neighbor search is a technique for quickly finding the data points most similar to a query point in a dataset. It is widely used in fields such as recommender systems, image recognition, and data clustering. Its core goal is to retrieve the closest samples from massive amounts of data by efficiently calculating distances or similarities (such as Euclidean distance and cosine similarity). High-dimensional nearest neighbor search (HD-NNS) is a core problem in the field of unstructured retrieval. Unstructured data refers to data that does not conform to a fixed format or schema. Unlike structured data (such as tables in relational databases), it cannot be organized and stored using predefined data models. Unstructured data comes in a variety of forms, including text, images, audio, video, and log files, making it difficult to process using traditional database management systems. To address this problem, a common approach is feature extraction. Taking image recognition as an example, key points are extracted from an image as feature dimensions. Each feature dimension is represented by a numerical value, and all feature dimensions are aggregated to form a feature vector. Since vectors can measure distance, when two images are similar, the distance between their feature vectors (usually Euclidean distance) is small; otherwise, the distance is large. It's important to note that the feature vectors generated by the same feature extraction method typically have the same dimension, denoted by d. Each d-dimensional feature vector can be considered a point in d-dimensional space. Therefore, the unstructured data retrieval problem is transformed into a nearest neighbor search problem in d-dimensional space. Nearest neighbor search has two basic elements: the query point q and the dataset D. These two elements are derived from real-world unstructured data retrieval scenarios.
[0025] Take image retrieval as an example. Shopping apps allow users to upload an image to search for similar or related products. In this process, the image uploaded by the user serves as the query image, and all product images in the shopping app constitute the dataset. During the preparation phase, before the image search function is launched, features are extracted from all product images. These features form dataset D. After dataset D is prepared, the image search function is launched. The user then uploads the query image, and features are extracted to obtain the query point q, which is then used to perform a nearest neighbor search. In the current era of big data, data volumes are increasing in volume and complexity, posing corresponding challenges to nearest neighbor search, including the sheer volume and high dimensionality of feature vectors. This surge in data volume exacerbates the limitations of traditional linear scanning methods. Linear scanning determines nearest neighbors by calculating the distance between the query point and all data points one by one, and its time complexity is linear (O(n)). In practical applications, such as image recognition, where real-time matching is required from a large image library, the time required for linear scanning can be unacceptable, leading to system response delays or even system failure. This practical need has driven stringent requirements for the real-time performance of search algorithms, forcing researchers to break through the bottleneck of linear complexity.
[0026] The discovery of the Curse of Dimensionality reveals the unique properties of data distribution in high-dimensional space, which has become another shackle for algorithm design. As the feature dimension increases, the sparsity of data points in high-dimensional space increases significantly, and the distance metric gradually loses its discriminatory power. Taking image recognition as an example, if only a single low-dimensional feature is used to distinguish similar people, the discrimination is almost zero. Although the use of high-dimensional features can improve the representation ability, it increases the computational complexity of the algorithm. Early tree-based index structures based on space partitioning performed well in low-dimensional space. By recursively partitioning the data space, they quickly narrowed the search range and achieved accurate nearest neighbor retrieval. However, as the dimension increases, the performance of these methods drops sharply, and the search efficiency is low.
[0027] The implementation of nearest neighbor search technology is usually divided into two stages: high-dimensional index construction and search algorithm execution. Index construction, as the core task of the offline stage, aims to reduce the complexity of online search by rearranging or structuring data. For example, if a library stacks millions of books in disorder, finding the target book can only rely on linear traversal; if the books are arranged in alphabetical order by title, binary search can be used to achieve efficient retrieval with logarithmic complexity (O(log n)). Similarly, high-dimensional indexes establish spatial relationships or similarity structures between data through pre-calculation, providing navigation paths for online searches. However, the optimization of index construction requires a balance between storage cost and query efficiency. Especially in high-dimensional scenarios, how to design an index structure that can compress the data size while maintaining spatial locality becomes a key problem.
[0028] In the field of high-dimensional indexing, tree-based space partitioning methods were once considered the mainstream solution. For example, KMeans clustering (an unsupervised learning algorithm) was used to construct a hierarchical index structure. This method recursively divides data into cluster subsets and prioritizes neighboring clusters during queries to accelerate matching. By adaptively selecting the optimal algorithm combination, it demonstrates high efficiency for medium-dimensional data. However, the KMeans algorithm suffers from two inherent flaws that severely restrict its applicability to high-dimensional, large-scale data.
[0029] First, KMeans cannot automatically determine the number of clusters (k value) and must rely on manual pre-setting. This limitation causes problems in many aspects: First, users need to have prior knowledge of the data distribution or select the k value through trial and error, which is not only time-consuming and highly subjective, but also may lead to inconsistent clustering results due to differences in experience between different users; second, dynamic changes in the data set (such as incremental updates or distribution shifts) require flexible adjustment of the k value, and static preset k values are difficult to adapt to such scenarios, resulting in the failure of the index structure. For example, in a system with a continuously expanding image library, a fixed k value may make it impossible to reasonably cluster the newly added data, thereby reducing search accuracy and speed.
[0030] Second, the computational cost of multiple rounds of iterations of KMeans is high and it is easy to fall into local optimality. The algorithm approaches the optimal cluster center by alternating the "distribution-update" steps, and its convergence speed is significantly affected by the selection of the initial center. If the initial center distribution is not good, the algorithm may converge to a suboptimal solution, resulting in a decrease in clustering quality. In addition, the computational overhead of the iterative process for large-scale data is extremely high, especially when processing tens of millions of high-dimensional data. The time consumption of a single iteration is difficult to meet the needs of real-time index updates. What is more serious is that KMeans has extremely poor adaptability to dynamic data streams and cannot efficiently maintain the index structure in scenarios where data is frequently updated, limiting its application prospects in real-time systems. In summary, the limitations of current nearest neighbor search methods in high-dimensional big data environments are becoming increasingly prominent. Based on this, the present invention provides a data search method to improve the efficiency of nearest neighbor search.
[0031] According to an embodiment of the present invention, an embodiment of a data search method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0032] In this embodiment, a data search method is provided. Figure 1 is a flow chart of a data search method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0033] Step S101: Acquire a target data set and generate initial clusters based on the target data set.
[0034] The target dataset is composed of features, a collection of feature vectors, such as image features and text features. For example, if there are a large number of image samples, features are extracted from each image, and the image feature vectors form the target dataset. After obtaining the target dataset, data preprocessing can be performed, including filling missing values, removing outliers, and normalizing the raw data to ensure consistent dimensionality across dimensions and avoid bias in distance calculations.
[0035] Add the target dataset D as the initial cluster to the set After initialization operation, the initial clustering is obtained
[0036]
[0037] Step S102: dividing the initial cluster into sub-clusters to obtain a target cluster index.
[0038] Data points in the same sub-cluster share the same farthest data point. Based on the initial cluster, sub-clusters are recursively divided, resulting in finer-grained sub-clusters. This results in a hierarchical index structure, the target cluster index. A size constraint is first imposed by setting a quantity threshold to ensure that the size of the sub-clusters is less than or equal to the quantity threshold. This quantity threshold can be set based on hardware cache capacity or query latency requirements.
[0039] After obtaining the initial cluster, let i = 1, for any set of data points (sub-cluster) in the initial cluster (in, Represents the i-th data point set), compare the size of the data point set with the quantity threshold, and the size of the data point set depends on the amount of data contained in the data point set. If the size of the data point set is less than or equal to the quantity threshold, no further division is required. If the size of the data point set is greater than the quantity threshold, the data point set needs to be further divided. First, calculate the farthest data point of each data point in the data point set, and classify each data point in the data point set by the farthest data point. Specifically, all data points in the data point set with the same farthest data point are classified into the same sub-cluster. A new sub-cluster is obtained by this division method, and the farthest point can be used as the representative point corresponding to the sub-cluster. Repeat the above operation for the newly generated sub-clusters until all sub-clusters meet the constraints. Suppose f sub-clusters are obtained, and f sub-clusters are added to the initial cluster. At the end. All sub-clusters obtained by division are recorded in The position number after and the representative point of each sub-cluster, so as to maintain All subclusters and the initial cluster constitute the target cluster index with a hierarchical structure.
[0040] Step S103 : Based on the cluster type of the sub-cluster in the target cluster index, search and obtain the target data point closest to the target query point in the target cluster index.
[0041] Cluster types include non-leaf nodes and leaf nodes. This step searches the target cluster index for the target data point closest to the target query point q. The target query point q is a feature vector. For example, in image search, if a target image is input, the target data point q is the feature vector obtained after feature extraction of the target image.
[0042] Before starting the search, a priority queue can be initialized to manage the cluster nodes to be processed. The priority of each element in the queue can be determined by the lower bound of the distance from the target query point to the cluster node. The smaller the distance lower bound, the higher the priority of the node. When the search starts, the root node of the target cluster index (i.e., the initial cluster) can be added to the queue, and the minimum possible distance between the target query point and the cluster center corresponding to the root node is calculated as the priority. Then enter the iterative processing stage, pop the node with the highest current priority from the queue for expansion, and the node can be considered a subcluster. Determine the cluster type of the subcluster. If the cluster type is a non-leaf node, further query the lower sub-clusters of the subcluster until the subcluster type is a leaf node. Specifically, when querying the lower-level sub-clusters of the sub-cluster of a non-leaf node, if the sub-cluster of the non-leaf node includes multiple lower-level sub-clusters of the same level, the distance between the target query point and the representative point corresponding to each lower-level sub-cluster can be calculated separately to obtain the representative point of the lower-level sub-cluster farthest from the target query point, and the lower-level sub-cluster corresponding to the representative point is used as the sub-cluster for the next query.
[0043] If the sub-cluster is a leaf node, query the data point closest to the target query point in the sub-cluster of the leaf node. When the search ends, the algorithm returns the data point with the smallest distance in the candidate list as the target data point of the target query point.
[0044] Optionally, the sub-cluster of the non-leaf node may store relevant sub-cluster division information, including representative point information of the sub-cluster and references to the sub-nodes.
[0045] This step distinguishes cluster types, efficiently searches for and locates the target data point closest to the target query point, and dynamically optimizes the search path to reduce computational overhead based on the hierarchical characteristics of the target cluster index and the combination of navigation information of non-leaf nodes and actual data of leaf nodes.
[0046] The data search method provided by the present invention is described below in conjunction with specific scenarios. Taking image search as an example, for a given target image, it is necessary to search for the image closest to the target image from the image library. Feature extraction is performed on the images in the image library in advance, and the feature vectors of each image constitute the target data set, which is used as the initial clustering. After obtaining the initial clustering, sub-clusters are divided, specifically constrained by the size of the sub-clusters. If the number of data points in the sub-cluster is less than or equal to the quantity threshold, there is no need to divide the sub-cluster; if the number of data points in the sub-cluster is greater than the quantity threshold, the sub-cluster is further divided, and the data points in the same lower-level sub-cluster obtained by division have the same farthest data point, until the size of the sub-cluster obtained by division meets the requirements, and the target cluster index is obtained.
[0047] Extract the feature vector of the target image to obtain the target query point, and query the data point closest to the target query point in the target cluster index as the target data point. The target data point can be a feature vector, and the image corresponding to the feature vector is the result of the image search requirement.
[0048] The data search method provided in this embodiment includes obtaining a target data set and generating an initial cluster based on the target data set; dividing the initial cluster into sub-clusters to obtain a target cluster index; wherein the farthest data point of the data points in the same sub-cluster is the same; based on the cluster type of the sub-cluster in the target cluster index, searching for the target data point closest to the target query point in the target cluster index, the cluster type includes non-leaf nodes and leaf nodes. In order to avoid large differences in the scales of different clusters, this method further divides the larger cluster into sub-clusters until all clusters are of similar scale, thereby obtaining the target cluster index and ensuring the hierarchical nature of the index. Searching based on the type of sub-cluster can avoid excessive search costs caused by too many lower-level sub-clusters of the factor cluster, thereby improving the accuracy and efficiency of the search.
[0049] In this embodiment, a data search method is provided, which includes the following steps:
[0050] Step S201: Acquire a target data set, and generate initial clusters based on the target data set.
[0051] For details, please see Figure 1 Step S101 of the illustrated embodiment will not be described in detail here.
[0052] Step S202: dividing the initial cluster into sub-clusters to obtain a target cluster index.
[0053] Specifically, step S202 includes:
[0054] Step S2021: classify at least one set of data points in the initial cluster based on a preset quantity threshold.
[0055] A preset number threshold can be used to constrain the size of the sub-clusters to ensure that the size of the sub-clusters is less than or equal to the number threshold. The preset number threshold can be set based on hardware cache capacity or query latency requirements. The data point set is classified by comparing the number of data points in the set with the preset number threshold.
[0056] Further, step S2021 includes: determining a first data point set whose number of data points is greater than a preset number threshold as a non-leaf node; and determining a second data point set whose number of data points is less than or equal to the preset number threshold as a leaf node.
[0057] The size of each data point set is compared with a preset number threshold. The first data point set with more data points than the preset number threshold is determined to be a non-leaf node, indicating that it needs to be further divided into substructures to refine the index hierarchy. The second data point set with less than or equal to the preset number threshold is marked as a leaf node, as a clustering unit that cannot be further divided.
[0058] This classification ensures the balance of the index structure and avoids the decline in search efficiency caused by a single cluster being too large or index depth redundancy caused by a single cluster being too small.
[0059] Step S2022: performing sub-clustering division on the first data point set whose cluster type is non-leaf nodes, and performing incremental processing on the second data point set whose cluster type is leaf nodes, to obtain a target cluster index.
[0060] For the first set of data points at non-leaf nodes, a specific sub-clustering algorithm is used to recursively decompose them into several sub-clusters. Each sub-cluster inherits some of the data from the parent node and forms a new index branch. During the partitioning process, sub-cluster information, such as representative points, is recorded. For the second set of data points at leaf nodes, incremental processing is performed to directly generate clustering results.
[0061] Specifically, the first data point set whose cluster type is non-leaf node is divided into sub-clusters, including: searching for the farthest data point of each data point in the first data point set; dividing the data points with the same farthest data point into the same sub-cluster, and the representative point of the sub-cluster is the corresponding farthest data point.
[0062] Traverse each data point within the first data point combination to determine the farthest data point within the set. The distance calculation method is not limited and can be determined by calculating Euclidean distance, cosine similarity, etc. Data points with the same farthest data point are divided into the same sub-cluster, and the common farthest data point is used as the representative point of the sub-cluster.
[0063] This division method makes the points of each sub-cluster naturally have consistent spatial distribution, thus forming a geometric structure centered on the representative point and covering a specific directional area, laying the foundation for subsequent hierarchical index search pruning.
[0064] Step S203 : Based on the cluster type of the sub-cluster in the target cluster index, search and obtain the target data point closest to the target query point in the target cluster index.
[0065] Specifically, step S203 includes:
[0066] Step S2031, incrementally searching for subclusters starting from the root cluster of the target cluster index;
[0067] Step S2032, determining the cluster type of the sub-cluster;
[0068] Step S2033: If there is a first sub-cluster whose cluster type is a leaf node, search the distance between the target query point and all data points in the first sub-cluster to obtain the target data point closest to the target query point.
[0069] The root cluster of the target cluster index is the cluster corresponding to the complete target dataset. The search starts from the root cluster and traverses the target cluster index layer by layer. The search path is selected based on the cluster type, and the current search pointer is maintained for hierarchical progression. By reading the tag information stored in each sub-cluster, it is determined whether the sub-cluster contains a structure that can be further divided or only stores the original sub-cluster. Cluster types include leaf nodes and non-leaf nodes.
[0070] If the first sub-cluster marked as a leaf node is found, all data points in the sub-cluster are traversed, and the distance between each data point and the target query point (such as Euclidean distance, cosine similarity) is calculated one by one, so as to determine the data point with the smallest distance to the target query point and use it as the target data point.
[0071] Step S2034: If there is a second sub-cluster whose cluster type is a non-leaf node, search the second sub-cluster for a third sub-cluster that is farthest from the target query point.
[0072] If a second sub-cluster marked as a non-leaf node is found, it indicates that the second sub-cluster still has branches that can be further searched. There may be multiple second sub-clusters. Therefore, the third sub-cluster farthest from the target query point is searched in the second sub-cluster to search for the target data point closest to the target query point in the third sub-cluster.
[0073] Furthermore, step S2034 includes: if there is a second sub-cluster whose cluster type is a non-leaf node, calculating the distance between the target query point and the representative point corresponding to the second sub-cluster; and determining the second sub-cluster corresponding to the representative point farthest from the target query point as the third sub-cluster.
[0074] The distance between the target query point and the representative point corresponding to each second sub-cluster is calculated, and the second sub-cluster corresponding to the representative point with the farthest distance is the third sub-cluster.
[0075] Step S2035: If the cluster type of the third sub-cluster is a leaf node, the distance between the target query point and all data points in the third sub-cluster is searched to obtain the target data point closest to the target query point.
[0076] The cluster type of the third sub-cluster is further determined. If the third sub-cluster is a leaf node, each data point in the third sub-cluster is traversed to calculate the distance between each data point and the target query point, thereby determining the target data point closest to the target query point.
[0077] If the third sub-cluster is a non-leaf node, the search for the lower sub-clusters of the third sub-cluster can continue until the target data point closest to the target query point is determined.
[0078] The data search method provided in this embodiment can automatically decide whether to continue dividing into sub-clusters according to the cluster size, can well adapt to the actual distribution pattern of data, and improve the efficiency and accuracy of nearest neighbor search.
[0079] In this embodiment, a data search method is provided that can be used for nearest neighbor search. First, the furthest neighbors of all member points in the data set are searched, and all data points with the furthest neighbors are grouped into the same cluster, thereby obtaining a target cluster index. Based on the target cluster index, the nearest target data point of a given target query point is searched. To prevent large differences in the scale of different clusters, the method further divides the larger cluster into subclusters until all clusters are of similar size, thereby obtaining a hierarchical cluster index. The complete process of this method is described below, and the method includes the following:
[0080] Algorithm 1: Generate a hierarchical clustering index for dataset D based on distant neighbor search.
[0081] Step a1, let Store all the clusters generated in the end and initialize them to get the initial clusters
[0082] Step a2, initializing i=1;
[0083] Step a3, if the i-th data point set If the number of data points in is greater than the preset threshold value τ, then the mark If it is a non-leaf node, skip to step a4 to divide it into sub-clusters; otherwise, Mark it as a leaf node, no longer divide it into sub-clusters, and jump to step a5;
[0084] Step a4 is Divide into subclusters;
[0085] Wherein, step a4 includes:
[0086] Step a41, for each The data point x in Search for the farthest neighbor of x in the search, the specific process includes calculating the data point x to The data point with the largest distance to the data point x is the farthest neighbor;
[0087] Step a42, in In , all data points with the same farthest neighbor are divided into the same sub-cluster, and each farthest neighbor is used as the representative point of the corresponding sub-cluster. Suppose that f sub-clusters are obtained, and these f sub-clusters are appended to end;
[0088] Step a43, in Record the above f sub-clusters in The position number after and the representative points of these f sub-clusters, so as to maintain to the connections between its subclusters;
[0089] Step a5, let i=i+1;
[0090] Step a6, if i is greater than The number of clusters in the cluster indicates that the clustering is completed. Return as the clustering result; otherwise, jump to step a3.
[0091] Algorithm 2: Given a target query point q, search for the nearest neighbor of q.
[0092] Step b1: Let S represent the cluster to be searched, and initialize
[0093] Step b2: Determine whether S is a leaf node. If S is a non-leaf node, jump to step b3; if S is a leaf node, jump to step b4;
[0094] Step b3: Calculate the distance from the target query point q to the representative points of all sub-clusters in S, find the sub-cluster C where the representative point with the farthest distance is located, set S = C, and jump to step b2;
[0095] Step b4, calculate the distance from the target query point q to all data points in S, and return the nearest point as the nearest neighbor result (i.e., the target data point).
[0096] The data search method provided by the present invention automatically determines the number of clusters by aggregating the farthest data points from all points. Therefore, it requires no human intervention or iteration, and the number of clusters can be automatically determined. This clustering method aligns with data distribution patterns and automatically determines whether to further sub-clusters based on cluster size. This method effectively adapts to actual data distribution patterns and improves the efficiency and accuracy of nearest neighbor searches.
[0097] The following is an analysis of the effectiveness of the present invention. Table 1 shows the number of farthest neighbors within several typical high-dimensional datasets. As can be seen, the proportion of farthest neighbors is very small. This means that the number of clusters will not be too large each time a cluster is divided. This results in a small number of representative points whose distance to q needs to be calculated during each search, resulting in a very high computational efficiency.
[0098] Table 1: The number of farthest neighbors in the dataset:
[0099]
[0100] Figure 2 Figure 1 is a diagram of a data search. Points F1, F2, F3, A, B, and C in the figure are all points in dataset D. Points F1, F2, and F3 are cluster representatives. After clustering, A and B are in the same cluster, with F1 as their farthest neighbor. C is in another cluster, with F2 as its farthest neighbor. This shows that points that are close together are more likely to share the same farthest neighbor. During the search, the distances from query point q to all cluster representatives are calculated. The cluster representative farthest from query point q is F1. Since cluster F1 is a leaf node, the distances from all member points in cluster F1 to q are then calculated to find q's nearest neighbor, A.
[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0102] The embodiment of the present application also provides a data search device, such as Figure 3 Shown, including:
[0103] An initialization module, configured to obtain a target data set and generate initial clusters based on the target data set;
[0104] A sub-clustering division module is used to divide the initial cluster into sub-clusters to obtain a target cluster index; wherein the farthest data point of the data points in the same sub-cluster is the same;
[0105] The search module is configured to search for a target data point closest to a target query point in the target cluster index based on a cluster type of a sub-cluster in the target cluster index, wherein the cluster type includes a non-leaf node and a leaf node.
[0106] In some optional implementations, the sub-clustering module includes:
[0107] a classification unit, configured to classify at least one set of data points in the initial cluster based on a preset quantity threshold;
[0108] The cluster division unit is used to perform sub-cluster division on a first data point set whose cluster type is a non-leaf node, and perform incremental processing on a second data point set whose cluster type is a leaf node to obtain a target cluster index.
[0109] In some optional embodiments, the classification unit includes:
[0110] A first subunit is configured to determine a first data point set whose number of data points is greater than the preset number threshold as a non-leaf node;
[0111] The second subunit is configured to determine a second data point set whose number of data points is less than or equal to the preset number threshold as a leaf node.
[0112] In some optional implementations, the clustering division unit includes:
[0113] a data point search subunit, configured to search the first data point set for the farthest data point of each data point;
[0114] The data point division subunit is used to divide the data points with the same farthest data point into the same sub-cluster, and the representative point of the sub-cluster is the corresponding farthest data point.
[0115] In some optional implementations, the search module includes:
[0116] A search unit, configured to incrementally search for subclusters starting from a root cluster of the target cluster index;
[0117] A type judgment unit, configured to judge the cluster type of the sub-cluster;
[0118] The first judgment unit is configured to, if the cluster type of the first sub-cluster is a leaf node, search the distance between the target query point and all data points in the first sub-cluster to obtain the target data point closest to the target query point.
[0119] In some optional implementations, the search module includes:
[0120] a second judgment unit, configured to search, if a cluster type of a second sub-cluster exists that is a non-leaf node, for a third sub-cluster that is farthest from the target query point in the second sub-cluster;
[0121] The data point query unit is configured to search for the distance between the target query point and all data points in the third sub-cluster if the cluster type of the third sub-cluster is a leaf node, and obtain the target data point closest to the target query point.
[0122] In some optional implementations, the second judgment unit includes:
[0123] a distance calculation subunit, configured to calculate the distance between the target query point and the representative point corresponding to the second subcluster if the cluster type of the second subcluster is a non-leaf node;
[0124] The sub-cluster determining sub-unit is configured to determine the second sub-cluster corresponding to the representative point farthest from the target query point as the third sub-cluster.
[0125] For the description of the features in the embodiment corresponding to the data search device, reference can be made to the relevant description of the embodiment corresponding to the data search method, which will not be repeated here.
[0126] The embodiment of the present application also provides a computer device, such as Figure 4 As shown, it includes a memory 10 and a processor 20, wherein the memory 10 stores a computer program, and the processor 20 is configured to run the computer program to execute the steps in any of the above data search method embodiments.
[0127] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data search method embodiments when run.
[0128] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0129] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned data search method embodiments are implemented.
[0130] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned data search method embodiments are implemented.
[0131] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0132] The above is a detailed introduction to a data search method, device, computer equipment and medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A data search method, characterized in that: include: Acquire a target data set, and generate initial clusters based on the target data set; Dividing the initial cluster into sub-clusters to obtain a target cluster index; wherein the farthest data points of the data points in the same sub-cluster are the same; Based on the cluster type of the sub-cluster in the target cluster index, a target data point closest to the target query point in the target cluster index is searched and obtained, where the cluster type includes non-leaf nodes and leaf nodes.
2. The method according to claim 1, characterized in that The sub-clustering of the initial cluster to obtain a target cluster index includes: Classifying at least one set of data points in the initial cluster based on a preset number threshold; Sub-clustering is performed on a first set of data points whose cluster type is a non-leaf node, and incremental processing is performed on a second set of data points whose cluster type is a leaf node to obtain a target cluster index.
3. The method according to claim 2, characterized in that The classifying at least one set of data points in the initial cluster based on a preset quantity threshold comprises: Determine a first data point set whose number of data points is greater than the preset number threshold as a non-leaf node; A second data point set whose number of data points is less than or equal to the preset number threshold is determined as a leaf node.
4. The method according to claim 2, characterized in that The sub-clustering division of the first data point set whose cluster type is non-leaf nodes includes: Searching for the farthest data point for each data point in the first set of data points; Data points with the same farthest data point are divided into the same sub-cluster, and the representative point of the sub-cluster is the corresponding farthest data point.
5. The method according to claim 1, wherein The searching for a target data point closest to a target query point in the target cluster index based on the cluster type of the sub-cluster in the target cluster index includes: Incrementally searching for subclusters starting from the root cluster of the target cluster index; Determining the cluster type of the sub-cluster; If there is a first sub-cluster whose cluster type is a leaf node, the distance between the target query point and all data points in the first sub-cluster is searched to obtain the target data point closest to the target query point.
6. The method according to claim 5, characterized in that After determining the cluster type of the sub-cluster, searching for a target data point closest to a target query point in the target cluster index based on the cluster type of the sub-cluster in the target cluster index further includes: If there is a second sub-cluster whose cluster type is a non-leaf node, searching the second sub-cluster for a third sub-cluster that is farthest from the target query point; If the cluster type of the third sub-cluster is a leaf node, the distance between the target query point and all data points in the third sub-cluster is searched to obtain the target data point closest to the target query point.
7. The method according to claim 6, characterized in that If the cluster type of the second sub-cluster is a non-leaf node, searching the second sub-cluster for a third sub-cluster farthest from the target query point includes: If there is a second sub-cluster whose cluster type is a non-leaf node, calculate the distance between the target query point and the representative point corresponding to the second sub-cluster; The second sub-cluster corresponding to the representative point farthest from the target query point is determined as the third sub-cluster.
8. A data search device, characterized in that: include: An initialization module, configured to obtain a target data set and generate initial clusters based on the target data set; A sub-clustering division module is used to divide the initial cluster into sub-clusters to obtain a target cluster index; wherein the farthest data point of the data points in the same sub-cluster is the same; The search module is configured to search for a target data point closest to a target query point in the target cluster index based on a cluster type of a sub-cluster in the target cluster index, wherein the cluster type includes a non-leaf node and a leaf node.
9. A computer device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data search method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the data search method according to any one of claims 1 to 7 are implemented.