Index construction method and device
By dividing the data into multiple data fragments and constructing sub-indexes, the problem of index construction time increasing with the data scale is solved, and efficient resource utilization and index construction efficiency are achieved.
Patent Information
- Application Number
- CN202510137934.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-09
AI Technical Summary
In the prior art, the index construction time increases with the increase in data scale, making it difficult to effectively expand to a large-scale distributed environment, resulting in low hardware resource utilization.
The data is divided into the target number of data shards based on multiple centroids of the target data, and the index generator is allocated for each data shard to build a sub-index, and finally the sub-index is integrated into the target index.
It realizes efficient and balanced resource utilization, improves the efficiency of sub-index construction, and shortens the time-consuming construction of target indexes.
Smart Images

Figure CN119961487A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present specification relate to the field of computer technology, and more particularly to an index construction method and device. Background Art
[0002] With the development of computer technology, indexing technology has been continuously improved. Indexing can be applied to recommendation systems, information retrieval, computer vision, and natural language processing. However, as the size of the data set continues to expand, the time to build the index is also getting longer and longer.
[0003] In the prior art, the approximate nearest neighbor search method is generally used to complete index construction. However, with the increase of data scale, the index construction time will also increase accordingly. Traditional index construction usually relies on a single machine or a small-scale cluster, which is difficult to effectively expand to a large-scale distributed environment, resulting in low hardware resource utilization. Therefore, a more effective index construction method is urgently needed to solve the above problems. Summary of the invention
[0004] In view of this, an embodiment of the present specification provides an index construction method. One or more embodiments of the present specification also relate to an index construction device, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.
[0005] According to a first aspect of an embodiment of this specification, there is provided an index construction method, including:
[0006] Dividing the target data into a target number of data shards based on multiple centroids of the target data, wherein target sub-data in the target data is divided into at least two data shards;
[0007] Allocate an index generator to the data set corresponding to each data shard in the target number of data shards, and use the index generator to construct a sub-index for each data set;
[0008] The sub-index of each data set is integrated into a target index corresponding to the target data.
[0009] According to a second aspect of an embodiment of this specification, there is provided an index construction device, including:
[0010] a partitioning module, configured to partition the target data into a target number of data fragments based on a plurality of centroids of the target data, wherein target sub-data in the target data is partitioned into at least two data fragments;
[0011] A construction module is configured to allocate an index generator to a data set corresponding to each data shard in the target number of data shards, and use the index generator to construct a sub-index of each data set;
[0012] The integration module is configured to integrate the sub-index of each data set into a target index corresponding to the target data.
[0013] According to a third aspect of an embodiment of this specification, a computing device is provided, including:
[0014] Memory and processor;
[0015] The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the above-mentioned index construction method are implemented.
[0016] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the above-mentioned index construction method are implemented.
[0017] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions, which implement the steps of the above-mentioned index construction method when executed by a processor.
[0018] An index construction method provided by an embodiment of the present specification divides the target data into a target number of data shards based on multiple centroids of the target data, that is, divides the target sub-data in the target data into at least two data shards. An index generator is allocated to the data set corresponding to each data shard in the target number of data shards, and the sub-index of each data set is constructed using the index generator, making full use of the index generator to achieve efficient and balanced utilization of resources, while improving the efficiency of constructing the sub-index. The sub-index of each data set is integrated into the target index corresponding to the target data, shortening the time consumption of constructing the target index. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic diagram of an index construction method provided by an embodiment of this specification;
[0020] Figure 2 is a flowchart of an index construction method provided by an embodiment of this specification;
[0021] Figure 3 is a process flow chart of a method for constructing a graph index provided by an embodiment of this specification;
[0022] Figure 4 This is a data preprocessing and sharding diagram of an index construction method provided by an embodiment of this specification;
[0023] Figure 5It is a schematic diagram of index construction of an index construction method provided by an embodiment of this specification;
[0024] Figure 6 It is a structural schematic diagram of an index building device provided by an embodiment of this specification;
[0025] Figure 7 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION
[0026] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.
[0027] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a", "said" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0028] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0029] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0030] First, the terms involved in one or more embodiments of this specification are explained.
[0031] ANNS: Approximate Nearest Neighbor Search, an algorithm for finding the data point closest to the query point in a high-dimensional data set, is used to improve the speed and efficiency of large-scale data retrieval.
[0032] DiskANN: Disk-based Approximate Nearest Neighbor, an algorithm for large-scale vector search. It achieves scalability, high accuracy, and low cost through efficient disk and memory management.
[0033] MapReduce: A programming model for processing large-scale data sets. It improves data processing efficiency by dividing the task into two phases, Map and Reduce, and executing them in parallel.
[0034] K-means clustering: A common clustering algorithm that divides the data set into several clusters, each represented by a centroid, for data sharding and load balancing.
[0035] Centroid: In clustering algorithms, the center point of each cluster represents the average position of all data points in the cluster.
[0036] Sharding: Sharding is a technology that divides large data sets into multiple smaller data blocks, each of which is processed on an independent computing node to improve parallel processing capabilities and system performance.
[0037] Min Heap: It is a sorted complete binary tree in which the data value of any non-terminal node is not greater than the value of its left and right child nodes.
[0038] Load Balancing: Load Balancing refers to the reasonable allocation of tasks and resources among multiple computing nodes to ensure that the load of each node is relatively balanced and avoid overloading of individual nodes.
[0039] Local Index: Local Index, the data index structure within each shard, is used to perform efficient retrieval operations on local data sets.
[0040] Global Index: Global Index is an overall data index structure that merges multiple local indexes to support global data retrieval across shards.
[0041] Approximate nearest neighbor search (ANNS) has been widely used in recommendation systems, information retrieval, computer vision, and natural language processing. For example, in image and video review systems, billions of data points are often required to perform fast similarity retrieval to detect duplicate content or abnormal data. As the size of the dataset increases (e.g., from 1 billion to 10 billion data points), the index construction time becomes unacceptable. In actual industrial environments, traditional index construction time can be as long as several days, which is difficult to meet the needs of daily updates. Uneven data distribution, waste of resources, and low utilization of memory and computing resources further aggravate the bottleneck of the data scale that large-scale retrieval systems can support.
[0042] Although existing index construction methods (such as DiskANN) can effectively support large-scale vector retrieval, their construction time is often unacceptable as the data scale increases. On a data set of tens of billions, the index construction time may take several days or even longer, which cannot meet the needs of large-scale real-time applications. Traditional index construction usually relies on a single machine or a small-scale cluster, which is difficult to effectively expand to a large-scale distributed environment, resulting in low utilization of hardware resources (such as memory, SSD). In addition, due to unbalanced sharding, some computing nodes are overloaded, resulting in resource waste and performance bottlenecks. When using traditional algorithms such as K-means for data sharding, it is easy for some shards to contain too much data, resulting in uneven load on computing nodes, further causing memory overflow (OOM) problems. This problem significantly increases the complexity and risk of index construction. Therefore, one or more embodiments of this specification provide an index construction method to solve the above problems.
[0043] Figure 1A schematic diagram of an index construction method provided according to an embodiment of the present specification is shown. When constructing a target index, the target data is obtained, the centroid of the target data is determined, and the target sub-data in the target data is divided into at least two data shards based on the centroid, that is, the target data is divided into n data shards, such as data shard 1, data shard 2...data shard n-1 and data shard n, and an index generator is allocated to the data set corresponding to each data shard in the target number of data shards, and the sub-index of each data set is constructed using the index generator. Make full use of the index generator to achieve efficient and balanced utilization of resources, while improving the efficiency of constructing sub-indexes. In practical applications, data shards 1-data shard n can be evenly distributed to index generators 1-index generators m for processing. For example, when the number of data shards is 100 and the number of index generators is 5, 100 data shards can be evenly distributed to 5 index generators. Each index generator generates a sub-index corresponding to the received data shard to obtain sub-indexes 1-n. Integrating the sub-indexes of n data sets into a target index corresponding to the target data shortens the time spent on building the target index. In this specification, an index building method is provided, and this specification also relates to an index building device, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail one by one in the following embodiments.
[0044] See also Figure 2 , Figure 2 A flowchart of an index construction method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.
[0045] Step 202: Divide the target data into a target number of data slices based on multiple centroids of the target data, wherein target sub-data in the target data is divided into at least two data slices.
[0046] Specifically, the target data can be any data that needs to be queried based on an index, such as commodity data or order data in an e-commerce system, data to be searched in a search engine, and user data or data to be recommended in a recommendation system. Specifically, the target data can also be media content data (including one or more of video, image, text, and audio) in a content distribution service or a resource browsing service. Resource browsing services can be browsing services such as video, graphics, and text, and resource browsing services can provide video, graphics, and text publishing functions and browsing functions. Data shards can be data blocks, and each data block is processed on an independent computing node to improve parallel processing capabilities and system performance. The centroid is the center point of a cluster, indicating the average position of all data points in the cluster. The multiple centroids of the target data are the center points corresponding to the multiple cluster clusters after the target data is clustered to obtain multiple cluster clusters. The target sub-data is the data contained in the set of target data. Clustering the target data is actually clustering the target sub-data and dividing the target sub-data into at least two data shards.
[0047] In practical applications, target data can also be medical data generated in the medical field, that is, clinical medical data, data collected by medical equipment, and real-time physiological data of patients, such as heart rate, blood pressure, blood sugar level, etc. Target data can also be logistics data generated in the logistics field, such as logistics flight data, logistics vehicle data, and cargo location data and status data. By analyzing physical data, it is possible to ensure that the cargo arrives at the destination safely and on time, while optimizing inventory management and transportation route planning.
[0048] Based on this, after obtaining the target data, the target sub-data in the target data data set can be clustered to obtain multiple clusters, and the centroids corresponding to the multiple clusters are the multiple centroids of the target data. Based on the multiple centroids of the target data, the target data is divided into a target number of data shards, which means that the target sub-data in the target data is divided into at least two data shards to facilitate the subsequent construction of the target index.
[0049] In practical applications, before the target data is divided based on the centroid, the target sub-data in the target data can be scattered and sorted. The data is scattered, partitioned and sorted according to the key (such as ID) of the target sub-data, so that the target sub-data with the same key are divided into one partition, and then the target sub-data are divided into at least two data shards according to the distance between each target sub-data and the centroid.
[0050] In one or more embodiments provided in this specification, before dividing the target sub-data into at least two data fragments based on the centroid of the target data, the method further includes:
[0051] A clustering algorithm is used to divide a plurality of target sub-data in the target data into a plurality of clusters, and a plurality of centroids corresponding to the plurality of clusters are used as a plurality of centroids of the target data.
[0052] The target sub-data is the vector expression of the original data. The original target sub-data can be converted by using algorithms such as principal component analysis and feature extraction to convert the target sub-data into a vector expression. The target sub-data can be clustered using a variety of clustering algorithms such as the K-means clustering algorithm. When clustering the target sub-data, clustering can be performed based on at least two centroids, and the data corresponding to each centroid is clustered into a cluster cluster.
[0053] It can be understood that the target data set includes multiple target sub-data. The original target sub-data is converted into target sub-data in vector expression form by using algorithms such as principal component analysis. The target sub-data is clustered by using a clustering algorithm, so that the target data is divided into at least two clusters. The centroids of at least two clusters are used as the centroids of the target data.
[0054] In practical applications, after clustering the target sub-data, the target data can be divided into multiple clusters, and the distance between the data points in each cluster and the centroid of the cluster is small, thereby ensuring that the data similarity in each cluster is high.
[0055] To summarize, by clustering the target sub-data, at least two data with high similarity can be divided into one cluster cluster, and at least two cluster clusters corresponding to the target data are obtained by clustering. The centroids of the at least two cluster clusters are used as the centroids of the target data, and subsequent data processing is performed on the at least two cluster clusters, which can reduce the computational pressure of unified data processing.
[0056] In one or more embodiments provided in this specification, dividing the target data into at least two data slices based on the centroid of the target data includes:
[0057] Partition the target data according to the key of the target sub-data to obtain the data to be divided corresponding to at least two partitions, so that the data to be divided of the same partition can be sent to the same Reducer for subsequent sharding processing; determine the target number of data shards and the data storage parameters corresponding to each data shard, the data storage parameters are determined based on the repetition factor, the data volume of the target data and the target number; for the sub-data to be divided in the data to be divided corresponding to each partition, divide the sub-data to be divided into at least two data shards according to the data storage parameters and the distance information between the sub-data to be divided and each centroid of the target data. It can be understood that after the target sub-data is divided into the corresponding partition, it is the sub-data to be divided in the data to be divided of the partition.
[0058] Among them, the target data data set contains multiple target sub-data, and the target sub-data can be vector data, which is stored in key-value pairs. The key of each target sub-data corresponds to an identification information, and the identification information can be the ID of the target sub-data, and the value is vector data. The target sub-data can be partitioned according to the key of the target sub-data, and the target sub-data of the key can be divided into one partition to achieve the purpose of ID deduplication. The data storage parameter refers to the data storage threshold of the data shard, which indicates the maximum data storage capacity of the data shard. The data storage parameter can be determined based on the repetition factor, the data volume of the target data, and the target number (i.e., the total number of shards). The repetition factor indicates the maximum number of shards to which each target sub-data can be allocated. For example, when the repetition factor is 3, it means that each target sub-data can be allocated to up to 3 data shards.
[0059] Based on this, the target data contains at least two target sub-data in the form of key-value pairs, where the key of the target sub-data is identification information and the value is vector data. By partitioning the target sub-data according to the identification information of the target sub-data, the target sub-data with the same identification information can be divided into one partition, so as to achieve the purpose of deduplication of identification, and finally obtain the data to be divided corresponding to at least two partitions. Determine the data storage parameters corresponding to the target number of data shards. Within the range of the data storage parameters, the sub-data to be divided is divided into at least two data shards according to the distance information between the sub-data to be divided and the centroid in the data to be divided, so that the amount of data of the sub-data to be divided divided in each data shard is relatively balanced.
[0060] Continuing with the above example, the target data contains target sub-data in a key-value pair structure, where the key of the target sub-data is ID and the value is vector data. The target data is scattered and sorted, and the target sub-data is partitioned according to the key of the target sub-data. The target sub-data with the same ID is divided into one partition. The number of partitions can be defined by the user and can be set to a multiple of the number of Mappers. The same key, that is, id, will be aggregated to ensure the uniqueness of the id (to achieve the purpose of id deduplication). The sub-data to be divided in the form of key-value pairs in these partitions will be assigned to the corresponding Reducer to ensure that each Reducer receives the same or almost the same number of sub-data to be divided.
[0061] In practical applications, data storage parameters can be dynamically adjusted according to actual needs. The dynamic adjustment formula is the following formula (1):
[0062] max_shard_data_num > k*tota l / (shard-1) (1)
[0063] Among them, max_shard_data_num indicates the maximum amount of data allowed to be allocated to each data shard, that is, the data storage parameter of the data shard. This is a threshold used to prevent the amount of data in a data shard from exceeding the storage range. k indicates the number of data shards to which a data point corresponding to a sub-data to be divided can be allocated, that is, the repetition factor. This means that each data point may be allocated to multiple shards, up to k. Total indicates the total amount of data of the target data. Shard indicates the total number of data shards, that is, the target number. During the allocation of the sub-data to be divided, the sub-data to be divided will be evenly distributed to these data shards. Shard-1 means reducing the number of data shards by 1, because the design of this formula needs to consider the situation that each data shard can "overflow" to other data shards to ensure the load balance of each data shard. This formula ensures that the maximum amount of data allocated to each data shard is proportional to the total amount of data of the sub-data to be divided, total l, and the number of data shards, and considering that in extreme cases, the amount of data allocated to the fullest data shard may be more than the average, so shard-1 is used to ensure load balance and avoid the risk of data overload.
[0064] In summary, at least two target sub-data are partitioned to achieve the purpose of deduplication of the target sub-data based on identification information. Within the range of data storage parameters, the sub-data to be divided is divided into at least two data shards, so that the data volume of the sub-data to be divided in each data shard is relatively balanced, which reduces the subsequent data processing pressure and improves the data processing efficiency.
[0065] In one or more embodiments provided in this specification, dividing the sub-data to be divided into the at least two data fragments according to the data storage parameter and the distance information between the centroid of the sub-data to be divided and the target data includes:
[0066] For the sub-data to be divided in the data to be divided, calculate the distance information between the multiple centroids of the target data and the sub-data to be divided, and determine the centroid sequence corresponding to the multiple centroids based on the distance information; select the target centroid in the centroid sequence based on the distance information; judge whether the target centroid meets the data storage conditions based on the data storage parameters; if so, store the sub-data to be divided to the data slice corresponding to the target centroid; if not, select candidate centroids other than the target centroid in the centroid sequence based on the distance information, and use the candidate centroid as the target centroid, and execute the step of judging whether the target centroid meets the data storage conditions based on the data storage parameters.
[0067] Among them, the distance information between multiple centroids and the sub-data to be divided can be calculated by square Euclidean distance, Manhattan distance, Chebyshev distance and cosine similarity. The centroid sequence refers to a sequence obtained by sorting at least two centroids in order of distance from small to large or from large to small according to the distance information. The target centroid can be a centroid with a smaller distance to the sub-data to be divided. When the centroids in the centroid sequence are arranged in order of distance from small to large, the first centroid in the centroid sequence can be selected as the target centroid; when the centroids in the centroid sequence are arranged in order of distance from large to small, the centroid at the end of the centroid sequence can be selected as the target centroid. The data storage condition refers to the quantity storage condition of the data, and the data storage parameter indicates the maximum storage data volume of the data shard corresponding to the centroid. Meeting the data storage condition means that the amount of data stored in the data shard corresponding to the target centroid has not reached the maximum storage data volume, and there is still data storage space in the data shard; on the contrary, not meeting the data storage condition means that the amount of data stored in the data shard corresponding to the target centroid has reached the maximum storage data volume, and the sub-data to be divided cannot be stored any more. The candidate centroid refers to a centroid whose distance between the centroid and the sub-data to be divided is greater than the distance between the target centroid and the sub-data to be divided. The candidate centroid may also be a centroid arranged after the target centroid in the centroid sequence.
[0068] Based on this, when calculating the distance information between multiple centroids of the target data and the sub-data to be divided, it is necessary to calculate each sub-data to be divided, and each sub-data to be divided can be traversed in turn to calculate the distance with multiple centroids, or it can be traversed randomly, and the present disclosure does not limit this. For example, select the first sub-data to be divided, calculate the distance information between the target number of centroids and the first sub-data to be divided, and sort the centroids based on the distance information. The centroids can be arranged in order from small to large according to the distance information, and a centroid sequence formed by arranging the target number of centroids is obtained. Select the target centroid that ranks first (with the smallest distance corresponding to the distance information) in the centroid sequence, determine whether the target centroid meets the data storage conditions based on the data storage parameters, and determine whether the data shard corresponding to the target centroid can continue to store the first sub-data to be divided; if the target centroid does not meet the data storage conditions, it means that the data shard corresponding to the target centroid is full and cannot continue to store other sub-data to be divided. At this time, candidate centroids other than the target centroid can be selected from the target number of centroids based on the distance information, that is, a centroid arranged after the target centroid is selected as the candidate centroid, and the candidate centroid is used as the target centroid, and the step of determining whether the target centroid meets the data storage conditions based on the data storage parameters is continued. If the candidate centroid as the target centroid meets the data storage conditions, it means that the data shard corresponding to the candidate centroid can continue to store the sub-data to be divided, then the first sub-data to be divided is stored in the data shard corresponding to the candidate centroid. If the candidate centroid still does not meet the data storage conditions, the next centroid is selected as the target centroid in the centroid sequence until the storage of the first sub-data to be divided is completed.
[0069] If the target centroid meets the data storage conditions, it means that the data shard corresponding to the target centroid can continue to store the sub-data to be divided, and the first sub-data to be divided is stored in the data shard corresponding to the target centroid. The second sub-data to be divided is selected from the multiple sub-data to be divided, and subsequent target centroid determination and data storage operations are performed until each sub-data to be divided is stored in the data shard.
[0070] Using the above example, the sub-data to be divided is evenly stored in at least two data shards. Each data shard allocates data by selecting the centroid closest to the data point of the sub-data to be divided. Each centroid corresponds to a data shard, and the distance between the centroid and the data point is calculated by the squared Euclidean distance. To avoid overloading the data shards, data storage parameters are set for the data shards. The data storage parameters represent the upper limit of the data storage capacity of the data shards.
[0071] In practical applications, an overlapping sharding algorithm based on dynamic load balancing and centroid distance can be used to determine data shards for the sub-data to be divided. When the number of sub-data to be divided (i.e., the total data volume of the target data) is 100 and the number of centroids is 10, the number of data shards is also 10. Each sub-data to be divided will be allocated to k=3 data shards. According to the above formula (1), it can be calculated that the number of sub-data to be divided allocated in each data shard needs to be greater than 33. Assuming that the data storage parameter max_shard_data_num is 33, the first 99 sub-data to be divided are divided into 9 data shards, and each data shard stores 33 sub-data to be divided. If max_shard_data_num=33, the first 9 data shards are full, and the last data cannot be sharded into k=3 shards. Therefore, max_shard_data_num needs to be greater than 33, for example, max_shard_data_num=34, then the last sub-data to be divided can be successfully allocated to any 3 data shards among the 10 data shards.
[0072] The pseudo code of the algorithm for allocating data slices to be divided is as follows:
[0073]
[0074] Output:
[0075] Data points are successfully assigned to at most k shards
[0076] #Step 1: Initialize the minimum heap and sort by centroid distance
[0077] heap = empty_min_heap()
[0078] #Step 2: Calculate the distance between the data point and all centroids and store the centroid information in the heap
[0079] for each centroid in centroids do
[0080] distance=compute_distance(val,centroid)
[0081] heap.push(centroid.index,distance)
[0082] #Step 3: Take the nearest centroid from the heap one by one and distribute it to at most k different shards
[0083] assigned_count = 0 # Used to track the number of assigned shards
[0084] while heap is not empty and assigned_count <k do
[0085] centroid = heap.pop() # Get the nearest centroid
[0086] shard_index=centroid.index
[0087] #If the shard is not full, allocate data
[0088] if shard_counts[shard_index] <max_shard_data_num then
[0089] assign_to_shard(val,doc_id,shard_index)
[0090] shard_counts[shard_index]+=1
[0091] assigned_count+=1#Update the number of assigned shards
[0092] #Step 4: If at least one shard is successfully allocated, return success, otherwise throw an error
[0093] if assigned_count>0then
[0094] return SUCCESS else
[0095] raise_error("No available shard with enough space")
[0096] return FAILURE
[0097] Among them, when initializing the minimum heap, a minimum heap is created to sort the centroids according to the distance between the centroids and the data points corresponding to the sub-data to be divided. The centroids closer to the data points will be processed first to ensure that the data points are assigned to the nearest data shards first. When calculating the distance to the centroid, the L2 distance between each centroid and the data point is calculated, and the index of the centroid and the corresponding distance are stored in the minimum heap to ensure that the centroids are sorted according to the distance to the data point, so that the nearest centroid is assigned first. When allocating the sub-data to be divided to the nearest k shards, the centroids closest to the data point are popped out one by one from the minimum heap, and the corresponding data shard index is obtained. If the amount of data stored in the data shard does not exceed max_shard_data_num, the data point is assigned to the data shard, and the load counter shard_counts of the data shard is updated. The assigned_count variable is used to control each data point to be assigned to a maximum of k data shards. Once it is successfully assigned to k data shards, the algorithm terminates. If at least one suitable data shard cannot be found for a data point after traversing all centroids, an error is thrown. This usually occurs in extreme cases where the amount of data stored in all data shards has reached max_shard_data_num.
[0098] In summary, by using the minimum heap, it is possible to ensure that the centroid closest to the data point corresponding to the sub-data to be divided is selected for data shard allocation each time. The time complexity of the insertion and pop-up operations of the minimum heap is O(log n), which can improve processing efficiency. Ensuring that data points can be allocated to a maximum of k data shards can improve the reliability and redundancy of the system, increase data redundancy and enhance the stability of retrieval. Through the max_shard_data_num data storage parameter, the maximum load of each data shard can be effectively controlled during dynamic sharding to avoid memory overload problems. By dynamically checking the current storage data volume of each data shard and selecting the data shard with a lighter load, it is ensured that the load of all data shards is balanced during the index construction process.
[0099] In one or more embodiments provided in this specification, determining a centroid sequence corresponding to the at least two centroids based on the distance information, and selecting a target centroid in the centroid sequence includes:
[0100] The multiple centroids are sorted based on the distance information to obtain the centroid sequence, and the centroid sequence is stored in a minimum heap; and the target centroid is extracted from the centroid sequence stored in the minimum heap based on the distance information.
[0101] Among them, the minimum heap is a sorted complete binary tree, in which the data value of any non-terminal node is not greater than the value of its left child node and right child node. By using the minimum heap, it can be ensured that the centroid closest to the data point of the sub-data to be divided is selected for data shard allocation each time. The time complexity of the insertion and pop-up operation of the centroid in the minimum heap is 0 (log n), which is suitable for processing a large number of centroids and large-scale data sharding problems. The target centroid is the selected centroid with a shorter distance to the i-th sub-data to be divided.
[0102] Based on this, multiple centroids are sorted based on distance information, and sorted from small to large or from large to small to obtain a centroid sequence. The centroids in the centroid sequence are stored in the minimum heap in order. The target centroid with the smallest distance to the sub-data to be divided is extracted from the centroid sequence stored in the minimum heap, and the sub-data to be divided is then divided into the data shards corresponding to the target centroid.
[0103] In summary, the minimum heap is used to store the centroid sequence, and the time complexity of the centroid insertion and pop-up operations in the minimum heap is 0 (log n), which improves the data processing efficiency.
[0104] Step 204: Allocate an index generator to the data set corresponding to each data shard in the target number of data shards, and use the index generator to construct a sub-index of each data set.
[0105] Specifically, after the target data is divided into the target number of data shards based on the multiple centroids of the target data, an index generator can be assigned to the data set corresponding to each data shard in the target number of data shards, and the index generator can be used to build a sub-index for each data set, wherein the index generator refers to the reducer, which is the processor corresponding to the Reduce stage in the MapReduce programming model. The sub-index refers to an index graph built based on the data set assigned to the index generator. There are at least two index generators, and each index generator builds a sub-index based on the assigned data set to obtain at least two sub-indexes. The sub-index is the local index corresponding to the target data.
[0106] Based on this, after the target data is divided into the target number of data shards based on the multiple centroids of the target data, an index generator is assigned to the data set corresponding to each data shard, and a data set is assigned to each index generator according to the data carrying capacity of the index generator, with priority given to assigning data sets to index generators with smaller loads and larger data capacities. The index generator is used to build a sub-index for each data set, and each index generator builds a sub-index based on the data set assigned to it.
[0107] In one or more embodiments provided in this specification, the data set allocation index generator corresponding to each data shard includes:
[0108] Determine multiple index generators and capacity information of each index generator; sort the data sets corresponding to each data shard in the target number of data shards according to the data set size to obtain a data set sequence; and assign an index generator to each data set based on the data set sequence and the capacity information of each index generator.
[0109] The capacity information of the index generator refers to the data capacity threshold of the index generator, which indicates the maximum data storage capacity of the index generator, that is, the maximum capacity of the index generator. The size of the data set can be the size of the amount of data in the data set, or the size of the storage space occupied by the amount of data in the data set. The data set sequence refers to the data set sequence obtained by sorting the data sets in at least two data shards according to the data size.
[0110] Based on this, the data in the same data shard is sent to the corresponding Mapper for data aggregation to obtain the data set corresponding to each data shard. Each data set is composed of homogeneous centroid data, and the number of data sets is equal to the number of centroids. Determine multiple index generator Reducers and the capacity information of each index generator. Sort the data sets corresponding to each data shard according to the size of the storage space occupied by the number of data sets to obtain a data set sequence. Based on the data set sequence and the capacity information of each index generator, assign index generators to each data set in turn, so that the data sets in the data set sequence are evenly distributed to multiple index generators.
[0111] In summary, an index generator is allocated to each data set based on the capacity information of the index generator, so as to achieve load balancing of the index generator and avoid overloading of the index generator.
[0112] In one or more embodiments provided in the present specification, the assigning of an index generator to each data set based on the data set sequence and the capacity information of each index generator includes: for a data set in the data set sequence, determining based on the capacity information of each index generator whether there is a target index generator among the multiple index generators that matches the size of the data set; if so, assigning the data set to the target index generator.
[0113] The target index generator is an index generator whose remaining data capacity among the multiple index generators matches the storage space size occupied by any data set in the data set sequence.
[0114] In practical applications, when n data sets are stored in a data set sequence, the first data set is selected in the data set sequence, the capacity information of each index generator is determined, and an index generator with a larger capacity and a lighter load is selected from the m index generators as the target index generator. At this time, it is necessary to determine whether there is a target index generator that matches the storage space size occupied by the first data set among the m index generators; if there is a target index generator that matches the storage space size occupied by the first data set among the m index generators, the first data set is allocated to the target index generator, and then the second data set is selected in the data set sequence, and the index generator is continuously allocated to the second data set until the allocation of index generators for the n data sets is completed; if there is no target index generator that matches the storage space size occupied by the first data set among the m index generators, an error message is generated, indicating that the index generator for the first data set has not been successfully matched, and the first data set cannot be allocated.
[0115] Continuing with the above example, after allocating the sub-data to be divided to at least two data shards, the sub-data to be divided in each data shard can be aggregated to obtain a set of key-value pairs corresponding to each data shard, that is, a data set. Sort the data sets by size to obtain a data set sequence, in which the data sets that occupy a larger storage space are arranged at the front. Take out the first data set in the data set sequence, that is, the data set that occupies a larger storage space. Determine the target reducer with a smaller load among at least two reducers and that can accommodate the first data set. Assign the first data set to the target reducer. And so on, until the target number of data sets are assigned to the reducers. Obtain the distribution results of the data sets and the maximum load among all reducers.
[0116] In practical applications, the greedy strategy-based task scheduling algorithm can be used to distribute the data sets in the data set sequence to multiple index generators (reducers). This algorithm is used to distribute the data sets composed of data in data shards to multiple reducers, each of which has a capacity limit. The goal of the algorithm is to ensure that no reducer exceeds its capacity and maximize the load balance between reducers, thereby minimizing the maximum load.
[0117] The pseudo code of the algorithm for assigning reducers to data sets in data shards is as follows:
[0118] Algorithm AssignShardsToReducers(K,s,N,r)
[0119] Input: K (number of shards), s (data shard size array), N (number of reducers), r (reducer capacity array)
[0120]
[0121] Among them, when sorting the data shards, the data shard size array s can be sorted in descending order to ensure that the data shards that occupy a larger storage space are processed first, thereby balancing the load of the reducer. Initialize multiple variables, oReducer_loads is used to record the current load of each reducer; oassignment is used to store the reducer assigned to each data shard; omax_load is used to determine the maximum load among all reducers. In addition, for each data shard, the algorithm can find the reducer with the smallest current load and can accommodate the data set of the data shard to ensure that the capacity of the reducer is not exceeded. If such a reducer is not found, the algorithm returns an error. After all data sets in the data set sequence are assigned, the maximum load among all reducers is determined. The algorithm returns the assignment array assignment and the maximum load max_load.
[0122] This algorithm helps to balance the load of reducers more effectively by sorting the data sets of data shards in descending order to ensure that larger data shards are processed first. Using a greedy strategy, the reducer with the smallest current load and capable of accommodating the shards is selected each time to ensure load balancing. The r array is used to ensure that the load of each reducer does not exceed its capacity. If the data sets of all data shards cannot be allocated without exceeding the capacity of any reducer, the algorithm returns an error. By dynamically checking the current load of each reducer and selecting the reducer with a lighter load, the load of all reducers is ensured to be balanced during the allocation process. Since the time to build an index is positively correlated with the size of the graph (amount of data / memory), balancing the load of the build tasks can minimize the impact of the long tail.
[0123] In summary, data sets are selected in the data set sequence in turn, and after matching the data sets with the capacity information of the index generator, the data sets are stored in the target index generator to achieve load balancing to multiple index generators and avoid overloading of the index generator.
[0124] In one or more embodiments provided in this specification, the construction of any sub-index includes:
[0125] A target data set is selected from a target number of data sets, and node data contained in the target data set and node relationships corresponding to the node data are determined; the node relationships are used as edges and the node data are used as nodes to construct a sub-index corresponding to the target data set.
[0126] Among them, node data refers to the data in the target data set that can be converted into graph nodes, and node relationships refer to the associations between node data in the target data set. Node relationships can be adjacent relationships between data points, or they can indicate that the meanings of node data are related. Both nodes and edges are elements that constitute a graph.
[0127] Based on this, a target dataset is selected from the target number of datasets, and a sub-index is constructed based on the target dataset. The key-value pairs contained in the target dataset are used as node data, the internal connections between the node data are determined, and the internal connections between the node data are used as node relationships. The node relationships are abstracted into edges, the node data are abstracted into nodes, the nodes are drawn, and the edges between the nodes are drawn to construct the sub-index corresponding to the target dataset.
[0128] Following the above example, after assigning the dataset to the reducer, it means assigning the sub-index construction task to the reducer. The reducer needs to perform the sub-index construction task and build the sub-index corresponding to the dataset based on a dataset. Drawing graph nodes based on the data in the form of key-value pairs in the dataset and drawing edges between graph nodes based on the relationship between key-value pairs to build the sub-index corresponding to the dataset can improve resource utilization efficiency and also improve the efficiency of sub-index construction.
[0129] In summary, by completing the construction of sub-indexes in sequence in the index generator, a number of sub-indexes corresponding to the number of data sets can be drawn.
[0130] Step 206: Integrate the sub-index of each data set into a target index corresponding to the target data.
[0131] Specifically, after allocating an index generator to the data set corresponding to each data shard and using the index generator to construct a sub-index for each data set, the sub-index for each data set can be integrated into a target index corresponding to the target data, wherein the target index corresponds to the target data and is the index corresponding to the target data. The target index is the global index corresponding to the target data. Resource query tasks refer to tasks for querying resources submitted by users.
[0132] Based on this, after allocating an index generator to the data set corresponding to each data shard and using the index generator to build a sub-index for each data set, the sub-indexes are integrated using a random neighbor selection strategy, and the sub-indexes of at least two data sets are integrated into the target index corresponding to the target data, thereby realizing the merging of at least two sub-indexes.
[0133] In actual applications, when assigning an index generator to the data set corresponding to each data shard, each centroid corresponds to one data shard, so the number of data shards can be equal to the number of centroids. When aggregating the data set corresponding to the same data shard, since there are a target number of data shards, the constructed data set is also the target number, and finally the target number of sub-indexes are constructed. Subsequently, the target number of sub-indexes can be merged to obtain the target index.
[0134] In one or more embodiments provided in this specification, integrating the sub-index of each data set into a target index corresponding to the target data includes:
[0135] An index integration strategy is determined, and the sub-indexes of each data set are integrated according to the index integration strategy, and a target index corresponding to the target data is determined according to the integration result.
[0136] The index integration strategy may be a random neighbor selection strategy for integrating at least two sub-indexes into one target index. The integration result may be a node connection result between at least two sub-indexes.
[0137] Based on this, determine the index integration strategy and at least two sub-indexes that need to be spliced and integrated. Determine the node relationship between the sub-indexes according to the index integration strategy, take the node relationship as the integration result, and obtain the target index corresponding to the target data by connecting two nodes with the node relationship, that is, merge at least two sub-indexes into the target index.
[0138] In practical applications, after determining the sub-index corresponding to each data shard, the random neighbor selection strategy can be used to merge only the sub-indexes. Create an empty graph G as the final target graph. The sub-index is represented by G1-Gn. Taking the construction of the target index based on G1 and G2 as an example, for each node n1 in G1, a node n2 is randomly selected from G2 as the neighbor of n1. "Random selection" can mean uniformly randomly selecting a node from all nodes in G2, or selecting nodes according to a certain probability distribution (for example, selecting according to the degree or label of the node). Similarly, for each node n2 in G2, a node n1 can also be randomly selected from G1 as the neighbor of n2. It should be noted that this selection is unidirectional, that is, n1 selecting n2 as a neighbor does not mean that n2 also selects n1 as a neighbor (unless they happen to select each other). Therefore, the final graph G may be a directed graph. After determining the neighbor relationship, add corresponding edges in G to connect these nodes. If a bidirectional neighbor relationship is selected, an undirected edge is added; if a unidirectional neighbor relationship is selected, a directed edge is added. In addition, in order to improve retrieval efficiency and accuracy, some optimization operations can be performed on the final graph G, such as removing duplicate edges, merging identical nodes, etc.
[0139] In summary, at least two sub-indexes are merged based on the index integration strategy to obtain a target index, which is used to perform subsequent resource detection tasks and improve the accuracy of resource detection task execution results.
[0140] In one or more embodiments provided in this specification, after integrating the sub-index of each data set into the target index corresponding to the target data, the method further includes:
[0141] Determine a resource query task in a resource browsing service; and perform a query according to the target index of the resource to be queried.
[0142] Using the above example, a resource query task can be submitted by a user who uses the resource browsing service. When a user has an image query requirement, a resource query task can be generated based on the image to be detected. By executing a resource query task.
[0143] In summary, the target index is used to assist in executing resource query tasks, improve the execution efficiency of resource query tasks, and improve the accuracy of resource queries.
[0144] An index construction method provided by an embodiment of the present specification divides the target data into a target number of data shards based on multiple centroids of the target data, that is, divides the target sub-data in the target data into at least two data shards. An index generator is allocated to the data set corresponding to each data shard in the target number of data shards, and the sub-index of each data set is constructed using the index generator, making full use of the index generator to achieve efficient and balanced utilization of resources, while improving the efficiency of constructing the sub-index. The sub-index of each data set is integrated into the target index corresponding to the target data, shortening the time consumption of constructing the target index.
[0145] In practical applications, the index construction method provided in this specification is applicable to building multiple types of indexes such as hash index, B-Tree index, spatial index, full-text index, etc. The index construction method is also applicable to building graph index. Graph index is a special type of index, which is mainly used to process graph data, that is, data that can be represented as a set of nodes (or vertices) and edges (or connections). In graph databases, graph indexes are used to speed up queries and operations on graph data.
[0146] The following combination Figure 3 , taking the application of the index construction method provided in this specification in the construction of a global graph index as an example, the index construction method is further described. Figure 3 A processing flow chart of a graph index construction method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0147] Step 302: Determine the original data associated with the resource query task.
[0148] When building a global graph index, obtain the original data used for resource query, and perform data preprocessing and sharding on the original data.
[0149] Step 304: clustering the original data to obtain clusters, and assigning the clusters to mapping partitions.
[0150] The process of preprocessing and slicing the original data is as follows: Figure 4 In the first process shown, the K-means clustering algorithm is used to cluster the original data id1, vector1, id2, vector2, id3, vector3, etc. to obtain cluster clusters. The original data contains id and vector data vector. Each piece of data in the original data is read, and each piece of data is processed according to the number of Mappers and processing functions defined by the user to obtain vector: processed_vec, and output: intermediate key-value pairs, such as {id1, processed_vec1}, {id2, processed_vec2},….
[0151] Step 306: pre-process the vector data in the cluster in the mapping partition to obtain key-value pair data.
[0152] Step 308: After breaking up and sorting the key-value pair data, intermediate key-value pair data is obtained.
[0153] In this embodiment, the original data is divided into three partitions, Mapper1-Mapper3. Mapper1 stores data such as id1 and raw_vec1, Mapper2 stores data such as id2 and raw_vec2, and Mapper3 stores data such as id3 and raw_vec3. The data in the three partitions are scattered and sorted. The intermediate key-value pairs will be scattered, partitioned and sorted according to the key, that is, the id. The number of partitions, that is, the number of Reducers, is defined by the user (usually a multiple of the number of Mappers). The same key, that is, the id, will be aggregated to ensure the uniqueness of the id (to achieve the purpose of id deduplication). Finally, these intermediate key-value pairs will be randomly sent to the number of Reducers defined by the user on average to ensure that each Reducer receives the same or almost the same number of intermediate key-value pairs.
[0154] Step 310: Allocate the intermediate key-value pair data to at least two data shards using an overlapping sharding algorithm with dynamic load balancing based on centroid distance.
[0155] The overlapping sharding algorithm with dynamic load balancing based on centroid distance distributes the intermediate key-value pair data to at least two data shards. The algorithm is used to distribute the data points corresponding to the intermediate key-value pairs to multiple shards, and each shard distributes the data by selecting the centroid closest to the data point. Each centroid corresponds to a shard, and the distance between the centroid and the data point is calculated by the squared Euclidean distance. In order to avoid overloading some shards, the algorithm dynamically adjusts the distribution so that the amount of data in each shard remains below the preset maximum value max_shard_data_num.
[0156] In practical applications, such as Figure 4 As shown, the intermediate key-value pairs are distributed to Reducer1-Reducer3. Reducer1 stores data such as Partition(id1, processed_vec1, k), where k represents the maximum number of shards to which a data point can be assigned. Reducer2 stores data such as Partition(id2, processed_vec2, k), and Reducer3 stores data such as Partition(id3, processed_vec3, k). Algorithm input: doc_id: unique identifier of the current data point. val: vector representation of the data point. centroids: centroid set, used to calculate the distance between the data point and the centroid of each shard. shard_counts: current data volume counter for each shard, used to track the load of each shard. max_shard_data_num: maximum data volume threshold for each shard to ensure that a shard does not carry too much data. k: maximum number of shards to which a data point can be assigned (i.e., duplicate distribution factor or duplication factor). The final algorithm outputs data such as id1, processed_vec1, c1; id1, processed_vec1, c2, id2, processed_vec2, c2; id2, processed_vec2, c3; id3, processed_vec3, ck, etc. The data points are successfully assigned to a maximum of k shards. If no shard is assigned, an error is returned and insufficient space is prompted. In other words, when assigning Reducers to intermediate key-value data, each centroid corresponds to a data shard, and the total number of data shards can be equal to the number of centroids corresponding to the target data.
[0157] When allocating the intermediate key-value data to at least two data shards, a minimum heap is created to sort the distance between the centroids and the data points. The centroids with the closest distance will be processed first to ensure that the data points are allocated to the closest shards first. For each centroid, the L2 distance between it and the data point is calculated, and the centroid index and the corresponding distance are stored in the minimum heap. This ensures that the centroids are sorted according to the distance to the data point, so that the closest centroids are allocated first. The nearest centroids are popped out from the minimum heap one by one, and the corresponding shard index is obtained. If the data volume of the shard does not exceed the data volume limit, the data point is allocated to the shard, and the load counter shard_counts of the shard is updated. Each data point is controlled to be allocated to at most k shards. Once k shards are successfully allocated, the algorithm terminates. If after traversing all centroids, at least one suitable shard cannot be found for the data point, an error is thrown. This usually occurs in extreme cases where the data volume of all shards has reached the data volume limit.
[0158] By using the minimum heap, it is possible to ensure that the centroid closest to the data point is selected for shard allocation each time. The time complexity of the insertion and pop-up operations of the minimum heap is 0 (log n), which is very suitable for handling a large number of centroids and large-scale data sharding problems. The algorithm ensures that data points can be allocated to at most k shards. This is crucial to improving the reliability and redundancy of the system, so that data points can be distributed to multiple shards, increasing data redundancy and improving retrieval stability. By setting the parameter parameter of the data volume upper limit, the maximum load of each shard can be effectively controlled during the dynamic sharding process to avoid memory overload problems. When the data volume of a shard reaches the upper limit, the algorithm automatically skips the shard and looks for the next available shard. By dynamically checking the current data volume of each shard and selecting the shard with a lighter load, it ensures that all shards are load balanced during the index building process.
[0159] Step 312: Perform homocentric data aggregation on the data in at least two data shards to obtain at least two data sets.
[0160] Step 314: Allocate at least two data sets to a plurality of reducers through a task scheduling algorithm based on a greedy strategy.
[0161] The input of the task scheduling algorithm based on the greedy strategy is the output of the overlapping sharding algorithm based on the dynamic load balancing based on the centroid distance, that is, id1, processed_vec1, c1; id1, processed_vec1, c2, id2, processed_vec2, c2; id2, processed_vec2, c3; id3, processed_vec3, ck and other data.
[0162] like Figure 5As shown, the data belonging to the same centroid is sent to the same Mapper for aggregation to obtain K centroid data sets to be constructed. Mapper1 stores c1: {(id1, vec1)......}; Mapper2 stores c2: {(id1, vec1), (id2, vec2)......}, and Mapper3 stores c3: {(id2, vec2)......}. Customized task scheduling is performed to assign the data sets stored in Mapper1-Mapper3 to reducers Reducer1-Reducer3. When the data of the same centroid is sent to the same Mapper for aggregation to obtain the data set to be constructed, since the number of centroids is equal to the number of data shards, that is, there are K centroids, K data sets to be constructed can be constructed.
[0163] This algorithm is used to distribute data shards to multiple reducers, each of which has a capacity limit. The goal of the algorithm is to ensure that no reducer exceeds its capacity and to maximize the load balance between reducers, thereby minimizing the maximum load. By sorting the shards in descending order, larger shards are processed first, which helps to balance the load more effectively. Using a greedy strategy, the reducer with the smallest current load and the capacity to accommodate the shards is selected each time to ensure load balance. By dynamically checking the current load of each reducer and selecting the reducer with a lighter load, the load of all reducers is ensured to be balanced during the distribution process. Since the time to build a graph index is positively correlated with the size of the graph (amount of data / memory), evenly distributing the load of the build task can minimize the impact of the long tail.
[0164] Step 316: Use the reducer to construct a subgraph index corresponding to the data set.
[0165] Build subgraph indexes in Reducer1-Reducer3 respectively. BuildGraph(c1,c2,...)→(g1,g2,...) in Reducer1; BuildGraph(c4,c5,...)→(g4,g5,...) in Reducer2; BuildGraph(c7,c8,...)→(g7,g8,...) in Reducer3. Distribute K data sets to be built to multiple reducers, and finally build K subgraph indexes, which are then merged to obtain a complete index, that is, the target graph index.
[0166] Step 318: Merge at least two subgraph indexes to obtain a target graph index.
[0167] Merge the subgraph indexes under each reducer to obtain the target graph index Merge(g1,g2,g3,3......gk)
[0168] In summary, this embodiment adopts the parallel index construction architecture of MapReduce, applies the MapReduce framework to the construction of ultra-large-scale vector indexes, and realizes efficient task allocation and processing by distributing data sharding and index construction tasks to multiple nodes of the cluster in parallel. The traditional single-machine index construction method is optimized to distributed parallel processing, which significantly shortens the construction time. At the same time, it supports dynamic expansion of cluster scale, so that the system can flexibly cope with data sets and hardware resources of different scales. A dynamic load balancing sharding algorithm based on heap sorting is adopted, and the data distribution method is dynamically adjusted in combination with factors such as the distance between the data point and the centroid and the current load of the shard. By giving priority to shards with lighter loads, data is evenly distributed to avoid memory overflow problems caused by too much shard data. By real-time monitoring and adjusting the shard load, a single shard is prevented from being overloaded, thereby improving the stability and reliability of the system. The resource utilization rate of each node in the cluster is improved, the uniform distribution of hardware resources (such as memory and storage) is ensured, resource waste is reduced, and the overall performance of the system is optimized. An embodiment of this specification proposes an adaptive resource scheduling mechanism that can dynamically adjust task allocation according to the memory and storage resource conditions of each node to ensure that the system can efficiently utilize hardware resources. This avoids bottlenecks caused by memory and storage resource limitations during task allocation and improves the parallel processing capabilities of the system. By reasonably allocating resources, resource contention between nodes in the cluster is reduced, further improving the processing efficiency of the system.
[0169] Corresponding to the above method embodiment, this specification also provides an index construction device embodiment, Figure 6 FIG. 1 shows a schematic diagram of a structure of an index building device provided by an embodiment of the present specification. Figure 6 As shown, the device comprises:
[0170] A partitioning module 602 is configured to partition the target data into a target number of data fragments based on multiple centroids of the target data, wherein target sub-data in the target data is partitioned into at least two data fragments;
[0171] A construction module 604 is configured to allocate an index generator to a data set corresponding to each data shard in the target number of data shards, and use the index generator to construct a sub-index of each data set;
[0172] The integration module 606 is configured to integrate the sub-index of each data set into a target index corresponding to the target data.
[0173] In an optional embodiment, the division module 602 is further configured to:
[0174] A clustering algorithm is used to divide a plurality of target sub-data in the target data into a plurality of clusters, and a plurality of centroids corresponding to the plurality of clusters are used as a plurality of centroids of the target data.
[0175] In an optional embodiment, the division module 602 is further configured to:
[0176] Partition the target data according to the key of the target sub-data to obtain data to be partitioned corresponding to at least two partitions;
[0177] Determine a target number of data shards and a data storage parameter corresponding to each data shard, wherein the data storage parameter is determined based on a repetition factor, a data volume of the target data, and a target number;
[0178] For the sub-data to be divided in the data to be divided corresponding to each partition, the sub-data to be divided is divided into at least two data fragments according to the data storage parameters and the distance information between the sub-data to be divided and each centroid of the target data.
[0179] In an optional embodiment, the division module 602 is further configured to:
[0180] For the sub-data to be divided in the data to be divided, calculating distance information between a plurality of centroids of the target data and the sub-data to be divided, and determining a centroid sequence corresponding to the plurality of centroids based on the distance information;
[0181] selecting a target centroid in the centroid sequence based on the distance information;
[0182] Determining whether the target centroid meets the data storage condition based on the data storage parameter;
[0183] If so, storing the sub-data to be divided into the data slice corresponding to the target centroid;
[0184] If not, based on the distance information, select candidate centroids other than the target centroid in the centroid sequence, and use the candidate centroids as the target centroid, and execute the step of determining whether the target centroid meets the data storage conditions based on the data storage parameters.
[0185] In an optional embodiment, the division module 602 is further configured to:
[0186] Sort the multiple centroids based on the distance information to obtain the centroid sequence, and store the centroid sequence in a minimum heap;
[0187] The target centroid is extracted from the centroid sequence stored in the minimum heap based on the distance information.
[0188] In an optional embodiment, the construction module 604 is further configured to:
[0189] Determine the target number of index generators and the capacity information of each index generator;
[0190] Sort the data sets corresponding to each data shard in the target number of data shards according to the data set size to obtain a data set sequence;
[0191] Based on the data set sequence and the capacity information of each index generator, an index generator is allocated to each data set.
[0192] In an optional embodiment, the construction module 604 is further configured to:
[0193] For a data set in the data set sequence, judging, according to the capacity information of each index generator, whether there is a target index generator matching the size of the data set among the target number of index generators;
[0194] If so, the data set is assigned to the target index generator.
[0195] In an optional embodiment, the construction module 604 is further configured to:
[0196] Selecting a target data set from at least two data sets, and determining node data contained in the target data set and node relationships corresponding to the node data;
[0197] The node relationship is used as an edge, and the node data is used as a node to construct a sub-index corresponding to the target data set.
[0198] In an optional embodiment, the integration module 606 is further configured to:
[0199] An index integration strategy is determined, and the sub-indexes of each data set are integrated according to the index integration strategy, and a target index corresponding to the target data is determined according to the integration result.
[0200] In an optional embodiment, the integration module 606 is further configured to:
[0201] Determine resource detection tasks in resource browsing business;
[0202] The target data to be detected is extracted in the resource detection task, and the target data to be detected is detected based on the target index to obtain a detection result.
[0203] An index building device provided by an embodiment of the present specification divides the target data into a target number of data shards based on multiple centroids of the target data, that is, divides the target sub-data in the target data into at least two data shards. An index generator is allocated to the data set corresponding to each data shard in the target number of data shards, and the sub-index of each data set is built using the index generator, making full use of the index generator to achieve efficient and balanced utilization of resources, while improving the efficiency of building the sub-index. The sub-index of each data set is integrated into the target index corresponding to the target data, shortening the time consumption of building the target index.
[0204] The above is a schematic scheme of an index building device of this embodiment. It should be noted that the technical scheme of the index building device and the technical scheme of the above index building method belong to the same concept, and the details not described in detail in the technical scheme of the index building device can be found in the description of the technical scheme of the above index building method.
[0205] Figure 7 The block diagram of a computing device 700 according to an embodiment of the present specification is shown. The components of the computing device 700 include but are not limited to a memory 710 and a processor 720. The processor 720 is connected to the memory 710 via a bus 730, and the database 750 is used to store data.
[0206] The computing device 700 also includes an access device 740 that enables the computing device 700 to communicate via one or more networks 760. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).
[0207] In one embodiment of the present specification, the above components of the computing device 700 and Figure 7 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 7 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0208] The computing device 700 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 700 may also be a mobile or stationary server.
[0209] The processor 720 is used to execute the following computer executable instructions, which, when executed by the processor, implement the steps of the above-mentioned index construction method.
[0210] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above index construction method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above index construction method.
[0211] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned index building method.
[0212] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the above index construction method belong to the same concept, and the details not described in detail in the technical scheme of the storage medium can be referred to the description of the technical scheme of the above index construction method.
[0213] An embodiment of the present specification also provides a computer program product, including a computer program or instructions, which implement the steps of the above index construction method when executed by a processor.
[0214] The above is a schematic scheme of a computer program product of this embodiment. It should be noted that the technical scheme of the computer program product and the technical scheme of the above index construction method belong to the same concept, and the details not described in detail in the technical scheme of the computer program product can be referred to the description of the technical scheme of the above index construction method.
[0215] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0216] The computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0217] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0218] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0219] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can understand and use this specification well.
Claims
1. An index construction method, characterized in that: include: Dividing the target data into a target number of data shards based on multiple centroids of the target data, wherein target sub-data in the target data is divided into at least two data shards; Allocate an index generator to the data set corresponding to each data shard in the target number of data shards, and use the index generator to build a sub-index for each data set; The sub-index of each data set is integrated into a target index corresponding to the target data.
2. The index construction method according to claim 1, characterized in that: Before dividing the target data into a target number of data slices based on the multiple centroids of the target data, the method further includes: A clustering algorithm is used to divide a plurality of target sub-data in the target data into a plurality of clusters, and a plurality of centroids corresponding to the plurality of clusters are used as a plurality of centroids of the target data.
3. The index construction method according to claim 1, characterized in that: The step of dividing the target data into a target number of data slices based on multiple centroids of the target data includes: Partition the target data according to the key of the target sub-data to obtain data to be partitioned corresponding to at least two partitions; Determine a target number of data shards and a data storage parameter corresponding to each data shard, wherein the data storage parameter is determined based on a repetition factor, a data volume of the target data, and a target number; For the sub-data to be divided in the data to be divided corresponding to each partition, the sub-data to be divided is divided into at least two data fragments according to the data storage parameters and the distance information between the sub-data to be divided and each centroid of the target data.
4. The index construction method according to claim 3, characterized in that: The step of dividing the sub-data to be divided into a target number of data slices according to the data storage parameter and the distance information between each centroid of the sub-data to be divided and the target data includes: For the sub-data to be divided in the data to be divided, calculating distance information between a plurality of centroids of the target data and the sub-data to be divided, and determining a centroid sequence corresponding to the plurality of centroids based on the distance information; selecting a target centroid in the centroid sequence based on the distance information; Determining whether the target centroid meets the data storage condition based on the data storage parameter; If so, storing the sub-data to be divided into the data slice corresponding to the target centroid; If not, based on the distance information, select candidate centroids other than the target centroid in the centroid sequence, and use the candidate centroids as the target centroid, and execute the step of determining whether the target centroid meets the data storage conditions based on the data storage parameters.
5. The index construction method according to claim 4, characterized in that: The selecting the target centroid in the centroid sequence based on the distance information comprises: Sort the multiple centroids based on the distance information to obtain the centroid sequence, and store the centroid sequence in a minimum heap; The target centroid is extracted from the centroid sequence stored in the minimum heap based on the distance information.
6. The index construction method according to claim 1, characterized in that: The data set allocation index generator corresponding to each data shard in the target number of data shards includes: Determine the target number of index generators and the capacity information of each index generator; Sort the data sets corresponding to each data shard in the target number of data shards according to the data set size to obtain a data set sequence; Based on the data set sequence and the capacity information of each index generator, an index generator is allocated to each data set.
7. The index construction method according to claim 6, characterized in that: The allocating an index generator to each data set based on the data set sequence and the capacity information of each index generator includes: For a data set in the data set sequence, judging, according to the capacity information of each index generator, whether there is a target index generator matching the size of the data set among the target number of index generators; If so, the data set is assigned to the target index generator.
8. The index construction method according to claim 1, characterized in that: The construction of any sub-index includes: Selecting a target data set from at least two data sets, and determining node data contained in the target data set and node relationships corresponding to the node data; The node relationship is used as an edge, and the node data is used as a node to construct a sub-index corresponding to the target data set.
9. The index construction method according to claim 1, characterized in that: The step of integrating the sub-index of each data set into a target index corresponding to the target data includes: An index integration strategy is determined, and the sub-indexes of each data set are integrated according to the index integration strategy, and a target index corresponding to the target data is determined according to the integration result.
10. The index construction method according to claim 1, characterized in that: After integrating the sub-index of each data set into the target index corresponding to the target data, the method further includes: Determine resource detection tasks in resource browsing business; The target data to be detected is extracted in the resource detection task, and the target data to be detected is detected based on the target index to obtain a detection result.
11. An index building device, characterized in that: include: a partitioning module, configured to partition the target data into a target number of data fragments based on a plurality of centroids of the target data, wherein target sub-data in the target data is partitioned into at least two data fragments; A construction module is configured to allocate an index generator to a data set corresponding to each data shard in the target number of data shards, and use the index generator to construct a sub-index of each data set; The integration module is configured to integrate the sub-index of each data set into a target index corresponding to the target data.
12. A computing device, characterized in that: include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the index construction method described in any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium, characterized in that: It stores computer executable instructions, which, when executed by a processor, implement the steps of the index building method described in any one of claims 1 to 10.
14. A computer program product, characterized in that The method comprises a computer program or an instruction, which, when executed by a processor, implements the steps of the index building method according to any one of claims 1 to 10.
Citation Information
Cited By
Efficient GPU neighbor graph index construction method based on data locality
CN120723941A
An efficient GPU-based construction method of proximity graph index based on data locality
CN120723941B
Cluster communication method and system based on artificial intelligence
CN121334170A