An index point offline sorting method, device, equipment and storage medium
By using the offline sorting method of index points in the search library, using the sliding window and the target correlation degree to sort the index points and saving them in the index memory, the inefficiency problem of randomly saving index points in the search library is solved, and more efficient search speed and index point sorting efficiency are achieved.
Patent Information
- Application Number
- CN202111490990.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-08
AI Technical Summary
In the search library of related applications, each index point is randomly stored in the index memory, which makes it a long time to obtain index points from the index memory during the search process, reducing the search efficiency.
Through an offline sorting method of index points, multiple index points iteratively sort through sliding windows and target correlations, and save them in index memory according to the sorting results.
It realizes that the associated index points are concentrated in the index memory, which reduces the number of memory access, improves the search speed, and improves the efficiency of index point sorting.
Smart Images

Figure CN114329135B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of search technology, and in particular, to a method, apparatus, device, and storage medium for offline sorting of index points. Background Art
[0002] With the development of Internet technology, the information in the Internet is constantly increasing. In order to facilitate the target object to quickly obtain the required information from the vast amount of information, the information search function has become an indispensable part of many applications.
[0003] In the related art, when performing a search, it is necessary to traverse the index points in the search library according to certain rules to obtain the search results. Among them, each index point represents a retrieved object, such as a retrieved article, a retrieved video, etc.
[0004] However, in the search library of related applications, each index point is randomly stored in the index memory. Therefore, during the search process, the time taken to obtain the index points from the index memory is relatively long, resulting in low search efficiency. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, device, and storage medium for offline sorting of index points to improve search efficiency.
[0006] On the one hand, embodiments of the present application provide a method for offline sorting of index points, the method comprising:
[0007] Select an index point from multiple index points of the search graph as an initial sorted index point and add it to the sliding window;
[0008] Determine the target correlation degree between each of the retained index points and the sorted index points in the sliding window;
[0009] Select a target index point from the retained index points based on each target correlation degree and add it to the sliding window;
[0010] Iteratively execute the step of determining the target correlation degree between each of the retained index points and the sorted index points in the sliding window until the multiple index points are added to the sliding window;
[0011] Take the order in which the multiple index points are added to the sliding window as the index point sorting result, and save the multiple index points in the index memory according to the index point sorting result.
[0012] On the one hand, embodiments of the present application provide an apparatus for offline sorting of index points, the apparatus comprising:
[0013] A selection module, configured to select one index point from multiple index points of a search graph as an initial sorted index point and add it to a sliding window;
[0014] A sorting module, configured to determine a target correlation degree between each of the retained index points and the sorted index points in the sliding window; select a target index point from each of the retained index points based on each target correlation degree and add it to the sliding window; iteratively execute the step of determining the target correlation degree between each of the retained index points and the sorted index points in the sliding window until all the multiple index points are added to the sliding window;
[0015] A storage module, configured to use the order in which the multiple index points are added to the sliding window as an index point sorting result, and save the multiple index points in an index memory according to the index point sorting result.
[0016] Optionally, the sorting module is further configured to:
[0017] After each iteration, if the number of sorted index points in the sliding window is greater than a preset threshold, remove the earliest added sorted index point from the sliding window, where the preset threshold is determined based on a central processing unit cache.
[0018] Optionally, the sorting module is specifically configured to:
[0019] Sort each of the index points according to each target correlation degree to obtain a target correlation sorting result;
[0020] Based on the target correlation sorting result, select a target index point from each of the index points as a sorted index point and add it to the sliding window.
[0021] Optionally, the sorting module is specifically configured to:
[0022] For each index point, respectively execute the following steps:
[0023] Obtain a historical correlation degree between an index point and the sorted index points in the sliding window during the previous iteration;
[0024] Determine a first sub-correlation degree between the one index point and a target index point newly added to the sliding window during the previous iteration;
[0025] Based on the first sub-correlation degree and the historical correlation degree, determine the target correlation degree between the one index point and the sorted index points in the sliding window.
[0026] Optionally, the sorting module is specifically configured to:
[0027] If, after the previous iteration, the earliest added sorted index point is removed from the sliding window, determine a second sub - association degree between the one index point and the earliest added sorted index point;
[0028] Based on the first sub - association degree, the second sub - association degree, and the historical association degree, determine a target association degree between the one index point and the sorted index points in the sliding window.
[0029] Optionally, the sorting module is specifically configured to:
[0030] For each index point, respectively perform the following steps:
[0031] Determine the association scores between each sorted index point in the sliding window and one index point;
[0032] Based on the obtained association scores, determine a target association degree between the one index point and the sorted index points in the sliding window.
[0033] Optionally, the sorting module is specifically configured to:
[0034] Obtain the historical association sorting results corresponding to each index point in the previous iteration;
[0035] Based on the respective target association degrees, adjust the historical association sorting results to obtain a target association sorting result.
[0036] Optionally, the index points in the target association sorting result are arranged in descending order of the target association degree;
[0037] The sorting module is specifically configured to:
[0038] Take the first index point in the target association sorting result as the target index point and add it to the sliding window.
[0039] Optionally, it further includes a search module;
[0040] The search module is specifically configured to:
[0041] Obtain candidate index points for the search condition from the index memory, and at least one other index point continuously saved with the candidate index point;
[0042] Save the at least one other index point in the central processing unit cache;
[0043] If the at least one other index point includes a neighbor index point of the candidate index point, obtain the neighbor index point of the candidate index point from the central processing unit cache;
[0044] Determine the search result of the search condition based on the candidate index points and the neighbor index points of the candidate index points.
[0045] On the one hand, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the above-mentioned index point offline sorting method are implemented.
[0046] On the one hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is enabled to execute the steps of the above-mentioned index point offline sorting method.
[0047] On the one hand, an embodiment of the present application provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is enabled to execute the steps of the above-mentioned index point offline sorting method.
[0048] In the embodiment of the present application, based on the target correlation degree between multiple index points, the multiple index points are sorted to obtain an index point sorting result. Then, according to the index point sorting result, the multiple index points are stored in the index memory, so as to realize centralized storage of associated index points in the index memory. Therefore, during the search process, multiple associated index points can be obtained for calculation each time the memory is accessed, thereby reducing the number of times of accessing the memory and further improving the search speed. Secondly, during the sorting process, the correlation degree between the unsorted index points and the sorted index points in the sliding window is calculated, avoiding calculating the correlation degree between the unsorted index points and all the sorted index points, reducing the process of repeatedly calculating the correlation degree, and thus improving the efficiency of index point sorting. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 It is a schematic structural diagram of a system architecture provided by an embodiment of the present application;
[0051] Figure 2 It is a schematic diagram of a search interface provided by an embodiment of the present application;
[0052] Figure 3A schematic diagram of a search result interface provided by an embodiment of the present application;
[0053] Figure 4 A schematic flowchart of an index point offline sorting method provided by an embodiment of the present application;
[0054] Figure 5 A schematic flowchart of a method for obtaining index points from memory provided by an embodiment of the present application;
[0055] Figure 6 A schematic flowchart of a method for obtaining index points from memory provided by an embodiment of the present application;
[0056] Figure 7 A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0057] Figure 8 A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0058] Figure 9 A schematic diagram of a common neighbor index point provided by an embodiment of the present application;
[0059] Figure 10 A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0060] Figure 11 A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0061] Figure 12 A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0062] Figure 13 A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0063] Figure 14a A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0064] Figure 14b A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0065] Figure 14c A schematic diagram of the processing result after the end of an iteration process provided by an embodiment of the present application;
[0066] Figure 15 A schematic diagram of a search scenario provided by an embodiment of the present application;
[0067] Figure 16Schematic flowchart of a method for constructing an HNSW graph provided by an embodiment of the present application;
[0068] Figure 17 Schematic diagram of an HNSW graph provided by an embodiment of the present application;
[0069] Figure 18 Schematic flowchart of an offline sorting method for index points provided by an embodiment of the present application;
[0070] Figure 19 Schematic structural diagram of an offline sorting device for index points provided by an embodiment of the present application;
[0071] Figure 20 Schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0072] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0073] For the convenience of understanding, the terms involved in the embodiments of the present invention will be explained below.
[0074] Navigable Small World: Abbreviated as NSW, used for approximate nearest neighbor search.
[0075] Hierarchical Navigable Small World: Abbreviated as HNSW, used for approximate nearest neighbor search.
[0076] Central Processing Unit: Abbreviated as CPU, the operation and control core of a computer system, and the final execution unit for information processing and program operation.
[0077] CPU Cache: A component used to reduce the average time required for a processor to access memory. When the processor issues a memory access request, it first checks whether the requested data is in the cache. If it exists (hit), the data is directly returned without accessing memory; if it does not exist (miss), the corresponding data in memory needs to be loaded into the cache first and then returned to the processor.
[0078] The design concept of the embodiments of the present application will be introduced below.
[0079] When performing a search, the related technology needs to traverse the index points in the search library according to certain rules to obtain search results, wherein each index point represents a retrieved object, such as a retrieved article, a retrieved video, etc.
[0080] However, in the search library of related applications, each index point is randomly stored in the index memory. Therefore, during the search process, it takes a long time to obtain the index point from the index memory, resulting in low search efficiency.
[0081] Through analysis, it is found that during the search process, the CPU of the computer device needs to traverse the index points in the search library according to certain rules. For each candidate index point traversed, it is necessary to obtain the candidate index point and the neighboring index points of the candidate index point from the index memory for calculation. When the CPU obtains the candidate index point from the index memory, in addition to obtaining the candidate index point, it will also obtain a part of other index points and save them in the CPU cache, where the part of the index points obtained are other index points that are continuously saved in the index memory with the candidate index point.
[0082] If the index points are sorted according to the association between them, and the index points are saved in the index memory according to the sorting results, then when the CPU obtains the candidate index points from the index memory, the extra index points are likely to be the neighbor index points of the candidate index points. In this way, after the CPU obtains the candidate index points from the index memory, it can directly obtain the neighbor index points of the candidate index points from the CPU cache, and then calculate the candidate index points and the neighbor index points of the candidate index points without obtaining the neighbor index points from the index memory, thereby effectively improving the search speed.
[0083] In view of this, an embodiment of the present application provides the following index point offline sorting method, the method comprising:
[0084] From the plurality of index points of the search graph, an index point is selected as an initial sorted index point to be added to the sliding window. Then, the target association degree between each of the retained index points and the sorted index points in the sliding window is determined. Based on each target association degree, a target index point is selected from each of the retained index points and added to the sliding window. The step of determining the target association degree between each of the retained index points and the sorted index points in the sliding window is iterated until a plurality of index points are added to the sliding window. The order in which the plurality of index points are added to the sliding window is used as the index point sorting result, and the plurality of index points are stored in the index memory according to the index point sorting result.
[0085] In the embodiments of the present application, based on the target correlation degree between multiple index points, the multiple index points are sorted to obtain the index point sorting result. Then, according to the index point sorting result, the multiple index points are stored in the index memory, so as to realize the centralized storage of associated index points in the index memory. Therefore, during the search process, multiple associated index points can be obtained for calculation each time the memory is accessed, thereby reducing the number of memory accesses and further improving the search speed. Secondly, during the sorting process, the correlation degree between the unsorted index points and the sorted index points within the sliding window is calculated, avoiding the calculation of the correlation degree between the unsorted index points and all sorted index points, reducing the process of repeatedly calculating the correlation degree, and thus improving the efficiency of index point sorting.
[0086] Reference Figure 1 , which is a system architecture diagram applicable to the embodiments of the present application. The system architecture includes at least a terminal device 101 and a server 102. The number of terminal devices 101 can be one or more, and the number of servers 102 can also be one or more. The present application does not make specific limitations on the number of terminal devices 101 and servers 102.
[0087] The target application with a pre-existing search function is in the terminal device 101, where the target application is a client application, a web version application, a mini-program application, etc. The terminal device 101 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart home appliance, a smart voice interaction device, a smart vehicle device, etc., but is not limited thereto.
[0088] The server 102 is the background server of the target application. The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device 101 and the server 102 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make limitations here.
[0089] The index point offline sorting method in the embodiments of the present application can be executed by the terminal device 101, or by the server 102, or by the interaction between the terminal device 101 and the server 102.
[0090] Taking the server 102 executing the index point offline sorting method in the embodiments of the present application as an example, the following steps are included:
[0091] Server 102 selects an index point from multiple index points of the search graph as the initial sorted index point and adds it to the sliding window. Then, it determines the target correlation degree between each of the remaining index points and the sorted index points in the sliding window. Based on each target correlation degree, it selects target index points from the remaining index points and adds them to the sliding window. The step of determining the target correlation degree between each of the remaining index points and the sorted index points in the sliding window is iteratively executed until multiple index points are added to the sliding window. The order in which the multiple index points are added to the sliding window is used as the index point sorting result, and the multiple index points are saved in the index memory according to the index point sorting result.
[0092] In practical applications, the index point offline sorting method in the embodiments of the present application can be applied to search scenarios such as article search, video search, and commodity search. Taking the article search scenario as an example:
[0093] Server 102 offline constructs an HNSW graph and sorts each index point in the HNSW graph by using the method in the embodiments of the present application to obtain an index point sorting result, where each index point represents an article. The multiple index points are saved in the index memory according to the index point sorting result.
[0094] The terminal device 101 displays a search interface of an instant messaging application, as Figure 2 shown. The search interface includes a search box 201, a search type 202, and recommended searches 203. The user enters a target term "XX Park" in the search box 201 and submits it. The terminal device 101 sends a search request carrying the target term to the server 102.
[0095] Server 102 searches for k nearest neighbor index points of the target term from the HNSW graph. In each search, it obtains candidate index points from the index memory and multiple other index points continuously saved with the current index node (in the first search, any index point at the top level of the HNSW graph is used as the candidate index point), and saves the multiple other index points in the CPU cache.
[0096] Since each index point in the HNSW graph is saved in the index memory according to the index point sorting result, it is highly probable that the multiple other index points continuously saved with the candidate index node will include the neighbor index points of the candidate index node.
[0097] For a candidate index point, calculate the distance between the candidate index point and the target entry. At the same time, obtain each neighbor index point of the candidate index point from the CPU cache, and calculate the distance between each neighbor index point and the target entry. Then, select the index point with the closest distance to the target entry from the candidate index point and the neighbor index points as the new candidate index point, and sequentially execute the subsequent search process until the search end condition is met, and select k nearest neighbor index points from each index node serving as the candidate index point.
[0098] Assume k = 2, and the articles corresponding to the k nearest neighbor index points are Article A and Article B respectively. Server 102 sends the relevant information of Article A and Article B to terminal device 101, and terminal device 101 displays the search result interface of the instant messaging application, as Figure 3 shown. The search result interface displays the relevant information of Article A in the first area 301 and the relevant information of Article B in the second area 302.
[0099] Based on Figure 1 the system architecture diagram shown, an embodiment of the present application provides a process of an index point offline sorting method, as Figure 4 shown. The process of this method is executed by a computer device, and the computer device can be Figure 1 the terminal device 101 and / or server 102 shown, and includes the following steps:
[0100] Step S401, select an index point from multiple index points of the search graph as the initial sorted index point and add it to the sliding window.
[0101] Specifically, the multiple index points of the search graph are unsorted index points. The search graph can be an HNSW graph, an NSW graph, etc. An index point can represent a video, or an article, or a commodity, etc. When an index point represents a video, the index point includes, but is not limited to, index data such as video category, video title, video details, etc. When an index point represents an article, the index point includes, but is not limited to, index data such as article category, article title, article details, etc. When an index point represents a commodity, the index point includes, but is not limited to, index data such as commodity category, commodity name, commodity price, etc.
[0102] In the initial stage of index point sorting, there are no sorted index points in the sliding window. Therefore, an index point can be randomly selected from multiple index points as the initial sorted index point and added to the sliding window, where the size of the sliding window can be fixed or dynamically changed.
[0103] Step S402, determine the target correlation degree between each retained index point and the sorted index points in the sliding window.
[0104] Step S403: Select target index points from the remaining index points based on respective target relevance degrees, and add them to the sliding window.
[0105] Specifically, the target relevance degree can also be referred to as the target tightness degree. The larger the target relevance degree is, the closer the relationship between the index point and the sorted index points in the sliding window is; the smaller the target relevance degree is, the more distant the relationship between the index point and the sorted index points in the sliding window is.
[0106] Iteratively execute the above-mentioned Step S402 and Step S403 until multiple index points in the search graph are added to the sliding window and then end.
[0107] Step S404: Take the order in which multiple index points are added to the sliding window as the index point sorting result, and save the multiple index points in the index memory according to the index point sorting result.
[0108] Optionally, after saving the multiple index points in the index memory according to the index point sorting result, during the search process, the computer device obtains candidate index points for the search condition and at least one other index point continuously saved with the candidate index point from the index memory, and then saves the at least one other index point in the central processing unit cache. If the at least one other index point includes a neighbor index point of the candidate index point, obtain the neighbor index point of the candidate index point from the central processing unit cache. If the at least one other index point does not include a neighbor index point of the candidate index point, obtain the neighbor index point of the candidate index point from the index memory. Determine the search result of the search condition based on the candidate index point and the neighbor index point of the candidate index point.
[0109] Specifically, the search condition can be text, image, audio, etc., and the number of at least one other index point can be the upper limit value of the index points that can be saved in the CPU cache. When the search graph is an HNSW graph or an NSW graph, multiple iterative searches are required during the search process. During each iterative search, the above method is used to obtain candidate index points and respective neighbor index points of the candidate index points, and then calculate the distances between the candidate index points and the neighbor index points of the candidate index points and the search condition respectively, and take the index point with the closest distance as the new candidate index point. Iteratively execute the subsequent search process until the search end condition is met to obtain the search result of the search condition.
[0110] For example, set the size of the CPU cache to 4 index points, and index point 7 is the neighbor index point of index point 1.
[0111] If the above-mentioned index points are not sorted, the above-mentioned index points will be randomly saved in the index memory, specifically as Figure 5 shown, where the storage positions of index point 1 and index point 7 in the memory are relatively far apart.
[0112] During the search process, when the CPU of the computer device retrieves index point 1 from the index memory, it will additionally retrieve index points 2, 3, 4, and 5 and store them in the CPU cache. Since the CPU cache does not contain index point 7, if the CPU needs to retrieve the neighbor index point of index point 1 (i.e., index point 7), it needs to retrieve it from the index memory.
[0113] After sorting the above-mentioned index points by using the index point offline sorting method in the embodiment of the present application, the obtained index point sorting result is: index point 5, index point 3, index point 1, index point 7, index point 6, index point 2, index point 4. According to the index point sorting result, the above-mentioned index points are stored in the index memory, specifically as Figure 6 shown, where the storage positions of index point 1 and index point 7 in the memory are adjacent.
[0114] During the search process, when the CPU retrieves index point 1 from the index memory, it will additionally index index points 7, 6, 2, and 4 and store them in the CPU cache. Since the CPU cache contains index point 7, if the CPU needs to retrieve the neighbor index point of index point 1 (i.e., index point 7), it can directly retrieve it from the CPU cache.
[0115] In the embodiment of the present application, based on the target correlation degree between multiple index points, the multiple index points are sorted to obtain an index point sorting result, and then according to the index point sorting result, the multiple index points are stored in the index memory, so as to realize centralized storage of associated index points in the index memory. Therefore, during the search process, multiple associated index points can be obtained for calculation each time the memory is accessed, thereby reducing the number of memory accesses and further improving the search speed. Secondly, during the sorting process, the correlation degree between the unsorted index points and the sorted index points in the sliding window is calculated, avoiding calculating the correlation degree between the unsorted index points and all the sorted index points, reducing the process of repeatedly calculating the correlation degree, and thus improving the efficiency of index point sorting.
[0116] Optionally, after each iteration, if the number of sorted index points in the sliding window is greater than a preset threshold, the earliest added sorted index point is removed from the sliding window.
[0117] Specifically, the preset threshold is determined based on the central processing unit cache, and the preset threshold is the upper limit value of the index points that can be stored in the CPU cache. After the number of sorted index points in the sliding window reaches the preset threshold, if one more index point is added to the sliding window, it will cause the sliding window to overflow. At this time, the earliest added sorted index point is removed from the sliding window. The sorted index points removed from the sliding window are stored in an array in the order of removal from the sliding window.
[0118] In the embodiments of the present application, the size of the sliding window is controlled within a certain range, and then by determining the target correlation degree between the unsorted index points and the sorted index points within the sliding window, the unsorted index points are sorted, without the need to determine the target correlation degree between the unsorted index points and all the sorted index points to sort the unsorted index points, avoiding the repeated calculation process, thereby improving the efficiency of index point sorting and avoiding the waste of computing resources.
[0119] Optionally, in each iteration process, according to each target correlation degree, each index point is sorted to obtain a target correlation sorting result. Then, based on the target correlation sorting result, a target index point is selected from each index point as a sorted index point and added to the sliding window.
[0120] Specifically, each index point can be sorted in descending order of the target correlation degree to obtain a target correlation sorting result. Then, the first index point in the target correlation sorting result is used as the target index point and added to the sliding window. Alternatively, each index point can be sorted in ascending order of the target correlation degree to obtain a target correlation sorting result. Then, the last index point in the target correlation sorting result is used as the target index point and added to the sliding window.
[0121] For example, it is set that the obtained multiple unsorted index points are index point 1, index point 2, index point 3, index point 4, index point 5, index point 6, and index point 7.
[0122] See Figure 7 , which is a schematic diagram of the processing result after the end of the 5th iteration process. Among them, the sorted index points include index point 5, index point 3, index point 1, index point 7, and index point 6. Among them, index point 5 and index point 3 have slid out of the sliding window, and index point 1, index point 7, and index point 6 are located within the sliding window. The unsorted index points include index point 2 and index point 4.
[0123] In the 6th iteration process, the target correlation degree 1 between index point 2 and the sorted index points (index point 1, index point 7, index point 6) within the sliding window is determined; the target correlation degree 2 between index point 4 and the sorted index points (index point 1, index point 7, index point 6) within the sliding window is determined.
[0124] Sorting each unsorted index point in descending order of the target correlation degree, the obtained target correlation sorting result is index point 2, index point 4. Then, index point 2 is added to the sliding window to obtain the processing result after the end of the 6th iteration process, specifically as Figure 8As shown, the sorted index points include index point 5, index point 3, index point 1, index point 7, index point 6, and index point 2. Among them, index point 5, index point 3, and index point 1 have slid out of the sliding window, and index point 7, index point 6, and index point 2 are within the sliding window. The unsorted index points include index point 4.
[0125] In the embodiment of the present application, according to each target correlation degree, each unsorted index point is sorted to obtain a target correlation sorting result for characterizing the tightness relationship between each unsorted index point and the sorted index points. Based on the target correlation sorting result, an unsorted index point is selected from each unsorted index point as a sorted index point and added to the sliding window, which can effectively add the unsorted index points with a high tightness with the sorted index points to the sliding window first. This also makes the index points continuously saved in the index memory be index points with a close relationship according to the sorting result, thereby improving the search speed during the search process.
[0126] Optionally, in each iteration process, the embodiment of the present application at least adopts the following implementation methods to respectively determine the target correlation degrees between each retained index point and the sorted index points in the sliding window:
[0127] After selecting an index point from multiple index points as the initial sorted index point and adding it to the sliding window, the first iteration process is executed.
[0128] In the first iteration process, the embodiment of the present application at least adopts the following implementation methods to determine the target correlation degrees between each retained index point and the initial sorted index point in the sliding window:
[0129] Specifically, for any two index points, based on the connection attribute and the number of common neighbor index points between the two index points, the association score between the two index points can be determined. Among them, the connection attribute includes: no connection, directed graph single connection, undirected graph connection, and directed graph double connection. The value corresponding to no connection is 0, the value corresponding to directed graph single connection is 1, and the value corresponding to undirected graph connection or directed graph double connection is 2. It should be noted that the embodiment of the present application can also use other values to represent the connection attribute. In this regard, the present application does not specifically limit it.
[0130] The common neighbor index point refers to the neighbor index points with the same two index nodes, and the connection edges between the neighbor index point and the two index nodes respectively point to the two index nodes.
[0131] For example, such as Figure 9As shown, index point 1 and index point 2 are respectively connected to index point 3. Among them, the connection edge between index point 3 and index point 1 points from index point 3 to index point 1, and the connection edge between index point 2 and index point 1 points from index point 2 to index point 1. Then, index point 3 is the common neighbor index point of index point 1 and index point 2.
[0132] Sum the value corresponding to the connection attribute between two index points and the number of common neighbor index points to obtain the association score between the two index points, as specifically shown in the following formula (1):
[0133] S(u, v) = S s (u, V) + S n (u, v) …………… (1)
[0134] Among them, S(u, v) represents the association score between index point s and index point v, and S s (u, v) represents the number of common neighbor index points between index point s and index point v, and S n (u, v) represents the value corresponding to the connection attribute between index point s and index point v.
[0135] In the first iteration process, use the above formula (1) to obtain the association scores between the index points and the initially sorted index points within the sliding window. Then use a transformation function to determine the target association degree between the index points and the sorted index points within the sliding window. The transformation function is specifically shown in the following formula (2):
[0136]
[0137] Among them, represents a permutation function, ω represents the size of the sliding window, represents the target association degree.
[0138] For the second iteration process and the iteration processes after the second iteration process, the embodiments of the present application at least adopt the following several implementation manners to determine the target association degrees between each retained index point and the sorted index points in the sliding window:
[0139] Implementation manner 1: For each index point, respectively execute the following steps:
[0140] Obtain the historical association degree between an index point and the sorted index points in the sliding window in the previous iteration process. Then determine the first sub-association degree between this index point and the target index point newly added to the sliding window in the previous iteration process. Then, based on the first sub-association degree and the historical association degree, determine the target association degree between this index point and the sorted index points in the sliding window.
[0141] Specifically, the historical relevance corresponding to an index point refers to the target relevance of the index point in the previous iteration process. If an index point is associated with the target index point newly added to the sliding window, the corresponding first sub-relevance is 1. If an index point is not associated with the target index point newly added to the sliding window, the corresponding first sub-relevance is 0. Of course, embodiments of the present application may also use other values to represent the first sub-relevance, and the present application does not make specific limitations in this regard.
[0142] Sum the historical relevance and the first sub-relevance of the index point to obtain the target relevance between the index point and the sorted index points in the sliding window.
[0143] For example, assume that the size of the sliding window is 3 index points, and the multiple unsorted index points obtained are index point 1, index point 2, index point 3, index point 4, index point 5, index point 6, and index point 7.
[0144] See Figure 10 , which is a schematic diagram of the processing result after the end of the 3rd iteration process. Among them, the sorted index points include index point 1, index point 3, and index point 6, and the target relevance sorting results of the remaining unsorted index points are: index point 5, index point 2, index point 4, index point 7, where index point 6 is the index point newly added to the sliding window.
[0145] Before adding index point 6 to the sliding window, the target relevance between index point 2 and the sorted index points (index point 1 and index point 3) in the sliding window is F 32 ; the target relevance between index point 4 and the sorted index points (index point 1 and index point 3) in the sliding window is F 34 ; the target relevance between index point 5 and the sorted index points (index point 1 and index point 3) in the sliding window is F 35 ; the target relevance between index point 6 and the sorted index points (index point 1 and index point 3) in the sliding window is F 36 ; the target relevance between index point 7 and the sorted index points (index point 1 and index point 3) in the sliding window is F 37 .
[0146] During the execution of the 4th iteration process, the target relevance obtained in the 3rd iteration process can be used as the historical relevance. For index point 5, first obtain the historical relevance F between index point 5 and the sorted index points in the sliding window in the previous iteration process 35 , and then determine the first sub-relevance T(5, 6) between index point 5 and index point 6. Sum the historical relevance F 35 and the first sub-relevance T(5, 6) to obtain the target relevance F corresponding to index point 5 in the 4th iteration process 45 .
[0147] For index point 2, first obtain the historical correlation degree F between index point 2 and the sorted index points in the sliding window during the previous iteration 32 , and then determine the first sub - correlation degree T(2, 6) between index point 2 and index point 6. Sum the historical correlation degree F 32 and the first sub - correlation degree T(2, 6) to obtain the target correlation degree F corresponding to index point 2 during the 4th iteration 42 .
[0148] For index point 4, first obtain the historical correlation degree F between index point 4 and the sorted index points in the sliding window during the previous iteration 34 , and then determine the first sub - correlation degree T(4, 6) between index point 4 and index point 6. Sum the historical correlation degree F 34 and the first sub - correlation degree T(4, 6) to obtain the target correlation degree F corresponding to index point 4 during the 4th iteration 44 .
[0149] For index point 7, first obtain the historical correlation degree F between index point 7 and the sorted index points in the sliding window during the previous iteration 37 , and then determine the first sub - correlation degree T(7, 6) between index point 7 and index point 6. Sum the historical correlation degree F 37 and the first sub - correlation degree T(7, 6) to obtain the target correlation degree F corresponding to index point 7 during the 4th iteration 47 .
[0150] Sort each unsorted index point in descending order according to the target correlation degree to obtain the target correlation sorting result as: index point 2, index point 5, index point 4, index point 7. Then add index point 2 to the sliding window and slide index point 1 out of the sliding window to obtain the index point sorting result after the 4th iteration, specifically as Figure 11 shown. The sorted index points include index point 1, index point 3, index point 6, index point 2, where index point 1 has been slid out of the sliding window, and index point 3, index point 6, index point 2 are within the sliding window. The unsorted index points include index point 5, index point 4, index point 7.
[0151] In the embodiment of the present application, since only one new unsorted index point is added to the sliding window in the previous iteration process, and the originally sorted index points in the sliding window remain unchanged, the historical correlation degree corresponding to the unsorted index point in the previous iteration process can be used. Only the first sub-correlation degree between the unsorted index point and the unsorted index point newly added to the sliding window needs to be calculated, and then the target correlation degree of the unsorted index point in the current iteration process can be obtained based on the historical correlation degree and the first sub-correlation degree, thereby reducing the amount of calculation and improving the efficiency of index point sorting.
[0152] Embodiment 2: For each index point, the following steps are respectively executed:
[0153] Obtain the historical correlation degree between an index point and the sorted index points in the sliding window in the previous iteration process. Then determine the first sub-correlation degree between this index point and the target index point newly added to the sliding window in the previous iteration process.
[0154] If the earliest added sorted index point is removed from the sliding window after the previous iteration process, then determine the second sub-correlation degree between this index point and the earliest added sorted index point. Then, based on the first sub-correlation degree, the second sub-correlation degree, and the historical correlation degree, determine the target correlation degree between this index point and the sorted index points in the sliding window.
[0155] Specifically, if an index point is associated with the sorted index point removed from the sliding window, the corresponding second sub-correlation degree is 1. If an index point is not associated with the sorted index point removed from the sliding window, the corresponding second sub-correlation degree is 0. Of course, in the embodiment of the present application, other values can also be used to represent the second sub-correlation degree, and the present application does not make specific limitations in this regard.
[0156] Sum the historical correlation degree and the first sub-correlation degree of this index point to obtain an intermediate correlation degree. Then subtract the second sub-correlation degree from the intermediate correlation degree to obtain the target correlation degree between this index point and the sorted index points in the sliding window in the current iteration process.
[0157] For example, referring to Figure 11 , it is a schematic diagram of the processing result after the end of the 4th iteration process. In this figure, the sorted index points include index point 1, index point 3, index point 6, and index point 2. Among them, index point 1 has slid out of the sliding window, and index point 3, index point 6, and index point 2 are located in the sliding window. The unsorted index points include index point 5, index point 4, and index point 7.
[0158] During the execution of the 5th iteration, the target correlation degrees obtained in the 4th iteration can all be used as historical correlation degrees. For index point 5, first obtain the historical correlation degree F between index point 5 and the sorted index points in the sliding window in the previous iteration 45 , then determine the first sub-correlation degree T(5, 2) between index point 5 and index point 2, and determine the second sub-correlation degree R(5, 1) between index point 5 and index point 1. Then sum the historical correlation degree F 45 and the first sub-correlation degree T(5, 2), and then subtract the second sub-correlation degree R(5, 1) to obtain the corresponding target correlation degree F of index point 5 in the 5th iteration 55 .
[0159] For index point 4, first obtain the historical correlation degree F between index point 4 and the sorted index points in the sliding window in the previous iteration 44 , then determine the first sub-correlation degree T(4, 2) between index point 4 and index point 2, and the second sub-correlation degree R(4, 1) between index point 4 and index point 1. Then sum the historical correlation degree F 44 and the first sub-correlation degree T(4, 2), and then subtract the second sub-correlation degree R(4, 1) to obtain the corresponding target correlation degree F of index point 4 in the 5th iteration 54 .
[0160] For index point 7, first obtain the historical correlation degree F between index point 7 and the sorted index points in the sliding window in the previous iteration 47 , then determine the first sub-correlation degree T(7, 2) between index point 7 and index point 2, and the second sub-correlation degree R(7, 1) between index point 7 and index point 1. Then sum the historical correlation degree F 47 and the first sub-correlation degree T(7, 2), and then subtract the second sub-correlation degree R(7, 1) to obtain the corresponding target correlation degree F of index point 7 in the 5th iteration 57 .
[0161] Sort the unsorted index points in descending order of the target correlation degree. The obtained target correlation sorting result is: index point 4, index point 5, index point 7. Then add index point 4 to the sliding window, and at the same time slide index point 3 out of the sliding window to obtain the index point sorting result after the 5th iteration. Specifically as Figure 12 shown, the sorted index points include index point 1, index point 3, index point 6, index point 2, index point 4, where index point 1 and index point 3 have slid out of the sliding window, and index point 6, index point 2, and index point 4 are located within the sliding window. The unsorted index points include index point 5 and index point 7
[0162] In the embodiment of the present application, since a new unsorted index point is added to the sliding window in the previous iteration process, and a sorted index point is removed from the sliding window at the same time, the originally sorted index points in the sliding window do not change. Therefore, the historical correlation degree corresponding to the unsorted index point in the previous iteration process can be used. Then, calculate the first sub-correlation degree between the unsorted index point and the newly added unsorted index point in the sliding window, and the first sub-correlation degree between the unsorted index point and the sorted index point removed from the sliding window. Then, based on the historical correlation degree, the first sub-correlation degree, and the second historical correlation degree, the target correlation degree of the unsorted index point in the current iteration process can be obtained, thereby reducing the amount of calculation and improving the efficiency of index point sorting.
[0163] Embodiment 3: For each index point, the following steps are respectively executed:
[0164] Determine the correlation scores of each sorted index point in the sliding window with an index point respectively, and then based on the obtained correlation scores, determine the target correlation degree between the index point and the sorted index points in the sliding window.
[0165] Specifically, in each iteration process, for each unsorted index point, based on the connection attribute and the number of common neighbor index points between the unsorted index point and each sorted index point in the sliding window, determine the correlation score between the unsorted index point and each sorted index point, as specifically shown in the above formula (1). Then, use the conversion function shown in the above formula (2) to determine the target correlation degree between the unsorted index point and the sorted index points in the sliding window based on the obtained correlation scores.
[0166] For example, refer to Figure 11 , which is a schematic diagram of the processing result after the end of the 4th iteration process. In this figure, the sorted index points include index point 1, index point 3, index point 6, and index point 2. Among them, index point 1 has slid out of the sliding window, and index point 3, index point 6, and index point 2 are located in the sliding window. The unsorted index points include index point 5, index point 4, and index point 7.
[0167] In the execution of the 5th iteration process, for index point 5, use the above formula (1) to respectively determine the correlation score S(5, 3) between index point 5 and index point 3, the correlation score S(5, 6) between index point 5 and index point 6, and the correlation score S(5, 2) between index point 5 and index point 2. Then, use the above formula (2) and the obtained correlation scores to determine the target correlation degree F corresponding to index point 5 in the 5th iteration process 55 .
[0168] For index point 4, using the above formula (1), respectively determine the association score S(4, 3) between index point 4 and index point 3, the association score S(4, 6) between index point 4 and index point 6, and the association score S(4, 2) between index point 4 and index point 2. Then, using the above formula (2) and the obtained association scores, determine the corresponding target association degree F of index point 4 in the 5th iteration process 54 。
[0169] For index point 7, using the above formula (1), respectively determine the association score S(7, 3) between index point 7 and index point 3, the association score S(7, 6) between index point 7 and index point 6, and the association score S(7, 2) between index point 7 and index point 2. Then, using the above formula (2) and the obtained association scores, determine the corresponding target association degree F of index point 7 in the 5th iteration process 57 。
[0170] Sort each unsorted index point in descending order of the target association degree to obtain the target association sorting result: index point 4, index point 5, index point 7. Then add index point 4 to the sliding window, and at the same time slide index point 3 out of the sliding window to obtain the index point sorting result after the 5th iteration process. Specifically, as Figure 12 shown, the sorted index points include index point 1, index point 3, index point 6, index point 2, index point 4. Among them, index point 1 and index point 3 have slid out of the sliding window, and index point 6, index point 2, and index point 4 are located within the sliding window. The unsorted index points include index point 5 and index point 7
[0171] In the embodiment of the present application, in each iteration process, based on the connection attributes and the number of common neighbor index points between the unsorted index points and the sorted index points in the sliding window, determine the association scores between the unsorted index points and the sorted index points in the sliding window. Then, based on the association scores and the conversion function, determine the target association degrees between the unsorted index points and the sorted index points in the sliding window, ensuring the accuracy of the target association degrees
[0172] Optionally, since in each iteration process, only one index point is newly added to the sliding window, and when the sliding window overflows, one sorted index point is removed, and the other sorted index points in the sliding window do not change. Therefore, compared with the previous iteration process, the change in the target association sorting result of each unsorted index point in this iteration process is not very large. Therefore, the historical association sorting results of each unsorted index point in the previous iteration process can be adjusted to obtain the target association sorting result in this iteration process
[0173] In view of this, in the embodiments of the present application, each index point is obtained, and the historical association sorting result corresponding to the previous iteration process is obtained. Then, based on each target association degree, the historical association sorting result is adjusted to obtain the target association sorting result.
[0174] Specifically, from the target association sorting result obtained in the previous iteration process, the target index points newly added to the sliding window in the previous iteration process are removed to obtain the historical association sorting result.
[0175] In the current iteration process, for each index point (hereinafter referred to as the target unsorted index point), when the first sub-association degree between the target unsorted index point and the target index point newly added to the sliding window is 1, the target association degree corresponding to the target unsorted index point is incremented by 1, and then it is determined whether there is a first replacement index point in the historical association sorting result. If so, the positions of the target unsorted index point and the first replacement index point in the historical association sorting result are replaced to complete one sorting; otherwise, the historical association sorting result is not adjusted.
[0176] Among them, when the sorting rule corresponding to the historical association sorting result is: sorting from largest to smallest according to the target association degree, the first replacement index point needs to meet the following conditions:
[0177] The target association degree of the first replacement index point is less than the target association degree of the target unsorted index point, and the target association degree of the unsorted index point ranked one position before the first replacement index point is greater than or equal to the target association degree of the target unsorted index point.
[0178] When the first sub-association degree between the target unsorted index point and the target index point newly added to the sliding window is 0, the historical association sorting result is not adjusted. After sorting for each unsorted index point, the sorting complexity is O(MlogN), where M represents the number of edges associated with the unsorted index points newly added to the sliding window, and N represents the number of each unsorted index point.
[0179] Similarly, when the second sub-association degree between the target unsorted index point and the sorted index point that slid out of the sliding window in the previous iteration process is 1, the target association degree corresponding to the target unsorted index point is decremented by 1, and then it is determined whether there is a second replacement index point in the historical association sorting result. If so, the positions of the target unsorted index point and the second replacement index point in the historical association sorting result are replaced to complete one sorting; otherwise, the historical association sorting result is not adjusted.
[0180] Among them, when the sorting rule corresponding to the historical association sorting result is: sorting from largest to smallest according to the target association degree, the second replacement index point needs to meet the following conditions:
[0181] The target correlation degree of the second replacement index point is greater than that of the target unsorted index point, and the target correlation degree of the unsorted index point ranked immediately after the second replacement index point is less than or equal to that of the target unsorted index point.
[0182] When the second sub-correlation degree between the target unsorted index point and the sorted index point that slid out of the sliding window in the previous iteration is 0, the historical correlation sorting result is not adjusted. After sorting for each unsorted index point, the sorting complexity is O(LlogN), where L represents the number of edges associated with the sorted index points that slid out of the sliding window, and N represents the number of each unsorted index point.
[0183] For example, referring to Figure 13 , it is a schematic diagram of the processing result after the end of the 4th iteration. In this figure, the sorted index points include index point 1, index point 3, index point 6, and index point 2, where index point 1 has slid out of the sliding window; index point 3, index point 6, and index point 2 are within the sliding window; the unsorted index points include: index point 5, index point 4, and index point 7. In the 4th iteration, the target correlation degree F 45 corresponding to index point 5 is 5, the target correlation degree F 44 corresponding to index point 4 is 5, and the target correlation degree F 47 corresponding to index point 7 is 4.
[0184] In the 4th iteration, index point 2 is newly added to the sliding window, then in the 5th iteration, the following operations are performed:
[0185] For index point 5, first obtain the historical correlation degree F 45 = 5 between index point 5 and the sorted index points in the sliding window in the previous iteration. Since the first sub-correlation degree between index point 5 and index point 2 is T(5, 2) = 0, therefore, after newly adding index point 2 to the sliding window, the historical correlation sorting result (historical correlation sorting result: index point 5, index point 4, index point 7) is not adjusted, and the target correlation degree F 55 corresponding to index point 5 in the 5th iteration is 5.
[0186] For index point 4, first obtain the historical correlation degree F 44 = 5 between index point 4 and the sorted index points in the sliding window in the previous iteration. Since the first sub-correlation degree between index point 4 and index point 2 is T(4, 2) = 1, then the target correlation degree F 54 corresponding to index point 4 in the 5th iteration is 5 + 1 = 6. At this time, as Figure 14a shows, the target correlation degrees corresponding to index point 5, index point 4, and index point 7 are respectively: 5, 6, and 4.
[0187] The first replacement index point in the historical association sorting result is obtained by binary search as index point 5. Then, the positions of replacement index point 5 and index point 4 in the historical association sorting result are swapped. The replacement result is as follows Figure 14b shown. The updated historical association sorting result after replacement is: index point 4, index point 5, index point 7, and their respective target association degrees are: 6, 5, 4.
[0188] For index point 7, first obtain the historical association degree F between index point 7 and the sorted index points in the sliding window during the previous iteration 47 = 4. Since the first sub - association degree between index point 7 and index point 2 is T(7, 2) = 0, therefore, after newly adding index point 2 to the sliding window, the historical association sorting result is not adjusted. During the 5th iteration, the target association degree F corresponding to index point 7 57 = 4.
[0189] In addition, during the 4th iteration, index point 1 is removed from the sliding window. Then, during the 5th iteration, the following operations are performed:
[0190] For index point 4, since the second sub - association degree between index point 4 and index point 1 is R(4, 1) = 0, therefore, after removing index point 1 from the sliding window, the historical association sorting result is not adjusted, and the target association degree F is not updated 54 .
[0191] For index point 5, since the second sub - association degree R(5, 1) = 1 between index point 5 and index point 1, then during the 5th iteration, subtract 1 from the target association degree F of index point 5 55 to obtain the updated target association degree F 55 = 5 - 1 = 4. At this time, as Figure 14c shown, the target association degrees corresponding to index point 4, index point 5, and index point 7 are: 6, 4, 4 respectively. The second replacement index point in the historical association sorting result is not obtained by binary search, so the historical association sorting result is not adjusted.
[0192] For index point 7, since the second sub - association degree between index point 7 and index point 1 is R(7, 1) = 0, therefore, after removing index point 1 from the sliding window, the historical association sorting result is not adjusted, and the target association degree F is not updated 57 .
[0193] After the adjustment is completed, the target association sorting result is: index point 4, index point 5, index point 7. Then, add index point 4 to the sliding window, and at the same time slide index point 3 out of the sliding window to obtain the index point sorting result after the 5th iteration. Specifically as Figure 12As shown, the sorted index points include index point 1, index point 3, index point 6, index point 2, and index point 4. Among them, index point 1 and index point 3 have slid out of the sliding window, and index point 6, index point 2, and index point 4 are within the sliding window. The unsorted index points include index point 5 and index point 7.
[0194] In the embodiments of the present application, in each iteration process, when a newly added unsorted index point is added to the sliding window or a sorted index point is removed from the sliding window, the target association degree of each associated unsorted index point is updated accordingly. After updating the target association degree once, a replacement method is used to adjust the historical association sorting result. Compared with the traditional sorting method, the sorting complexity is reduced, thereby improving the sorting efficiency.
[0195] It should be noted that in the embodiments of the present application, the method for adjusting the historical association sorting result to obtain the target association sorting result is not limited to the one described above. For each unsorted index point, based on the first sub-association degree between the unsorted index point and the newly added unsorted index point, and the second sub-association degree with the removed sorted index point, after determining the target association degree of the unsorted index point, it is then determined whether it is necessary to adjust the historical association sorting result. If adjustment is required, the position of the unsorted index point and other unsorted index points in the historical association sorting result is replaced to complete the sorting. Regarding this, the present application does not make specific limitations.
[0196] To better explain the embodiments of the present application, the following introduces an off-line sorting method for index points provided by the embodiments of the present application in combination with a search scenario. The search scenario includes an off-line data processing stage and an on-line search stage, which are executed by a computer device. The computer device can be Figure 1 the terminal device 101 and / or the server 102 shown. As Figure 15 shown, the off-line data processing stage includes constructing an HNSW graph and saving index points, and the on-line search stage includes obtaining search conditions, search condition analysis, and on-line search. The following will elaborate on each stage:
[0197] First, the process of constructing the HNSW graph is introduced, including the following steps: initializing the HNSW graph, and then adding index points to the initialized HNSW graph in an iterative manner until the iteration end condition is met to obtain the constructed HNSW graph. The iteration end condition can be that all index points are added to the HNSW graph. Each iteration process is as Figure 16 shown, including the following steps:
[0198] Step S1601, obtain the index point q to be added.
[0199] Step S1602, determine the target level i where the index point q falls through a random function.
[0200] Among them, the HNSW graph includes multiple levels, each level corresponding to a level number. From top to bottom, the corresponding level numbers of the multiple levels decrease in sequence, and the index points included in each level increase in sequence. The target level i is a level in the HNSW graph, where 0 ≤ i ≤ k. Here, 0 represents the level number corresponding to the bottommost level in the HNSW graph, and k represents the level number corresponding to the topmost level in the HNSW graph.
[0201] Step S1603, determine whether i is less than k. If so, execute step S1604; otherwise, execute step S1608.
[0202] Step S1604, set the current processing level Kc = k.
[0203] Step S1605, find the index point x closest to the index point q in the current processing level.
[0204] Step S1606, enter the next level via the index point x, and set the current processing level Kc = Kc - 1.
[0205] Step S1607, determine whether Kc is equal to i. If so, execute step S1608; otherwise, execute step S1605.
[0206] Step S1608, find the set of neighbor index points of the index point q in the current processing level Kc.
[0207] Step S1609, insert the index point q into the current processing level Kc, and connect the index point q to each neighbor index point in the set of neighbor index points.
[0208] Step S1610, set the current processing level Kc = Kc - 1.
[0209] Step S1611, determine whether Kc is equal to 0. If so, execute step S1612; otherwise, execute step S1608.
[0210] Step S1612, the addition of the index point q is completed.
[0211] In one example, the HNSW graph obtained by adopting the above method is as Figure 17, the HNSW graph includes three levels. Level 2 is the topmost level, Level 1 is the middle level, and Level 0 is the bottommost level. Among them, Level 0 includes 8 index points, namely Index Point 1, Index Point 2, Index Point 3, Index Point 4, Index Point 5, Index Point 6, Index Point 7, and Index Point 8. Level 1 includes Index Point 1, Index Point 2, Index Point 6, and Index Point 8, which has 4 fewer index points than Level 2. Level 0 includes Index Point 1 and Index Point 6, which has 2 fewer index points than Level 1. During the search process, for the search condition, starting from Index Point 6 in Level 2, perform a layer-by-layer search on each level in the order from top to bottom until Index Point 3 in Level 0 to obtain the search result. It should be understood that the above Figure 17 is only an exemplary illustration. In actual applications, the number of index points included in each layer and the connection relationship between index points are not limited to this.
[0212] In the embodiments of the present application, all the index points in the bottommost level of the above HNSW graph can be sorted to obtain an index point sorting result. According to the index point sorting result, then all the index points in the bottommost level of the HNSW graph are saved in the index memory, as Figure 18 shown, including the following steps:
[0213] Step S1801, obtain all the index points in the bottommost level of the HNSW graph as unsorted index points.
[0214] Step S1802, select an unsorted index point from multiple unsorted index points as an initial sorted index point and add it to the sliding window.
[0215] Step S1803, determine the target correlation degree between each remaining unsorted index point and the sorted index points in the sliding window.
[0216] Step S1804, sort each unsorted index point according to each target correlation degree to obtain a target correlation sorting result.
[0217] Step S1805, based on the target correlation sorting result, select an unsorted index point from each unsorted index point as a sorted index point and add it to the sliding window.
[0218] Step S1806, determine whether the number of remaining unsorted index points is 0. If so, execute Step S1809; otherwise, execute Step S1807.
[0219] Step S1807, determine whether the number of sorted index points in the sliding window is greater than a preset threshold. If so, execute Step S1808; otherwise, execute Step S1803.
[0220] Step S1808, remove the earliest added sorted index point from the sliding window, and execute Step S1803.
[0221] Step S1809, obtain the index point sorting result, and save multiple unsorted index points in the index memory according to the index point sorting result.
[0222] During the search process, assume that each neighboring index point represents an article, and the lowest level in the HNSW graph includes n index points. Based on the search condition, starting from the top level of the HNSW graph, perform a layer-by-layer search on each level in the top-down order until k nearest neighbor index points that meet the search condition are found among the n index points in the lowest level, and the articles corresponding to the k nearest neighbor index points are used as the search result of the search condition.
[0223] During each search process, obtain candidate index points from the index memory, as well as multiple other index points continuously saved with the current index node (when performing the first search, use any index point in the top level of the HNSW graph as the candidate index point), and save the multiple other index points in the CPU cache.
[0224] For the candidate index point, calculate the distance between the candidate index point and the search condition. At the same time, determine whether the CPU cache contains the neighbor index points of the candidate index point. If so, obtain each neighbor index point of the candidate index point from the CPU cache; otherwise, obtain each neighbor index point of the candidate index point from the index memory. Then calculate the distance between each neighbor index point and the search condition. Then select the index point with the shortest distance to the search condition from the candidate index point and the neighbor index points as the new candidate index point, and sequentially execute the subsequent search process.
[0225] In the embodiments of the present application, based on the target correlation degree between multiple index points in the HNSW graph, sort the multiple index points to obtain the index point sorting result, and then save the multiple index points in the index memory according to the index point sorting result, so as to realize the centralized storage of associated index points in the index memory. Therefore, during the HNSW graph search process, multiple associated index points can be obtained for calculation each time the memory is accessed, thereby reducing the number of memory accesses, further improving the HNSW graph search speed, and optimizing the HNSW graph search performance. Secondly, during the sorting process, calculate the correlation degree between the index point and the sorted index points in the sliding window, avoiding calculating the correlation degree between the index point and all sorted index points, reducing the process of repeatedly calculating the correlation degree, and thus improving the efficiency of index point sorting.
[0226] In addition, to prove the above effects of the index point offline sorting method in the embodiments of the present application, the present application conducted tests using index point data at the ten-million level. The test results show that when using the index point offline sorting method in the embodiments of the present application to sort index point data at the ten-million level, it only takes about 40 minutes, which is 75 times more efficient than the existing index point offline sorting method. Moreover, after saving the index points in the index memory according to the index point sorting results, the search time is reduced by 15% compared to before, and the recall rate is not affected.
[0227] Based on the same technical concept, the embodiments of the present application provide a structural schematic diagram of an index point offline sorting device, as Figure 19 shown. The device 1900 includes:
[0228] A selection module 1901, configured to select an index point from multiple index points of a search graph as an initial sorted index point and add it to a sliding window;
[0229] A sorting module 1902, configured to determine the target correlation degrees between each of the remaining index points and the sorted index points in the sliding window; select a target index point from each of the remaining index points based on the respective target correlation degrees and add it to the sliding window; iteratively execute the step of determining the target correlation degrees between each of the remaining index points and the sorted index points in the sliding window until all the multiple index points are added to the sliding window;
[0230] A storage module 1903, configured to use the order in which the multiple index points are added to the sliding window as the index point sorting result, and save the multiple index points in the index memory according to the index point sorting result.
[0231] Optionally, the sorting module 1902 is further configured to:
[0232] After each iteration, if the number of sorted index points in the sliding window is greater than a preset threshold, remove the earliest added sorted index point from the sliding window, and the preset threshold is determined based on the central processing unit cache.
[0233] Optionally, the sorting module 1902 is specifically configured to:
[0234] Sort the respective index points according to the respective target correlation degrees to obtain a target correlation sorting result;
[0235] Based on the target correlation sorting result, select a target index point from the respective index points as a sorted index point and add it to the sliding window.
[0236] Optionally, the sorting module 1902 is specifically configured to:
[0237] For each index point, the following steps are respectively executed:
[0238] Obtain the historical correlation degree between an index point and the sorted index points in the sliding window during the previous iteration;
[0239] Determine the first sub-correlation degree between the index point and the target index point newly added to the sliding window during the previous iteration;
[0240] Based on the first sub-correlation degree and the historical correlation degree, determine the target correlation degree between the index point and the sorted index points in the sliding window.
[0241] Optionally, the sorting module 1902 is specifically configured to:
[0242] If the earliest added sorted index point is removed from the sliding window after the previous iteration, determine the second sub-correlation degree between the index point and the earliest added sorted index point;
[0243] Based on the first sub-correlation degree, the second sub-correlation degree and the historical correlation degree, determine the target correlation degree between the index point and the sorted index points in the sliding window.
[0244] Optionally, the sorting module 1902 is specifically configured to:
[0245] For each index point, the following steps are respectively executed:
[0246] Determine the correlation scores between each sorted index point in the sliding window and an index point;
[0247] Based on the obtained correlation scores, determine the target correlation degree between the index point and the sorted index points in the sliding window.
[0248] Optionally, the sorting module 1902 is specifically configured to:
[0249] Obtain the historical correlation sorting results corresponding to the index points during the previous iteration;
[0250] Based on the target correlation degrees, adjust the historical correlation sorting results to obtain the target correlation sorting results.
[0251] Optionally, the index points in the target correlation sorting results are arranged in descending order of the target correlation degree;
[0252] The sorting module 1902 is specifically configured to:
[0253] Take the first index point in the target association sorting result as the target index point and add it to the sliding window.
[0254] Optionally, it further includes a search module 1904;
[0255] The search module 1904 is specifically configured to:
[0256] Obtain candidate index points of the search condition from the index memory, and at least one other index point continuously saved with the candidate index points;
[0257] Save the at least one other index point in the central processing unit cache;
[0258] If the at least one other index point includes a neighbor index point of the candidate index point, obtain the neighbor index point of the candidate index point from the central processing unit cache;
[0259] Determine the search result of the search condition based on the candidate index point and the neighbor index point of the candidate index point.
[0260] In the embodiments of the present application, based on the target association degree between multiple index points, multiple index points are sorted to obtain an index point sorting result, and then according to the index point sorting result, multiple index points are saved in the index memory, so as to realize that associated index points are centrally saved in the index memory. Therefore, during the search process, multiple associated index points can be obtained for calculation each time the memory is accessed, thereby reducing the number of times of accessing the memory, and further improving the search speed. Secondly, during the sorting process, the association degree between the index point and the sorted index points in the sliding window is calculated, avoiding calculating the association degree between the index point and all sorted index points, reducing the process of repeatedly calculating the association degree, and thus improving the efficiency of index point sorting.
[0261] Based on the same technical concept, the embodiments of the present application provide a computer device, which can be Figure 1 the terminal device and / or server shown in the figure, such as Figure 20 shown in the figure, including at least one processor 2001 and a memory 2002 connected to at least one processor. In the embodiments of the present application, the specific connection medium between the processor 2001 and the memory 2002 is not limited. Figure 20 Take the example that the processor 2001 and the memory 2002 are connected through a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.
[0262] In the embodiments of the present application, the memory 2002 stores instructions executable by at least one processor 2001. By executing the instructions stored in the memory 2002, at least one processor 2001 can execute the steps of the above index point offline sorting method.
[0263] Among them, the processor 2001 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and lines, and realize the sorting of index points by running or executing the instructions stored in the memory 2002 and calling the data stored in the memory 2002. Optionally, the processor 2001 may include one or more processing units. The processor 2001 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 2001. In some embodiments, the processor 2001 and the memory 2002 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.
[0264] The processor 2001 can be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0265] The memory 2002, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 2002 can include at least one type of storage medium. For example, it can include flash memory, hard disks, multimedia cards, card-type memories, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memories, magnetic disks, optical discs, and so on. The memory 2002 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer device, but is not limited thereto. The memory 2002 in the embodiments of the present application can also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0266] Based on the same inventive concept, embodiments of the present application provide a computer-readable storage medium storing a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the above-mentioned index point offline sorting method.
[0267] Based on the same inventive concept, embodiments of the present application provide a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer device, the computer device is caused to execute the steps of the above-mentioned index point offline sorting method.
[0268] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0269] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and combinations of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0270] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0271] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0272] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0273] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. An off-line sorting method for index points, characterized in that, it includes: Select an index point from multiple index points of a search graph as an initial sorted index point and add it to a sliding window; Determine the target correlation degrees between each of the remaining index points and the sorted index points in the sliding window; Based on each target correlation degree, select a target index point from the remaining index points and add it to the sliding window; Iteratively execute the step of determining the target correlation degrees between each of the remaining index points and the sorted index points in the sliding window until the multiple index points are added to the sliding window; Take the order of adding the multiple index points to the sliding window as the index point sorting result, and save the multiple index points in an index memory according to the index point sorting result.
2. The method according to claim 1, characterized in that, it further includes: After each iteration process, if the number of sorted index points in the sliding window is greater than a preset threshold, remove the earliest added sorted index point from the sliding window, and the preset threshold is determined based on a central processing unit cache.
3. The method according to claim 1, characterized in that, The step of selecting a target index point from the remaining index points based on each target correlation degree and adding it to the sliding window includes: Sort the index points according to each target correlation degree to obtain a target correlation sorting result; Based on the target correlation sorting result, select a target index point from the index points as a sorted index point and add it to the sliding window.
4. The method according to claim 3, characterized in that, The step of determining the target correlation degrees between each of the remaining index points and the sorted index points in the sliding window includes: For each index point, respectively execute the following steps: Obtain the historical correlation degree between an index point and the sorted index points in the sliding window in the previous iteration process; Determine the first sub-correlation degree between the index point and the target index point newly added to the sliding window in the previous iteration process; Based on the first sub-correlation degree and the historical correlation degree, determine the target correlation degree between the index point and the sorted index points in the sliding window.
5. The method according to claim 4, characterized in that, The step of determining the target correlation degree between an index point and the sorted index points in the sliding window based on the first sub-correlation degree and the historical correlation degree includes: If the earliest added sorted index point is removed from the sliding window after the previous iteration process, determine the second sub-correlation degree between the index point and the earliest added sorted index point; Based on the first sub-correlation degree, the second sub-correlation degree and the historical correlation degree, determine the target correlation degree between the index point and the sorted index points in the sliding window.
6. The method according to claim 3, characterized in that, The step of determining the target correlation degrees between each of the remaining index points and the sorted index points in the sliding window includes: For each index point, respectively execute the following steps: Determine the association scores of each of the sorted index points in the sliding window with an index point respectively; Based on the obtained association scores, determine the target association degree between the index point and the sorted index points in the sliding window.
7. The method according to claim 3, wherein, the sorting the respective index points according to each target association degree to obtain a target association sorting result includes: obtain the historical association sorting results corresponding to the respective index points in the previous iteration process; adjust the historical association sorting results based on the respective target association degrees to obtain a target association sorting result.
8. The method according to claim 3, wherein, the index points in the target association sorting result are arranged in descending order of the target association degree; the selecting a target index point from the respective index points as a sorted index point and adding it to the sliding window based on the target association sorting result includes: taking the first index point in the target association sorting result as the target index point and adding it to the sliding window.
9. The method according to any one of claims 1 to 8, wherein, after saving the multiple index points in the index memory according to the index point sorting result, further includes: obtain candidate index points for the search condition from the index memory, and at least one other index point continuously saved with the candidate index points; save the at least one other index point in the central processing unit cache; if the at least one other index point includes a neighbor index point of the candidate index point, obtain the neighbor index point of the candidate index point from the central processing unit cache; determine a search result for the search condition based on the candidate index point and the neighbor index point of the candidate index point.
10. An index point offline sorting device, wherein, includes: a selection module, configured to select an index point from multiple index points of a search graph as an initial sorted index point and add it to a sliding window; a sorting module, configured to determine the target association degree between each of the remaining index points and the sorted index points in the sliding window; select a target index point from the remaining index points based on each target association degree and add it to the sliding window; iteratively execute the step of determining the target association degree between each of the remaining index points and the sorted index points in the sliding window until the multiple index points are added to the sliding window; a storage module, configured to use the order of adding the multiple index points to the sliding window as an index point sorting result, and save the multiple index points in an index memory according to the index point sorting result.
11. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, when the processor executes the program, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, wherein, It stores a computer program executable by a computer device, and when the program runs on the computer device, it causes the computer device to execute the steps of any one of claims 1 to 9.
13. A computer program product, characterized in that the computer program product includes a computer program stored on a computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer device, it causes the computer device to execute the steps of the method according to any one of claims 1-9.
Citation Information
Patent Citations
Extensible partition method for associated flow graph data
CN104820705A
Webpage content extraction method and device
CN107741942A