String Retrieval Method, Device, Computer Equipment and Computer Readable Storage Medium

By using the layer neighbor graph index structure and incremental k-center point clustering algorithm in string retrieval, the problem of high computational complexity of character string k-nearest neighbor search based on edit distance in the prior art is solved, and efficient string retrieval is achieved.

CN119719434BActive Publication Date: 2025-05-30BERGMEIS (SHENZHEN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510213414.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In the prior art, the string k-nearest neighbor retrieval based on edit distances has high computational complexity, which limits its fast response and high accuracy requirements in practical applications.

Method used

The index structure of the layer neighbor graph is adopted, and the number of neighbors of each node is limited by building the layer directed graph and corresponding table, the number of nodes traversed during the search process is reduced, and the index structure is updated through the incremental k-central point clustering algorithm.

Benefits of technology

It significantly reduces the time complexity of string retrieval, making its time complexity close to the unrelated constant level, far lower than the linear level of the original search algorithm, meeting the needs of fast response and high accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719434B_ABST
    Figure CN119719434B_ABST
Patent Text Reader

Abstract

The embodiments of the present application disclose a string retrieval method, device and computer-readable storage medium. The method includes: obtaining a given reference string set to perform an index structure construction operation; according to the obtained search instruction, performing a k-nearest neighbor search operation on the index structure, and returning an approximate search result of the k-nearest neighbor. The index structure is composed of a neighbor graph and a correspondence table; the neighbor graph is composed of a directed graph, each layer of the directed graph is composed of nodes corresponding to the reference string, and some nodes in each layer are defined as center points, and the nodes of this layer are divided into clusters of limited size; the correspondence table stores each center point and its cluster and the nearest neighbor center point set information. The retrieval process implemented based on the constructed index structure requires much fewer edit distances to be calculated than the existing methods. Therefore, the present application provides a string k-nearest neighbor similar retrieval algorithm for edit distance, which has a significantly lower time complexity than the prior art and can achieve string retrieval more efficiently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of natural language processing, and particularly relates to a string retrieval method, device, computer device, and computer-readable storage medium. Background Art

[0002] Edit distance is a commonly used measure of the literal proximity between two strings, defined as the sum of the operation costs of converting one string to another by operations such as adding characters, deleting characters, and modifying characters. Among them, adding, deleting, and modifying characters are three common operations, each with its own preset unit cost. The string k-nearest neighbor retrieval problem based on edit distance is a typical problem in the field of natural language processing, generally defined as finding k reference strings with the smallest edit distances to a query string from a given set of reference strings under given unit costs of adding, deleting, and modifying characters, and listing these k reference strings in ascending order of edit distance. This problem has wide applications in scenarios such as text correction, auto-completion, and speech recognition.

[0003] Although calculating the string k-nearest neighbor based on edit distance has wide applications, this problem has a high computational complexity, which limits the practicality of the solutions to this problem. On the one hand, using the optimal dynamic programming method to calculate the edit distance between two strings with lengths m and n respectively, the time complexity reaches O(mn), that is, at the level of the square of the string length. On the other hand, to find the k nearest neighbors of a given query string among M reference strings, it is usually necessary to first calculate M edit distances and then select the smallest k from these M edit distances. Since the dynamic programming algorithm for edit distance is difficult to further optimize, to reduce the computational complexity of the string k-nearest neighbor based on edit distance, it is necessary to introduce some index structures to reduce the number of edit distance calculations. In the field of vector k-nearest neighbor retrieval, there is a class of efficient approximate k-nearest neighbor retrieval algorithms. The basic idea of these algorithms is to introduce a specific index structure to ensure that approximate k-nearest neighbors can be found on the premise of calculating the proximity between the query vector and only part rather than all reference vectors. However, the existing methods for calculating the string k-nearest neighbor based on edit distance (including the adapted methods for approximate k-nearest neighbor calculation in the field of vector retrieval) all have efficiency or effectiveness problems, and more efficient methods need to be designed to meet the requirements of fast response and high accuracy in practical applications. How to improve the retrieval efficiency of strings is a technical problem that needs to be urgently solved by those skilled in the art.

[0004] The foregoing description is for the purpose of providing general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] Based on this, it is necessary to propose a string retrieval method, device, computer device, and computer-readable storage medium for the above problems, which can more efficiently achieve the string retrieval efficiency.

[0006] The present application solves its technical problems by adopting the following technical solutions:

[0007] The present application provides a string retrieval method, including the following steps: when there is no index structure, obtain a given set of reference strings for index structure construction operations; the index structure consists of a neighbor graph and a correspondence table; the neighbor graph consists of layers of directed graphs, where is a preset integer; each layer of the directed graph consists of nodes, and one node corresponds to one reference string; the layer of the directed graph stores all the given reference strings; for any integer not less than 1 and not greater than , the set of reference strings stored in the layer of the directed graph is a subset of the set of reference strings stored in the layer of the directed graph; some nodes in each layer of the directed graph are marked as central points, and the central points are determined according to the index structure construction operations. The central points divide all the nodes in the directed graph into multiple non-overlapping subsets, and each subset is called a cluster. The number of nodes in the cluster does not exceed , where is a preset integer; in each layer of the directed graph, there are directed edges between some nodes, and the directed edges are determined according to the index structure construction operations; the node pointed to by the directed edge is called the neighbor of the node that initiates the directed edge, and the number of neighbors of each node does not exceed ; each layer of the neighbor graph has a unique correspondence table, and the correspondence table is a two-dimensional table composed of rows and columns; each row corresponds to a central point in the corresponding layer of the neighbor graph; there are three columns in total: the first column stores the reference string corresponding to the central point; the second column stores the reference strings corresponding to the other nodes in the cluster where the central point is located except the central point; the third column stores the set of reference strings corresponding to the set of the nearest neighbor central points of the central point, and the set of the nearest neighbor central points of the central point is determined according to the index structure construction operations; when the set of reference strings corresponding to the index structure changes, generate add, delete, and modify instructions according to the change situation of the set of reference strings, and perform add, delete, and modify operations on the index structure according to the add, delete, and modify instructions; according to the obtained retrieval instruction, perform -nearest neighbor retrieval operation on the index structure, and return -the retrieval result output by the nearest neighbor retrieval operation; the content of the retrieval instruction includes the query string, the expected result set size and the retrieval distance, and the retrieval distance is the edit distance from a preset string to another string; the retrieval operation includes: marking the query string as a test node; randomly selecting in the first layer of the directed graph A node without a deletion mark, marked as the entry node of the first layer. The deletion mark of the node is determined according to the index structure. If the number of nodes without a deletion mark in the first-layer directed graph is less than , then all nodes without a deletion mark are marked as the entry nodes of the first-layer directed graph; start processing layer by layer from the first layer of the neighbor graph until the layer of the neighbor graph is processed. The processing operations include: after determining the set of entry nodes of the -layer directed graph, first mark the set of entry nodes as the node set at the 0th moment ; then according to the gradually increasing moment , successively determine the node set at the th moment from the node set at the th moment . Mark the union of the neighbor sets of all nodes in as the first union set, and mark the union of the first union set and as the second union set. Then determine as the set composed of the nodes without a deletion mark with the smallest retrieval distance to the test node in the second union set; if the number of nodes without a deletion mark in the second union set is less than , then determine as the set composed of all nodes without a deletion mark in the second union set; when = , mark as ; if , then mark as the set of entry nodes of the -layer directed graph and continue the processing of the layer; if , then mark as the retrieval result and output.

[0008] In an alternative embodiment of the present application, the index structure construction operation includes: determining the probability that the minimum layer for each reference string to be added to the neighbor graph is the l th layer :

[0009] , where is a preset real number greater than 1;

[0010] According to , randomly determine the minimum layer l for each given reference string to be added to the neighbor graph, and successively add each given reference string as a node to the neighbor graph from the th layer to the In the directed graph of each layer; when adding the corresponding node of a certain reference string to the directed graph of each layer, mark the node to be added as an inserted node; regard the first inserted node as the first center point and form a cluster by itself; in the subsequent insertion process, calculate the specified distance value between the newly inserted node and the existing center points one by one. The specified distance value is a preset literal distance metric value between two nodes that satisfies non-negativity, symmetry, and the triangle inequality; mark the cluster where the center point with the smallest specified distance value is located as the insertion cluster, and add the inserted node to the insertion cluster; when the number of nodes in the insertion cluster exceeds , split the insertion cluster into two new clusters. The center points of the two new clusters are two different nodes in the insertion cluster, satisfying that the sum of the deviation values of the other nodes in the insertion cluster to the two new clusters is the smallest. The deviation value of each node to the two new clusters is the minimum of the specified distance values between the node and the center points of the two new clusters. Then, call the other nodes in the extended cluster except the center points of the two new clusters as the nodes to be processed, calculate the specified distance values between the nodes to be processed and the center points of the two new clusters, and add the nodes to be processed to the new cluster with the smaller specified distance value; if the specified distance values between the nodes to be processed and the center points of the two new clusters are the same, add the nodes to be processed to the new cluster with fewer nodes; after all the given reference strings are added as nodes to the neighbor graph, determine the clusters included in the directed graph of each layer of the neighbor graph; calculate the node with the smallest sum of the specified distance values to other nodes in each cluster and mark it as the center point of each cluster; for each integer that satisfies , determine the neighbor relationship based on all the center points and other nodes in the th layer of the neighbor graph to construct the corresponding table for the th layer; summarize the neighbor graph and its corresponding tables for each layer to obtain the index structure.

[0011] In an optional embodiment of the present application, determine the neighbor relationship based on all the center points and other nodes in the th layer of the neighbor graph to construct the corresponding table for the th layer, including: calculating the specified distance values between the center points in the th layer; according to the specified distance values between the center points, determine other center points with the smallest specified distance values for each center point, and mark the set composed of other center points as the set of the nearest neighbor center points of the center point; for each node in the l th layer, obtain the center point of the cluster where the node is located, and all the center points in the set of the nearest neighbor center points of the center point of the cluster where the node is located, a total of center points; mark the set composed of center points as the set of the nearest neighbor center points of the node; obtain the node with the smallest specified distance value to the node in the cluster where all the center points in the set of the nearest neighbor center points of each node are located a number of other nodes, and mark the set composed of a number of other nodes as the neighbor set of the node; Add a directed edge from the node to a number of other nodes in the layer directed graph; According to the reference string corresponding to each center point in the layer, the reference strings corresponding to other nodes in the cluster where the center point is located, and the set of nearest neighbor center points of the center point, determine the correspondence table of the layer.

[0012] In an alternative embodiment of the present application, the add / delete / modify instruction includes an operation type and an operation object, and the operation type is one of add, delete, and modify; When the operation type is add, the operation object is the reference string to be added; When the operation type is delete, the operation object is the reference string to be deleted; When the operation type is modify, the operation object is the reference string before change and the reference string after change.

[0013] In an alternative embodiment of the present application, performing add / delete / modify operations on the index structure according to the add / delete / modify instruction includes: When the operation type of the add / delete / modify instruction is add, obtain the reference string to be added corresponding to the operation object, and mark the reference string to be added as the insert node; Determine the probability that the minimum layer for the insert node to join the neighbor graph is the layer :

[0014] , is a preset real number greater than 1;

[0015] According to randomly determine the minimum layer for the insert node to join , and add the insert node to the directed graph from the layer to the layer of the neighbor graph; When adding the insert node to each layer of the directed graph, calculate the specified distance value between the insert node and all center points in the directed graph, mark the cluster where the center point with the smallest specified distance value is located as the extended cluster, and add the insert node to the extended cluster;

[0016] When the number of nodes in the extended cluster exceeds When splitting the extended cluster into two new clusters, the central points of the two new clusters are two different nodes in the extended cluster, such that the sum of the deviation values of the other nodes in the extended cluster to the two new clusters is minimized. The deviation value of each node to the two new clusters is the minimum of the specified distance values between the node and the central points of the two new clusters. The other nodes in the extended cluster except the central points of the two new clusters are called nodes to be processed, and the specified distance values between the nodes to be processed and the central points of the two new clusters are calculated. The nodes to be processed are added to the new cluster with the smaller specified distance value. If the specified distance values of the nodes to be processed to the central points of the two new clusters are the same, the nodes to be processed are added to the new cluster with fewer nodes. The specified distance values between the central points of the two new clusters and other central points are calculated, and each new cluster's central point is respectively determined to be the one with the smallest specified distance value among the following other central points, and the set composed of the following other central points is marked as the nearest neighbor central point set of the central point of the new cluster. For each node in the two new clusters, the central point of the corresponding new cluster and all the central points in the nearest neighbor central point set of the central point of the new cluster are obtained, a total of the following central points; and the set composed of the following central points is marked as the nearest neighbor central point set of each node in the new cluster. The following other nodes with the smallest specified distance value to the node within the cluster where all the central points in the nearest neighbor central point set of each node are located are obtained, and the set composed of the following other nodes is marked as the neighbor set of the node. The row where the central point of the extended cluster is located is deleted from the corresponding table, and according to the node sets, central points of the two new clusters, and the nearest neighbor central point set of the central point of the new cluster, the corresponding row information of the central points of the two new clusters is added to the corresponding table. Each node in the neighbor set of the inserted node that does not belong to any of the new clusters is marked as a neighbor to be processed. First, the inserted node is added to the neighbor set of the neighbor to be processed, and then the neighbor with the largest specified distance value is deleted from the neighbor set of the neighbor to be processed.

[0017] When the number of nodes in the extended cluster after adding the inserted node does not exceed the following, the central point of the cluster where the inserted node is located and all the central points in the nearest neighbor central point set of the central point of the cluster are obtained, a total of the following central points; and the set composed of the following central points is marked as the nearest neighbor central point set of the inserted node. The following other nodes with the smallest specified distance value to the inserted node within the cluster where all the central points in the nearest neighbor central point set of the inserted node are located are obtained; and The set composed of other nodes is marked as the neighbor set of the inserted node; update the information in the second column of the row where the center point of the expanded cluster is located in the corresponding table according to the node set of the expanded cluster; mark each node in the neighbor set of the inserted node as a pending neighbor; first add the inserted node to the neighbor set of the pending neighbor, and then delete the neighbor with the largest specified distance value from the neighbor set of the pending neighbor.

[0018] In an alternative embodiment of the present application, performing addition, deletion, and modification operations on the index structure according to the addition, deletion, and modification instructions includes: when the operation type of the addition, deletion, and modification instruction is deletion, obtaining the reference string to be deleted corresponding to the operation object, and marking the reference string to be deleted as a deletion node; processing the directed graph and the corresponding table containing the deletion node layer by layer for the neighbor graph; determining the cluster where the deletion node is located according to the corresponding table, marking the cluster where the deletion node is located as a shrinking cluster, and marking the deletion node with a deletion mark;

[0019] When all nodes in the shrinking cluster have deletion marks, delete all nodes in the shrinking cluster from the directed graph, and delete the row where the center point of the shrinking cluster is located from the corresponding table; for each center point in the directed graph, if the set of the nearest neighbor center points of the center point contains the center point of the shrinking cluster, then re-determine the other center points with the smallest specified distance value from the center point in the directed graph, and update the set of the nearest neighbor center points of the center point to the set composed of these other center points; for each node in the directed graph, if the neighbor set of the node contains at least one node in the shrinking cluster, then re-obtain the center point of the cluster where the node is located, and all center points in the set of the nearest neighbor center points of the center point of the cluster where the node is located, a total of center points; update the set composed of center points to the set of the nearest neighbor center points of the node; re-obtain the other nodes with the smallest specified distance value from the node within the cluster where all center points in the set of the nearest neighbor center points of the node are located, and update the set composed of these other nodes to the neighbor set of the node;

[0020] When there are still nodes without deletion marks in the shrinking cluster, no further processing is performed.

[0021] In an alternative embodiment of the present application, performing addition, deletion, and modification operations on the index structure according to the addition, deletion, and modification instructions includes: when the operation type of the addition, deletion, and modification instruction is modification, obtaining the reference string before the change and the reference string after the change corresponding to the operation object; first marking the reference string before the change as a deletion node to perform the addition, deletion, and modification operation with the operation type of deletion, and then marking the reference string after the change as an inserted node to perform the addition, deletion, and modification operation with the operation type of addition.

[0022] The present application also provides a string retrieval device, including: an index construction module, configured to obtain a given set of reference strings for index structure construction operations when there is no index structure; an index maintenance module, configured to generate addition, deletion, and modification instructions according to the change situation of the set of reference strings corresponding to the index structure when the set of reference strings changes, and perform addition, deletion, and modification operations on the index structure according to the addition, deletion, and modification instructions; a retrieval module, configured to perform - a nearest neighbor retrieval operation and return the retrieval result output by the retrieval operation.

[0023] The present application also provides a computer device, including a processor and a memory: the processor is configured to execute a computer program stored in the memory to implement the method as described above.

[0024] The present application also provides a computer-readable storage medium storing a computer program, which implements the method as described above when executed by a processor.

[0025] Adopting the embodiments of the present application has the following beneficial effects:

[0026] The present application is based on The key of the index structure of the layer neighbor graph lies in restricting the number of neighbors of each node, so that the retrieval process can be performed in the finite neighborhood of the entry node set rather than the entire node set, thereby controlling the number of nodes traversed in the retrieval process. In order to limit the number of neighbors of the inserted node to a preset threshold within, directly adapting the index construction method of HNSW requires finding the node with the smallest specified distance value from all the added nodes to the inserted node. Therefore, directly adapting the index construction method of HNSW requires calculating times of the specified distance value, where is the number of reference strings. In the string retrieval method based on this index structure, the most costly operation is the operation of calculating the edit distance between two strings. Therefore, the time complexity of the string retrieval method can be measured by the number of calculations of the edit distance. Since each layer in the layer neighbor graph needs to calculate times of the edit distance, where is the number of layers of the neighbor graph, is the number of nodes included in the retrieval result, is the maximum number of nodes in a cluster, is the average number of iterations at the convergence of each layer. Therefore, the above string retrieval method needs to calculate times of the edit distance in total. Relatively speaking, the original string retrieval method needs to calculate times of the edit distance. Since generally Very small (usually between 1 and 9), and is much smaller than a constant, so the string retrieval method of this application can be regarded as an approximate retrieval algorithm at the constant level independent of , and its time complexity is significantly lower than the linear level of the original retrieval algorithm.

[0027] The above description is only an overview of the technical solution of this application. In order to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of this application more obvious and understandable, the following specifically gives preferred embodiments and, in conjunction with the drawings, details are described in detail. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings

[0028] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0029] Among them:

[0030] Figure 1 is a schematic flowchart of a string retrieval method provided by an embodiment;

[0031] Figure 2 is an internal functional module relationship diagram of a string retrieval device provided by an embodiment;

[0032] Figure 3 is a schematic block diagram of the structure of a computer device provided by an embodiment. Detailed Embodiments

[0033] The following will clearly and completely describe the technical solutions in the embodiments of this application in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of this application.

[0034] Edit distance is a commonly used measure of the literal proximity between two strings, defined as the sum of the operation costs of converting one string to the other by operations such as adding characters, deleting characters, and modifying characters. Among them, adding, deleting, and modifying characters are three common operations, each with its own preset unit cost. When the unit costs of adding characters and deleting characters are different, the edit distance is asymmetric, that is, the edit distance from string A to string B is not necessarily equal to the edit distance from B to A. The problem of string k-nearest neighbor retrieval based on edit distance is a typical problem in the field of natural language processing, generally defined as finding k reference strings with the smallest edit distances to the query string from a given set of reference strings under given unit costs of adding, deleting, and modifying characters and a given query string, and listing these k reference strings in ascending order of edit distance. This problem has a wide range of applications in scenarios such as text correction, auto-completion, and speech recognition. Among them, in the text correction scenario, calculating the string k-nearest neighbor based on edit distance can be used to detect spelling mistakes and correct grammar mistakes; in the auto-completion scenario, calculating the string k-nearest neighbor based on edit distance can be used to auto-complete the words input by the user, that is, to perform auto-completion by finding the word corresponding to the smallest edit distance; in the speech recognition scenario, calculating the string k-nearest neighbor based on edit distance can identify the words in the speech and find the most likely text match, thereby improving the accuracy of speech recognition.

[0035] Although calculating the string k-nearest neighbor based on edit distance has a wide range of applications, this problem has a high computational complexity, which limits the practicality of the solutions to this problem. On the one hand, using the optimal dynamic programming method to calculate the edit distance between two strings with lengths m and n respectively, the time complexity reaches O(mn), that is, at the square level of the string length. On the other hand, to find the k nearest neighbors of a given query string among M reference strings, usually M edit distances need to be calculated first, and then the smallest k of these M edit distances are selected. Since the dynamic programming algorithm for edit distance is difficult to further optimize, to reduce the computational complexity of the string k-nearest neighbor based on edit distance, some indexing structures need to be introduced to reduce the number of edit distance calculations.

[0036] In the field of k-nearest neighbor retrieval of vectors, there is a class of efficient approximate k-nearest neighbor retrieval algorithms. The basic idea of this class of algorithms is to introduce a specific index structure to ensure that the approximate k-nearest neighbors can be found on the premise of calculating the proximity between the query vector and some rather than all reference vectors. For example, the popular Hierarchical Navigable Small World (HNSW) algorithm constructs a multi-layer neighbor graph as the index structure, where the nodes of the neighbor graph are reference vectors, and the number of neighbors of each layer of nodes is limited within a preset threshold; when calculating the approximate k-nearest neighbors of a given query vector, the HNSW algorithm starts from the sparsest top-layer neighbor graph and refines the approximate k-nearest neighbors of the query vector layer by layer downward, that is, taking the approximate k-nearest neighbors output by the adjacent upper layer as the entry point to find more accurate approximate k-nearest neighbors in the current layer until the bottom-layer neighbor graph containing all reference vectors is processed. This algorithm only needs to traverse a small part of the vector nodes in each layer of the multi-layer neighbor graph and calculate the proximity between them and the query vector, so it does not need to calculate the proximity between all reference vectors and the query vector.

[0037] There are two relatively intuitive methods to adapt the approximate k-nearest neighbor retrieval algorithm in the vector retrieval field to the calculation of string k-nearest neighbor based on the edit distance.

[0038] One method is to first vectorize the strings, then construct a vector-based index structure, and finally call a similar approximate k-nearest neighbor retrieval algorithm to find the approximate k-nearest neighbor based on the edit distance. For example, algorithms like HNSW require that two connected vector nodes in the neighbor graph be as close as possible in the specified vector proximity, but two vector nodes that are as close as possible in the vector proximity are not necessarily close in the edit distance, because the edit distance considers literal symbols and its value depends on the preset unit costs of three operations: insertion, deletion, and modification, while the vector proximity considers the semantics of words. This results in that in the multi-layer neighbor graph constructed by vector nodes, it is not necessarily possible to gradually find more accurate approximate nearest neighbors based on the edit distance through the neighbor relationship.

[0039] Another method does not introduce the vector representation of strings. Instead, during the index construction process and the string retrieval process, it directly operates at the symbol level of the strings and replaces the vector proximity with the edit distance. Since the edit distance is asymmetric, when replacing the vector proximity with the edit distance, the asymmetry of the proximity between nodes must be considered, which will inevitably increase the complexity of graph-based index construction and string retrieval. In addition, this method also exacerbates the high-cost problem of index construction. Taking the HNSW algorithm as an example, this algorithm constructs an index structure by adding vector nodes one by one to a multi-layer neighbor graph. When adding a single vector node, it is necessary to calculate the proximity between it and the existing vector nodes in the graph. Therefore, a total of M(M - 1) / 2 vector proximities need to be calculated. Since the computational complexity of vector proximity is linear with respect to the vector dimension, while the computational complexity of edit distance is quadratic with respect to the string length, directly using the multi-layer neighbor graph construction algorithm in HNSW to construct an index structure based on edit distance will result in a higher time cost.

[0040] It can be seen that the existing methods for calculating the string k-nearest neighbors based on edit distance (including the adaptation methods for approximate k-nearest neighbor calculation in the field of vector retrieval) all have efficiency or effectiveness problems, with high computational time complexity and low efficiency. For this reason, this application proposes a string retrieval method to meet the requirements of fast response and high accuracy in string retrieval in practical applications. To clearly describe the method provided in this embodiment, please refer to Figure 1 , which includes steps S110 to S130.

[0041] Step S110: When there is no index structure, obtain a given set of reference strings for index structure construction operations.

[0042] In one embodiment, the retrieval method provided in this application reduces the computational time complexity during the retrieval process by introducing an index structure. This index structure consists of a neighbor graph and a corresponding table. Before retrieval, if it is detected that this index structure does not exist, it is necessary to obtain a given set of reference strings for index structure construction operations to complete the construction of this index structure. The specific composition and construction process of the index structure will be described in detail later.

[0043] The neighbor graph in the index structure consists of layers of directed graphs, where is a preset integer. Among them, each layer of the directed graph consists of nodes, and one node corresponds to one reference string. The th layer of the directed graph stores all the given reference strings; for any integer not less than 1 and not greater than , the set of reference strings stored in the th layer of the directed graph is the th layer of the directed graph is the A subset of the set of reference strings stored in the layer directed graph. Specifically, the set of reference strings stored in the layer directed graph and the set of reference strings stored in the layer directed graph, the ratio of these two sets is random and depends on the minimum layer probability of addition. The minimum layer probability of addition will be calculated and determined for each reference string during the index structure construction operation, which will be described in detail later. Some nodes in each layer directed graph are marked as central points, and the central points are determined according to the index structure construction operation. The central points divide all the nodes in the directed graph into multiple non-overlapping subsets, and each subset is called a cluster. The number of nodes in the cluster does not exceed , which is a preset integer. In each layer directed graph, there are directed edges between some nodes, and the directed edges are determined according to the index structure construction operation; the node pointed to by the directed edge is called the neighbor of the node that initiates the directed edge. The number of neighbors of each node does not exceed , that is, the preset integer described above, which will be mentioned many times later and will not be separately explained. The two nodes with a directed edge indicate that the reference strings corresponding to the two nodes are literally close.

[0044] Each layer of the neighbor graph has a unique corresponding table, and the corresponding table is a two-dimensional table composed of rows and columns. Each row corresponds to a central point in the corresponding layer of the neighbor graph; there are three columns in total: the first column stores the reference string corresponding to the central point; the second column stores the reference strings corresponding to the other nodes in the cluster where the central point is located except the central point; the third column stores the set of reference strings corresponding to the set of the nearest neighbor central points of the central point, and the set of the nearest neighbor central points of the central point is determined through the index structure construction operation. The set of the nearest neighbor central points is used to represent several central points that are closest to a certain node / central point. There are central points in the set of the nearest neighbor central points of each node; there are central points in the set of the nearest neighbor central points of each central point.

[0045] In one embodiment, the index structure construction operation operates on the given reference string. As described above, the index structure is composed of a neighbor graph and a corresponding table, and the corresponding construction operation is to construct the neighbor graph and the corresponding table respectively.

[0046] The method provided by this application is implemented through the index structure, and the index structure is necessary. Therefore, if it is determined before retrieval that the required index structure of this application does not exist, the corresponding index structure needs to be constructed. Assume that the set of reference strings given for constructing the index structure includes reference strings. When constructing the neighbor graph, each reference string needs to be processed one by one to construct the neighbor graph layer by layer. First, determine that the minimum layer for each reference string to be added to the neighbor graph is thel Probability of layer , is calculated as follows:

[0047] (1)

[0048] In the above formula, is a preset real number greater than 1.

[0049] According to randomly determine the minimum layer for each given reference string to be added to the neighbor graph l , and one by one add the given reference string as a node to the directed graph from the rd layer to the th layer of the neighbor graph, so that the th layer directed graph stores all the given reference strings, and the set of reference strings stored in the th layer directed graph is a subset of the set of reference strings stored in the th layer directed graph. Based on the definition of , the total probability of the reference string being added to the th layer is approximately:

[0050] (2)

[0051] The total probability of being added to the l + 1 layer is approximately:

[0052] (3)

[0053] Thus, the ratio of the number of nodes in the th layer to the number of nodes in the th layer is approximately equal to:

[0054] (4)

[0055] This application uses an incremental k-medoids clustering algorithm to partition the node set of each layer into disjoint clusters. That is, when adding the corresponding node of a certain reference string to the directed graph of each layer, the node to be added is marked as an insertion node. The first insertion node is regarded as the first center point and forms a cluster by itself. In the subsequent insertion process, the specified distance values between the newly inserted nodes and the existing center points are calculated one by one. The cluster where the center point with the smallest specified distance value is located is marked as the insertion cluster, and the insertion node is added to the insertion cluster. The specified distance value is a preset literal distance metric value between two nodes that satisfies non-negativity, symmetry, and the triangle inequality. Among them, the specified distance value between two strings can be the difference between the maximum lengths of the two strings and the length of their longest common substring, or the edit distance with all operation unit costs set to 1, etc. Requiring the specified distance value to satisfy symmetry is to reduce the computational amount of the specified distance value by half, while requiring the specified distance value to satisfy non-negativity and the triangle inequality is to ensure that the specified distance value from each node to the neighbor of its neighbor is as close as possible to the specified distance value from the node to its neighbor, so as to ensure that there is also a high probability of finding the nearest neighbor in the neighborhood of the neighbor.

[0056] When the number of nodes in the insertion cluster exceeds the insertion cluster is split into two new clusters. The center points of the two new clusters are two different nodes in the insertion cluster, satisfying that the sum of the deviation values from the other nodes in the insertion cluster to the two new clusters is the smallest. The deviation value of each node to the two new clusters is the minimum of the specified distance values between the node and the center points of the two new clusters. Specifically, the original insertion cluster is marked as the nodes selected as the center points of the two new clusters are marked as and the nodes in the original insertion cluster except and are marked as . and need to minimize the following formula:

[0057] (5)

[0058] In the above formula, is the specified distance value calculation function. It should be noted that the specified distance value introduced in the process of index construction and subsequent update in the technical solution of this application does not limit the specific calculation formula. As long as the specified distance value is defined at the string symbol level and satisfies non-negativity, symmetry, and the triangle inequality, no matter what specific calculation formula is used, it should not be considered to exceed the technical scope protected by this application.

[0059] Nodes other than the central points of the two new clusters in the extended cluster are called nodes to be processed, and the specified distance values between the nodes to be processed and the central points of the two new clusters are calculated. The nodes to be processed are added to the new cluster with the smaller specified distance value. That is to say, for the inserted cluster each non or node , if , then is added to 's containing cluster; if , then is added to 's containing cluster.

[0060] If the specified distance values between the node to be processed and the central points of the two new clusters are the same, the node to be processed is added to the new cluster with fewer nodes. That is to say, if , then is added to the cluster with fewer elements among 's containing cluster and 's containing cluster.

[0061] After all the given reference strings are added as nodes to the neighbor graph, the clusters included in each layer of the directed graph of the neighbor graph are determined. The node with the smallest sum of the specified distance values to other nodes in each cluster is recalculated and marked as the central point of each cluster. That is to say, the new central point in the cluster where the original central point is located needs to minimize the following formula:

[0062] (6)

[0063] For each integer satisfying , neighbor relationships are determined based on all the central points and other nodes in the th layer of the neighbor graph to construct the corresponding table for the th layer.

[0064] In one embodiment, to determine neighbor relationships based on all the central points and other nodes in the th layer of the neighbor graph to construct the corresponding table for the th layer, the set of nearest neighbor central points of each node and central point needs to be determined.

[0065] The processing steps first determine the set of nearest neighbor central points of each central point, specifically by calculating the specified distance values between the central points in the th layer. Based on the specified distance values between the central points, for each central point, other central points with the smallest specified distance values are determined, and these The set composed of other center points is marked as the nearest neighbor center point set of the center point.

[0066] Then, determine the nearest neighbor center point set of each node. For each node in the l -th layer, obtain the center point of the cluster where the node is located, and all the center points in the nearest neighbor center point set of the center point of the cluster where the node is located, a total of center points. Mark the set composed of these center points as the nearest neighbor center point set of the node.

[0067] Next, determine the neighbor set of each node. Obtain the other nodes with the smallest specified distance value within the clusters where all the center points in the nearest neighbor center point set of each node are located, and mark the set composed of these other nodes as the neighbor set of the node. According to the neighbor set, the directed edges between nodes can be determined, that is, add directed edges from the node to other nodes in the directed graph of the -th layer.

[0068] Based on the nearest neighbor center point set and neighbor set determined above, taking the -th layer in the neighbor graph as an example, determine the correspondence table of the -th layer according to the reference string corresponding to each center point in the -th layer, the reference strings corresponding to other nodes in the cluster where the center point is located, and the nearest neighbor center point set of the center point. Thus, determine the correspondence table of the neighbor graph for each layer.

[0069] Finally, summarize the neighbor graph and its correspondence table for each layer to obtain the index structure.

[0070] In the above index construction method, the most costly operation is the operation of calculating the specified distance value between two strings. Therefore, the time complexity of the index construction method can be measured by the number of calculations of the specified distance value. When calculating the of each reference string, there is no need to calculate the specified distance value; while starting from adding the first reference string, it is necessary to calculate the specified distance value. When adding a certain reference string, that is, inserting a node, calculations of the specified distance value between the inserted node and the existing center points are required for positioning the insertion cluster, and calculations of the specified distance value are required for the splitting of the insertion cluster. Therefore, calculations of the specified distance value are required to add all the reference strings to the corresponding clusters. Then, calculations of the specified distance value are required to adjust the center points of all the clusters, and a total of Calculate the specified distance value for the second time. Finally, within the cluster where the central point in the set of nearest neighbor central points of each node is located, to determine the total number of times of calculating the specified distance value. Therefore, the above index construction method requires L in each layer of the layer neighbor graph, a total of times of calculating the specified distance value, so the overall method requires much less than , so the above time complexity is significantly less than that of the index construction method directly adapting HNSW in terms of times of calculating the specified distance value.

[0071] Step S120: When the reference string set corresponding to the index structure changes, generate add, delete, and modify instructions according to the change situation of the reference string set; perform add, delete, and modify operations on the index structure according to the add, delete, and modify instructions.

[0072] In one embodiment, it can be understood that the index structure is used to record reference strings. When the targeted reference string changes, the index structure needs to change accordingly. Therefore, it is necessary to generate add, delete, and modify instructions according to the change situation of the reference string to perform add, delete, and modify operations on the index structure to continuously maintain the index structure. Among them, the add, delete, and modify instructions include an operation type and an operation object, and the operation type is one of add, delete, and modify. For different operation types, the corresponding operation objects are also different: when the operation type is add, the operation object is the reference string to be added; when the operation type is delete, the operation object is the reference string to be deleted; when the operation type is modify, the operation objects are the reference string before change and the reference string after change. How to perform add, delete, and modify operations on the index structure will be described one by one later.

[0073] In one embodiment, performing add, delete, and modify operations on the index structure according to the add, delete, and modify instructions includes: when the operation type of the add, delete, and modify instruction is add, obtain the reference string to be added corresponding to the operation object, and mark the reference string to be added as an insertion node. The subsequent processing process is similar to the index structure construction operation. Generally speaking, it is to determine the minimum layer where the insertion node joins, and add the insertion node layer by layer in the directed graph from the minimum layer to layer, and update the corresponding table according to the situation after insertion.

[0074] First, determine the probability that the minimum layer where the insertion node joins the neighbor graph is the layer , and the calculation method refers to Equation (1). According to randomly determine the minimum layer where the insertion node joins , and add the insertion node to the neighbor graph from the layer to the in the directed graph of the layer.

[0075] Then determine the cluster where the insertion node is specifically inserted, and update the subsequent set of nearest neighbor central points and neighbor set. Specifically, when adding an insertion node in the directed graph of each layer, calculate the specified distance value between the insertion node and all central points in the directed graph, mark the cluster where the central point with the smallest specified distance value is located as the extended cluster, and add the insertion node to the extended cluster. The calculation method of the specified distance value has been described above and will not be elaborated further below.

[0076] During the insertion process, since the number of nodes in a cluster cannot exceed , so it is necessary to divide the insertion process into two methods according to whether it exceeds .

[0077] When the number of nodes in the extended cluster after adding the insertion node exceeds , split the extended cluster into two new clusters. Determine the central points of the two new clusters, specifically two different nodes in the extended cluster, so that the sum of the deviation values of other nodes in the extended cluster to the two new clusters is the smallest. The deviation value of each node to the two new clusters is the minimum of the specified distance values between the node and the central points of the two new clusters.

[0078] Call the other nodes in the extended cluster except the central points of the two new clusters as the nodes to be processed, and calculate the specified distance values between the nodes to be processed and the central points of the two new clusters. Add the nodes to be processed to the new cluster with the smaller specified distance value; if the specified distance values between the nodes to be processed and the central points of the two new clusters are the same, add the nodes to be processed to the new cluster with fewer nodes.

[0079] After that, update the set of nearest neighbor central points of the central points of the new clusters, and the set of nearest neighbor central points of each node in the new clusters. Calculate the specified distance values between the central points of the two new clusters and other central points, and respectively determine the other central points with the smallest specified distance value for the central points of each new cluster according to the specified distance values, and mark the set composed of these other central points as the set of nearest neighbor central points of the central points of the new clusters. For each node in the two new clusters, obtain the central point of the new cluster corresponding to each node, and all the central points in the set of nearest neighbor central points of the central points of the new clusters, a total of central points. Mark the set composed of these central points as the set of nearest neighbor central points of each node in the new clusters.

[0080] Then update the neighbor set of the nodes in the new clusters, and keep the number of neighbor sets of each node as . Obtain the other nodes with the smallest specified distance value from the node in the cluster where all the central points in the set of nearest neighbor central points of each node are located, and The set composed of other nodes is marked as the neighbor set of the node. In addition, each node in the neighbor set of the inserted node that does not belong to any new cluster is marked as a neighbor to be processed. First, the inserted node is added to the neighbor set of the neighbor to be processed, and then the neighbor with the largest specified distance value is deleted from the neighbor set of the neighbor to be processed.

[0081] According to the above processing, the neighbor graph is updated, and then the corresponding table is updated. Delete the row where the center point of the extended cluster is located from the corresponding table, and add the corresponding row information of the center points of the two new clusters to the corresponding table according to the node sets, center points of the two new clusters, and the set of the nearest neighbor center points of the center points of the new clusters.

[0082] When the number of nodes in the extended cluster after adding the inserted node does not exceed , since there is no need to perform clustering, the set of nearest neighbor center points, neighbor set, and corresponding table are directly updated.

[0083] Determine the set of nearest neighbor center points of the inserted node: Obtain the center point of the cluster where the inserted node is located and all the center points in the set of the nearest neighbor center points of the center point of the cluster, a total of center points. Mark the set composed of these center points as the set of the nearest neighbor center points of the inserted node.

[0084] Determine the neighbor set of the inserted node: Obtain the other nodes with the smallest specified distance value from the inserted node within the clusters where all the center points in the set of the nearest neighbor center points of the inserted node are located. Mark the set composed of these other nodes as the neighbor set of the inserted node. In addition, mark each node in the neighbor set of the inserted node as a neighbor to be processed. First, the inserted node is added to the neighbor set of the neighbor to be processed, and then the neighbor with the largest specified distance value is deleted from the neighbor set of the neighbor to be processed to maintain the number of elements in the neighbor set of this neighbor as .

[0085] Update the corresponding table: Update the information in the second column of the row where the center point of the extended cluster is located in the corresponding table according to the node set of the extended cluster.

[0086] The above incremental method of adding a single reference string requires times of calculating the specified distance value between the inserted cluster and the existing center points when positioning the inserted cluster in each layer of the -layer neighbor graph. The splitting of the inserted cluster requires times of calculating the specified distance value, and determining the set of the nearest neighbor center points of the center points of the two new clusters after splitting requires times of calculating the specified distance value. Therefore, adding the inserted node to the corresponding cluster requires at most Calculate the specified distance value for the first time. Then, obtaining the neighbor set of the inserted node requires calculations of the specified distance value. Adjusting the neighbor sets of each neighbor of the inserted node requires a total of calculations of the specified distance value. Therefore, processing the inserted node in each layer of the layer neighbor graph requires a total of calculations of the specified distance value. Thus, adding a single reference string requires a total of calculations of the specified distance value.

[0087] In one embodiment, when the operation type of the add / delete / modify instruction is delete, obtain the reference string to be deleted corresponding to the operation object. Perform add / delete / modify operations on the index structure according to the add / delete / modify instruction, which specifically includes:

[0088] Mark the reference string to be deleted as a delete node, and layer by layer process the directed graph and the corresponding table containing the delete node in the neighbor graph. Similarly, it specifically includes the update of the nearest neighbor center point set, the neighbor set, and the corresponding table.

[0089] When processing a certain layer, determine the cluster where the delete node is located according to the corresponding table, mark the cluster where the delete node is located as a shrinking cluster, and mark the delete node with a delete flag.

[0090] When all nodes in the shrinking cluster have delete flags, perform the following processing steps.

[0091] Update the corresponding table: Delete all nodes in the shrinking cluster from the directed graph, and delete the row where the center point of the shrinking cluster is located from the corresponding table.

[0092] Update the nearest neighbor center point set of the center point: For each center point in the directed graph, if the nearest neighbor center point set of the center point contains the center point of the shrinking cluster, then re-determine the other center points in the directed graph with the smallest specified distance value from the center point, and update the nearest neighbor center point set of the center point to be the set composed of these other center points.

[0093] Update the nearest neighbor center point set of the node: For each node in the directed graph, if the neighbor set of the node contains at least one node in the shrinking cluster, then re-obtain the center point of the cluster where the node is located, and all the center points in the nearest neighbor center point set of the center point of the cluster where the node is located, a total of center points. Update the set composed of these center points to be the nearest neighbor center point set of the node.

[0094] Update the neighbor set: For each node in the directed graph, if the neighbor set of the node contains at least one node in the shrinking cluster, then re-obtain the node with the smallest specified distance value within the cluster where all the center points in the nearest neighbor center point set of the node are located, a number of other nodes, and update the set composed of a number of other nodes to the neighbor set of the node, so as to maintain the number of elements in the neighbor set of the neighbor as .

[0095] When there are still nodes without deletion marks in the contracted cluster, no further processing is performed.

[0096] The above incremental method for deleting a single reference string can directly locate the contracted cluster in each layer of the layer neighbor graph according to the correspondence table between the center point and its cluster, and mark the deleted node with a deletion mark without directly deleting the deleted node from the multi-layer neighbor graph, so that the index structure can be ensured to remain unchanged and no additional calculation of the specified distance value is required. Only when the contracted cluster exists and all nodes in the contracted cluster have deletion flags, the deletion operation of the contracted cluster is performed. The number of different center points containing the center point of the contracted cluster in the set of the nearest neighbor center points of the center point does not exceed , so a total of times of calculating the specified distance value are required when re-determining the set of the nearest neighbor center points of these center points. The number of different nodes containing at least one node in the contracted cluster in the neighbor set does not exceed , and when updating the neighbor sets of these nodes, the set of the nearest neighbor center points of the updated nodes of these nodes needs to be considered, and at most only one in the original set of the nearest neighbor center points of each of these nodes is replaced (that is, the center point of the contracted cluster is replaced by other center points), so only the nodes within the cluster where the other center point replacing the center point of the contracted cluster is located (not exceeding ones) are used to update the neighbor set, and the updated neighbor sets of these nodes can be obtained. That is to say, a total of times of calculating the specified distance value are required to update the neighbor sets of these no more than nodes. In short, when deleting a single reference string triggers the deletion of the contracted cluster, a total of times of calculating the specified distance value are required, otherwise no calculation of the specified distance value is required. Although the time cost of removing the entire contracted cluster is relatively high, that is, times of calculating the specified distance value are required, but since this removal operation only occurs occasionally and can be completed independently by a background thread, the impact on the string retrieval performance is not significant.

[0097] In one embodiment, performing addition, deletion, and modification operations on the index structure according to the addition, deletion, and modification instructions includes: when the operation type of the addition, deletion, and modification instruction is modification, obtaining the reference string before change and the reference string after change corresponding to the operation object; first marking the reference string before change as a deleted node to perform the addition, deletion, and modification operation with the operation type of deletion, and then marking the reference string after change as an inserted node to perform the addition, deletion, and modification operation with the operation type of addition.

[0098] In one embodiment, the specific execution methods for the add / delete / update operations with the operation types of modification and deletion have been described in detail above and will not be elaborated here. Modifying a single reference string is equivalent to deleting the original reference string and adding the modified reference string. Therefore, modifying a single reference string requires no more than times of calculating the specified distance value.

[0099] It should be noted that in the retrieval operations described above, the applicable edit distance may not be limited to the edit distance with only three operations of addition, deletion, and modification. The operation of modifying characters can be further divided into operations such as replacing characters with the same or similar pronunciation, replacing characters with similar glyphs, and replacing general characters, and different unit costs can be set for each. In addition, the applicable edit distance also allows operations such as introducing modification of words or phrases and swapping adjacent two characters. Adopting different edit distance settings should not be considered as exceeding the technical scope protected by this application.

[0100] At the same time, this application does not describe the specific implementation method for the operation of modifying nodes in the index update process. The technical solution only points out that the operation of modifying nodes is equivalent to deleting the old nodes and adding new nodes. This equivalence means that there are multiple specific feasible implementation methods, including (1) performing the deletion operation first and then the addition operation, (2) performing the addition operation first and then the deletion operation, and (3) marking the deletion first and then performing the addition operation, and finally performing the removal operation of shrinking the cluster under certain conditions, etc. These specific implementation methods are very intuitive and easy to think of, so limiting the specific implementation method should not be considered as exceeding the technical scope protected by this application.

[0101] Step S130: According to the obtained retrieval instruction, perform - the nearest neighbor retrieval operation on the index structure, and return - the retrieval result output by the nearest neighbor retrieval operation.

[0102] In one embodiment, when the retrieval instruction is obtained, perform - the nearest neighbor retrieval operation on the index structure to implement string retrieval, that is, calculate the reference strings with the smallest edit distance to the query string. The content of the retrieval instruction includes the query string, the expected result set size and the retrieval distance. Among them, the retrieval distance is a preset edit distance from one string to another string, which is different from the specified distance used in the index construction process mentioned later, but is an edit distance defined separately.

[0103] - The nearest neighbor retrieval operation specifically includes: marking the query string as a test node and randomly selecting A node without a deletion mark, marked as the entry node of the first layer, and the deletion mark of the node is determined according to the index structure. The deletion mark will be specifically described in the add, delete, and modify operations for the index structure, which will be elaborated later. For now, it suffices to know that some nodes in the index structure will be marked with deletion marks. If the number of nodes without deletion marks in the first-layer directed graph is less than , then all nodes without deletion marks are marked as the entry nodes of the first-layer directed graph.

[0104] Start processing layer by layer from the first layer of the neighbor graph until the processing of the th layer of the neighbor graph is completed. The processing operations include: after determining the set of entry nodes of the th layer directed graph, first mark the set of entry nodes as the node set at time 0 . Then, according to the gradually increasing time , successively determine the node set at the th time from the node set at the th time . Mark the union of the neighbor sets of all nodes in as the first union set, and mark the union of the first union set and as the second union set. Then, determine as the set consisting of the nodes without deletion marks with the smallest retrieval distance to the test node in the second union set. If the number of nodes without deletion marks in the second union set is less than , then determine as the set consisting of all nodes without deletion marks in the second union set.

[0105] will necessarily converge within a finite time. When = , that is, when converges, mark as .

[0106] Judge whether the current processing layer has reached the maximum layer of the neighbor graph of the index structure: If , that is, the current layer is not the maximum layer of the neighbor graph, then mark as the set of entry nodes of the th layer directed graph and continue the processing of the th layer. Correspondingly, if the current processing layer is the maximum layer of the neighbor layer of the index structure, that is, when , mark as the retrieval result and output.

[0107] The most costly operation in the above string retrieval method is the operation of calculating the edit distance between two strings. Therefore, the time complexity of the string retrieval method can be measured by the number of times of calculating the edit distance. Since the edit distance needs to be calculated times for each layer in the layer neighbor graph, where is the number of layers of the neighbor graph, is the number of nodes included in the retrieval result, is the maximum number of nodes in a cluster, is the average number of iterations when each layer converges. Therefore, the above string retrieval method needs to calculate times of edit distance in total. Relatively speaking, the original string retrieval method needs to calculate times of edit distance, where is the number of reference strings. Since generally is very small (usually between 1 and 9), and is a constant much smaller than , the string retrieval method of the present application can be regarded as a constant-level approximate retrieval algorithm independent of , and its time complexity is significantly lower than the linear level of the original retrieval algorithm.

[0108] The present application proposes a method for accelerating the construction and update of the index by using incremental center point clustering. Based on the key point of the index structure of the layer neighbor graph is to limit the number of neighbors of each node, so that the retrieval process can be carried out in the finite neighborhood of the entry node set rather than the entire node set, thereby controlling the number of nodes traversed in the retrieval process. In order to limit the number of neighbors of the inserted node to the preset threshold or less, directly adapting the index construction method of HNSW requires finding the nodes with the smallest specified distance value from the inserted node among all the added nodes. Therefore, directly adapting the index construction method of HNSW requires calculating times of the specified distance value, and adding a node to the index requires calculating times of the specified distance value, where is the number of reference strings. The present application maintains a correspondence table between the center point of each layer and the set of the nearest neighbor center points of its cluster and the center point during the index construction and incremental addition of nodes, so as to accelerate finding the nodes with the smallest specified distance value from the inserted node among the added nodes. Specifically, when looking for these N nodes, the technical solution of the present application only needs to consider at most nodes within the cluster where the nearest neighbor center point of the inserted node is located, and determining the set of the nearest neighbor center points of the inserted node only requires calculating The specified distance value between the center point of the cluster where the node is inserted next time and other center points. Therefore, when adding a node to the index, only the number of specified distance values needs to be calculated. When constructing the multi-layer neighbor graph, the technical solution of this application completes the calculation of the specified distance values between any two center points in each layer at one time. Therefore, the index construction process is controlled within the calculation of the number of specified distance values. Since generally is much smaller than and is a relatively large constant, the time cost of the technical solution of this application in the index construction process and the incremental node addition process is significantly lower than the time cost of directly adapting HNSW.

[0109] In addition, in the string retrieval method based on this index structure, the operation with the highest cost is the operation of calculating the edit distance between two strings. Therefore, the time complexity of the string retrieval method can be measured by the number of edit distance calculations. Since each layer in the multi-layer neighbor graph needs to calculate the number of edit distances, where is the number of layers of the neighbor graph, is the number of nodes included in the retrieval result, is the maximum number of nodes in a cluster, is the average number of iterations when each layer converges. Therefore, the above string retrieval method needs to calculate a total of the number of edit distances. Relatively speaking, the original string retrieval method needs the number of times (i.e., linear times measured by ). Since generally is very small (usually between 1 and 9), and is much smaller than , the string retrieval method of this application can be regarded as a constant-level approximate retrieval algorithm independent of , and its time complexity is significantly lower than the linear level of the original retrieval algorithm.

[0110] Figure 2The internal functional module relationship diagram of the string retrieval device in an embodiment is shown. The string retrieval device 20 includes: an index construction module 21, an index maintenance module 22, and a retrieval module 23. The index construction module 21 is configured to, when there is no index structure, obtain a given set of reference strings for index structure construction operations. The index maintenance module 22 is configured to, when the set of reference strings corresponding to the index structure changes, generate add, delete, and modify instructions according to the change situation of the set of reference strings, and perform add, delete, and modify operations on the index structure according to the add, delete, and modify instructions. The retrieval module 23 is configured to, according to the obtained retrieval instruction, perform - nearest neighbor retrieval operations and return the retrieval results output by the retrieval operations.

[0111] Figure 3 The internal structure diagram of a computer device in an embodiment is shown. The computer device may specifically be a terminal or a server. As Figure 3 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor may implement the string retrieval method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor may execute the string retrieval method. Those skilled in the art can understand that Figure 3 the structure shown in

[0112] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0114] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0115] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A character string retrieval method, characterized in that: The steps include: When the index structure does not exist, obtaining a given reference string set to perform an index structure construction operation; the index structure is composed of a neighbor graph and a correspondence table; The neighbor graph is given by Layer directed graph, is a preset integer; each layer of the directed graph is composed of nodes, and one node corresponds to one reference string; No. The directed graph stores all given reference strings; for not less than 1 and not greater than Any integer , No. The reference string set stored in the directed graph of the first layer is A subset of the reference string set stored in the directed graph; Part of the nodes of each layer of the directed graph are marked as center points, and the center points are determined according to the index structure construction operation. The center points divide all the nodes in the directed graph into multiple non-overlapping subsets, each of which is called a cluster, and the number of nodes in the cluster does not exceed , is a preset integer; in each layer of the directed graph, there are directed edges between some of the nodes, and the directed edges are determined according to the index structure construction operation; the nodes pointed by the directed edges are called neighbors of the nodes initiating the directed edges, and the number of neighbors of each node does not exceed ; Each layer of the neighbor graph has a unique corresponding table, and the corresponding table is a two-dimensional table consisting of rows and columns; each row corresponds to a center point of the corresponding layer in the neighbor graph; the columns have three columns: the first column stores the reference string corresponding to the center point; the second column stores the reference string corresponding to other nodes in the cluster where the center point is located except the center point; the third column stores the reference string set corresponding to the nearest neighbor center point set of the center point, and the nearest neighbor center point set of the center point is determined by the index structure construction operation; When the reference string set corresponding to the index structure changes, generating an add, delete, or modify instruction according to the change of the reference string set, and performing add, delete, or modify operations on the index structure according to the add, delete, or modify instruction; According to the obtained search instruction, execute -Nearest neighbor search operation, returning the - The search result output by the nearest neighbor search operation; the search instruction content includes the query string, the expected result set size and a search distance, wherein the search distance is a preset edit distance from one character string to another character string; Said -Nearest neighbor search operations include: Mark the query string as a test node; randomly select A node without a deletion mark is marked as the entry node of the first layer. The deletion mark of the node is determined according to the index structure. If the number of nodes without a deletion mark in the directed graph of the first layer is less than , then all nodes without deletion marks are marked as entry nodes of the directed graph described in the first layer; The processing operation is performed layer by layer starting from the first layer of the neighbor graph until the first layer of the neighbor graph The layer processing is completed, and the processing operation includes: In determining the After the entry node set of the directed graph is obtained, the entry node set is first marked as the node set at time 0. ; Then according to the gradually increasing time , in turn by Time Node Set Determine Time Node Set ,Will The union of the neighbor sets of all nodes in is marked as the first union, and the first union is combined with The union of is marked as the second union, then Determine the node with the smallest search distance to the test node in the second union The number of nodes without deletion marks in the second union is less than , then Determine a set consisting of all nodes without deletion marks in the second union; when = Time will Mark as ; like , then Marked as The entry node set of the directed graph at the layer and continue Layer processing; like , then Mark as the search result and output.

2. The character string search method according to claim 1, wherein: The index structure building operation includes: Determine the minimum level at which each reference string is added to the neighbor graph l Probability of layer : ; is a preset real number greater than 1; according to Randomly determine the minimum level at which each given reference string is added to the neighbor graph l , and add the given reference strings as nodes to the first node of the neighbor graph one by one Layer to In the directed graph of layers; When adding a corresponding node of the reference string in each layer of the directed graph, marking the node to be added as an insertion node; The first inserted node is regarded as the first center point and forms a cluster by itself; subsequently, in the insertion process, the specified distance value between the newly inserted node and the existing center point is calculated one by one, and the specified distance value is a preset literal distance measurement value between the two nodes that satisfies non-negativity, symmetry and triangle inequality; Marking the cluster where the center point with the smallest specified distance value is located as an insertion cluster, and adding the insertion node to the insertion cluster; When the number of nodes in the inserted cluster exceeds , split the inserted cluster into two new clusters, the center points of the two new clusters are two different nodes in the inserted cluster, satisfying that the sum of the deviation values ​​from other nodes in the inserted cluster to the two new clusters is the minimum, and the deviation value of each node to the two new clusters is the minimum value of the specified distance value between the node and the center points of the two new clusters; the other nodes in the extended cluster except the center points of the two new clusters are called to-be-processed nodes, and the specified distance values ​​between the to-be-processed nodes and the center points of the two new clusters are calculated; the to-be-processed nodes are added to the new cluster with the smaller specified distance value; if the specified distance values ​​between the to-be-processed nodes and the center points of the two new clusters are the same, then the to-be-processed nodes are added to the new cluster with fewer nodes; After all given reference strings are added as nodes to the neighbor graph, the clusters included in each layer of the directed graph of the neighbor graph are determined; the node with the smallest sum of specified distance values ​​from other nodes in each cluster is calculated and marked as the center point of each cluster; For satisfaction Every integer , according to the neighbor graph All the center points and other nodes in the layer determine the neighbor relationship to build the first the correspondence table of layers; The neighbor graph and the corresponding table of each layer thereof are aggregated to obtain the index structure.

3. The character string search method according to claim 2, wherein: The neighbor graph is All the center points and other nodes in the layer determine the neighbor relationship to build the first The corresponding table of layers includes: Calculate the The specified distance value between the center points in the layer; according to the specified distance value between the center points, determine for each center point other center points with the smallest specified distance value, and The set of other center points is marked as the nearest neighbor center point set of the center point; For the l For each node in the layer, obtain the center point of the cluster where the node is located, and all the center points in the nearest neighbor center point set of the center point of the cluster, a total of center point; The set of center points is marked as the nearest neighbor center point set of the node; Get the cluster of all the center points in the nearest neighbor center point set of each node, with the smallest distance value to the node. other nodes, The set of other nodes is marked as the neighbor set of the node; Add the node to the directed graph of the layer directed edges to other nodes; According to The reference string corresponding to each of the center points in the layer, the reference strings corresponding to other nodes in the cluster where the center point is located, and the set of nearest neighboring center points of the center point are determined. The correspondence table of the layer.

4. The character string search method according to claim 1, wherein: The add, delete and modify instruction includes an operation type and an operation object, and the operation type is one of add, delete and modify; When the operation type is adding, the operation object is the reference string to be added; When the operation type is deletion, the operation object is the reference character string to be deleted; When the operation type is modification, the operation objects are the reference character string before modification and the reference character string after modification.

5. The character string search method according to claim 4, wherein: The performing the add, delete, and modify operations on the index structure according to the add, delete, and modify instructions includes: When the operation type of the add, delete or modify instruction is add, obtaining a reference string to be added corresponding to the operation object, and marking the reference string to be added as an insert node; Determine the minimum level at which the inserted node is added to the neighbor graph Probability of layer : ; is a preset real number greater than 1; according to Randomly determine the minimum layer to which the inserted node is added , and add the inserted node to the neighbor graph Layer to In the directed graph of layers; When adding the inserted node to each layer of the directed graph, calculating the specified distance value between the inserted node and all the center points in the directed graph, marking the cluster where the center point with the smallest specified distance value is located as an extended cluster, and adding the inserted node to the extended cluster; When the number of nodes in the extended cluster after adding the inserted node exceeds , split the extended cluster into two new clusters, the center points of the two new clusters are two different nodes in the extended cluster, so that the sum of the deviation values ​​from other nodes in the extended cluster to the two new clusters is minimized, and the deviation value of each node to the two new clusters is the minimum value of the specified distance value between the node and the center points of the two new clusters; the other nodes in the extended cluster except the center points of the two new clusters are called to-be-processed nodes, and the specified distance values ​​between the to-be-processed nodes and the center points of the two new clusters are calculated; the to-be-processed nodes are added to the new cluster with the smaller specified distance value; if the specified distance values ​​between the to-be-processed nodes and the center points of the two new clusters are the same, then the to-be-processed nodes are added to the new cluster with fewer nodes; Calculate the specified distance value between the center point of the two new clusters and other center points, and determine the center point of each new cluster with the minimum specified distance value according to the specified distance value. other center points, The set of other center points is marked as the nearest neighbor center point set of the center point of the new cluster; For each node in the two new clusters, obtain the center point of each node corresponding to the new cluster and all the center points in the nearest neighbor center point set of the center point of the new cluster, a total of center point; The set of center points is marked as the nearest neighbor center point set of each node in the new cluster; Get the cluster of all the center points in the nearest neighbor center point set of each node, with the smallest distance value to the node. other nodes, The set of other nodes is marked as the neighbor set of the node; Deleting the row where the center point of the extended cluster is located from the correspondence table, and adding corresponding row information of the center points of the two new clusters to the correspondence table according to the node sets, center points of the two new clusters and the nearest neighbor center point sets of the center points of the new clusters; Mark each node in the neighbor set of the inserted node that does not belong to any of the new clusters as a neighbor to be processed, first add the inserted node to the neighbor set of the neighbor to be processed, and then delete the neighbor with the largest specified distance value from the neighbor set of the neighbor to be processed; When the number of nodes in the extended cluster after adding the inserted node does not exceed , obtain the center point of the cluster where the inserted node is located, and all the center points in the nearest neighbor center point set of the center point of the cluster, a total of center point; The set of center points is marked as the nearest neighbor center point set of the inserted node; Get the cluster of all the center points in the nearest neighbor center point set of the inserted node, which has the smallest specified distance value from the inserted node other nodes; The set consisting of other nodes is marked as the neighbor set of the inserted node; The second column information of the row where the center point of the extended cluster is located in the corresponding table is updated according to the node set of the extended cluster; each node in the neighbor set of the inserted node is marked as a neighbor to be processed; the inserted node is first added to the neighbor set of the neighbor to be processed, and then the neighbor with the largest specified distance value is deleted from the neighbor set of the neighbor to be processed.

6. The character string search method according to claim 4, wherein: The performing the add, delete, and modify operations on the index structure according to the add, delete, and modify instructions includes: When the operation type of the add, delete or modify instruction is deletion, obtaining a reference string to be deleted corresponding to the operation object, and marking the reference string to be deleted as a deletion node; Processing the directed graph and the corresponding table containing the deleted node layer by layer for the neighbor graph; determining the cluster where the deleted node is located according to the corresponding table, marking the cluster where the deleted node is located as a contracted cluster, and marking the deleted node with a deletion mark; When all nodes in the contracted cluster have deletion marks, all nodes in the contracted cluster are deleted from the directed graph, and the row where the center point of the contracted cluster is located is deleted from the corresponding table; For each center point in the directed graph, if the center point of the shrinking cluster is included in the set of the nearest neighboring center points of the center point, the center point in the directed graph with the smallest specified distance value to the center point is re-determined. other center points, and update the nearest neighbor center point set of the center point to be composed of these The set of other center points; For each node in the directed graph, if the neighbor set of the node includes at least one node in the contracted cluster, the center point of the cluster where the node is located and all the center points in the nearest neighbor center point set of the center point of the cluster where the node is located are re-obtained, and a total of center point; The set of center points is updated as the nearest neighbor center point set of the node; the cluster of all center points in the nearest neighbor center point set of the node with the smallest specified distance value to the node is re-obtained. other nodes, The set consisting of other nodes is updated as the neighbor set of the node; When there are still nodes without deletion marks in the contracted cluster, no further processing is performed.

7. The character string search method according to claim 4, wherein: The performing the add, delete, and modify operations on the index structure according to the add, delete, and modify instructions includes: When the operation type of the add, delete or modify instruction is modification, obtaining a reference string before modification and a reference string after modification corresponding to the operation object; First, the reference string before the change is marked as a deletion node to perform the addition, deletion and modification operation with the operation type of deletion, and then the reference string after the change is marked as an insertion node to perform the addition, deletion and modification operation with the operation type of addition.

8. A character string search device, characterized in that: The device refers to the character string retrieval method according to any one of claims 1 to 7, and the device includes: An index building module, used for obtaining a given reference string set to perform an index structure building operation when no index structure exists; An index maintenance module, configured to generate an add, delete, or modify instruction according to the change of the reference string set when the reference string set corresponding to the index structure changes, and perform an add, delete, or modify operation on the index structure according to the add, delete, or modify instruction; A retrieval module is used to execute a search for the index structure according to the obtained search instruction. -Nearest neighbor search operation, returning the search result output by the search operation.

9. A computer device, characterized in that: including a processor and a memory; The processor is configured to execute the computer program stored in the memory to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Distributed vector indexing and retrieval method and system using memory and disk in mixed mode

    CN118467544A

  • Read-write method of solid state disk and computer readable storage medium

    CN119200962A