Reference distance similarity search

By using the bin's hierarchical database and similarity searcher in the associated memory array, the problem of inefficient similarity search in n-dimensional space is solved, fast and accurate similarity search is achieved, and storage utilization is improved.

CN112199408BActive Publication Date: 2025-05-27GSI TECHNOLOGY INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202010522234.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-01
Filing Date
2020-06-10
Publication Date
2025-05-27
Estimated Expiration
2040-06-10

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency in fast and accurate similarity searches in n-dimensional spaces, especially in similarity searches in large data sets.

Method used

Fast similarity search is achieved by using a hierarchical database of bins in an associative memory array, using an order vector to represent the original vector, and operating in multiple columns simultaneously through a similarity searcher.

Benefits of technology

Improves the efficiency and accuracy of similarity search, reduces search time, provides O(1) search complexity, and improves storage utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112199408B_ABST
    Figure CN112199408B_ABST
Patent Text Reader

Abstract

A similarity search system includes a database of original vectors, a hierarchical database of bins, and a similarity searcher. The hierarchical database of bins is stored in an associative memory array, each bin being identified by a rank vector representing at least one original vector, and the dimension of the rank vector being less than the dimension of the original vector. The similarity searcher searches the database for at least one similar bin, the rank vectors of the similar bins being similar to the rank vector representing the query vector, and the similarity searcher provides at least one original vector similar to the query vector represented by the bin.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 871,212, filed on Jul. 8, 2019, and U.S. Provisional Patent Application No. 63 / 003,314, filed on Apr. 1, 2020, both of which are incorporated herein by reference. Technical Field

[0003] The present invention generally relates to associative computing and, in particular, to data mining algorithms using associative computing. Background Art

[0004] Data mining is a computational process for discovering patterns in large datasets. Data mining uses different techniques to analyze datasets. A frequently required computation in data mining is fast and accurate similarity search in an n - dimensional space, where each item stored in a large dataset in the space is represented by a vector of n floating - point numbers. The purpose of similarity search is to quickly identify items in the dataset that are similar to a specific query item, which is also represented by a vector of n floating - point numbers.

[0005] Throughout the document, a space containing L vectors of dimension S is represented as E={E1, E2……El}(|E| = L), the query vector is represented as Q (which is also of dimension S), and a general vector in space E is represented as Ei(0 < i < L). The purpose of the search is to find a subset of K vectors Ei∈E (K << L) that are most similar to Q (i.e., have the smallest distance from Q).

[0006] One of the state - of - the - art solutions for finding the set of K items Ei that are most similar to query Q is to utilize a K - nearest neighbor search algorithm using a distance function (e.g., L2 distance, cosine distance, Hamming distance, etc.). Summary of the Invention

[0007] According to an embodiment of the present invention, a similarity search system is provided. The system includes a database of original vectors, a hierarchical database of bins, and a similarity searcher. The hierarchical database of bins is stored in an associative memory array, each bin is identified by an order vector representing at least one original vector, and the dimension of the order vector is less than the dimension of the original vector. The similarity searcher searches the database for at least one similar bin, the order vectors of these similar bins are similar to the order vector representing the query vector, and the similarity searcher provides at least one original vector similar to the query vector represented by the bin.

[0008] Additionally, according to an embodiment of the present invention, the bins of the hierarchical database are stored in columns of the associative memory array, and the similarity searcher operates on multiple columns simultaneously.

[0009] In addition, according to a preferred embodiment of the present invention, the hierarchical database is arranged by level, and each level is stored in a different part of the associated memory array.

[0010] In addition, according to a preferred embodiment of the present invention, the system includes a hierarchical database builder for building a hierarchical database of bins based on a database of original vectors.

[0011] Furthermore, according to a preferred embodiment of the present invention, the hierarchical database builder includes a reference vector definer, an order vector creator, and a bin creator. The reference vector definer defines a set of reference vectors in terms of the dimensions of the original vectors. The order vector creator calculates the distance from each original vector to each reference vector and creates an order vector that includes the IDs of the reference vectors sorted by the distance of the reference vectors from the original vector, and the bin creator creates bins identified by the order vectors representing at least one original vector.

[0012] Additionally, according to a preferred embodiment of the present invention, the hierarchical database builder clusters order vectors representing different original vectors sharing an order vector into a single bin.

[0013] In addition, according to a preferred embodiment of the present invention, the hierarchical database includes at least two levels, and the bins in one level are associated with the bins in a lower level.

[0014] In addition, according to a preferred embodiment of the present invention, the similarity searcher starts the search in the first level of the hierarchical database and continues the search in the lower levels for bins associated with the bins found in the first level.

[0015] According to an embodiment of the present invention, a method for finding a set of vectors similar to a query vector in a database of original vectors is provided. The method includes: accessing a set of reference vectors; creating a query order vector associated with the query vector using the reference vectors, the query order vector having a dimension smaller than that of the query vector. The method further includes searching for at least one similar bin in a hierarchical database of bins stored in an associated memory array, wherein each bin represents at least one original vector and is identified by an order vector created using the set of reference vectors, and the order vectors of the at least one similar bin are similar to the query order vector. The method further includes providing at least one original vector similar to the query vector represented by the similar bin.

[0016] In addition, according to a preferred embodiment of the present invention, the hierarchical database stores the bins in columns of the associated memory array, and the searching step operates on multiple columns simultaneously.

[0017] Furthermore, according to a preferred embodiment of the present invention, the method includes arranging the hierarchical database by level, each level in a different part of the associated memory array.

[0018] Additionally, according to a preferred embodiment of the present invention, the method includes constructing a hierarchical database of bins based on a database of original vectors.

[0019] Furthermore, according to a preferred embodiment of the present invention, the step of constructing the hierarchical database includes: defining a set of reference vectors in terms of the dimensions of the original vectors; calculating the distance from each original vector to each reference vector and creating an order vector that includes the IDs of the reference vectors sorted by the distance of the reference vectors from the original vector. The method also includes a bin creator for creating bins identified by the order vectors representing at least one original vector.

[0020] Moreover, according to a preferred embodiment of the present invention, the method further includes clustering order vectors representing different original vectors sharing an order vector into a single bin.

[0021] In addition, according to a preferred embodiment of the present invention, the hierarchical database includes at least two levels, and the bins in one level are associated with the bins in a lower level.

[0022] Additionally, according to a preferred embodiment of the present invention, the step of searching includes starting the search in the first level of the hierarchical database and continuing the search in the bins in the lower level that are associated with the bins found in the first level. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] What is regarded as the subject matter of the present invention is particularly pointed out and distinctly claimed in the concluding portion of the specification. However, the present invention, both as to its organization and method of operation, together with its objects, features, and advantages, may best be understood by reference to the following detailed description when read in conjunction with the accompanying drawings, in which:

[0024] Figure 1A-1E is a schematic diagram explaining the concept of distance similarity used by a system constructed and operated according to an embodiment of the present invention;

[0025] Figure 2 is a schematic diagram of a process for constructing a hierarchical database of vectors using the concept of distance similarity vectors implemented by a system constructed and operated according to an embodiment of the present invention;

[0026] Figure 3A is by Figure 2 a schematic diagram of an exemplary hierarchical database created by the process of;

[0027] Figure 3B is Figure 3A a schematic diagram of the arrangement of the hierarchical database of in an association processing unit (APU);

[0028] Figure 4 is implemented by a system constructed and operated according to an embodiment of the present invention Figure 2Schematic diagram of a hierarchical database builder for a process;

[0029] Figure 5 Schematic diagram of a process for finding a set of vectors similar to a query vector in a hierarchical database implemented by a system constructed according to an embodiment of the present invention;

[0030] Figure 6 Is constructed and operated according to an embodiment of the present invention to implement Figure 5 Schematic diagram of a similarity searcher for a process; and

[0031] Figure 7A and Figure 7B Are two alternative schematic diagrams of a similarity search system constructed and operated according to an embodiment of the present invention.

[0032] It will be appreciated that, for simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, for clarity, the dimensions of some of the elements may be exaggerated relative to other elements. Additionally, where considered appropriate, reference numerals may be repeated between the figures to indicate corresponding or similar elements. Detailed Description

[0033] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present invention. However, those skilled in the art will recognize that the present invention may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the present invention.

[0034] The applicant has recognized that associative memory devices such as those described in U.S. Patent No. 9,558,812 assigned to a co-applicant of the present invention can be efficiently utilized to provide accurate similarity search, such similarity search providing a set of the K most similar records to a query record Q with a latency of less than 100 microseconds. In devices such as those described in U.S. Patent No. 9,558,812, searches can be performed in parallel on multiple columns, thus providing an O(1) search complexity.

[0035] The applicant has further recognized that similarity searches performed on datasets in lower dimensional spaces can improve storage utilization while maintaining a high level of accuracy and the same latency. Additionally, when using distance similarity search instead of standard similarity search, the similarity search can provide sufficient results, which can further improve its performance.

[0036] Distance similarity search is based on heuristics, i.e., if the distance between vector E and vector Q is small (i.e., the vectors are similar to each other), then the distance between vector E and reference vector P is also similar to the distance between vector Q and vector P. In other words, when vector E is similar to reference vector P and vector Q is similar to the same reference vector P, vector E is similar to vector Q.

[0037] It can be appreciated that using an alternative dataset of vectors of lower-dimensional natural numbers OV instead of the original dataset of vectors of higher-dimensional floating-point numbers can improve storage requirements and computational performance. In the alternative database, each vector OVi can store the IDs of the reference vectors sorted by the distance of the reference vector from vector Ei (which means the position of the reference vector in the original space). The number of reference vectors can determine the dimension of the new space and can be set to be less than the number of original features of vector Ei.

[0038] The distance similarity concept is shown in Figure 1A-1E and is now referred to Figure 1A-1E . The distance between each vector Ei and a set of predetermined reference vectors Pj can be pre-computed, and a new vector OVi can be created that has the indices of the reference vectors Pj sorted by the distance of the reference vector Pj from Ei. For each new vector Q, a vector OVq (the indices of the reference vector P sorted by the distance of the reference vector P from Q) can be created and compared with the OVi vectors in the dataset to find the most similar vectors, from which the most similar Ei can be immediately determined.

[0039] Figure 1A Figure shows two vectors E1 and E2 and a query vector Q for which the most similar vector (E1 or E2) should be determined. First, a database of smaller dimension can be created by using a set of reference vectors (P1, P2, and P3), added to the space as shown in Figure 1B , the distance between each vector (E1 and E2) and the reference vectors (P1, P2, and P3) can be calculated, and for each vector E (E1 and E2), a new vector OV (OV1 and OV2) can be created that has the IDs of the reference vectors sorted by the distance (between the vector and the reference vector), as shown in Figure 1C and Figure 1D . In Figure 1CIn [context], the distances between E1 and each of the reference vectors (P1, P2, and P3) can be calculated and the distances indicated as D1-1 (distance between E1 and P1), D1-2 (distance between E1 and P2), and D1-3 (distance between E1 and P3). The calculated distances can be sorted, and a new distance vector OV1 can be created using the IDs of the reference vectors P from closest to farthest. In Figure 1C In [context], the smallest distance is D1-1, then D1-2, then D1-3, so the vector OV1 is [1, 2, 3]. In Figure 1D In [context], the same process can be performed for E2, and the resulting vector OV2 is [3, 2, 1]. The same process can be performed for each vector Ei in a large dataset.

[0040] Figure 1E Shows the same process performed for query Q, with the result that the vector OVq is [1, 2, 3]. The resulting vector OVq can be compared with all other OV vectors (OV1 and OV2). In this example, the most similar vector is OV1 (which is [1, 2, 3]), meaning that the query vector Q is more similar to vector E1 than to vector E2.

[0041] It can be recognized that the dimension of the original vector Ei may be large, and the data stored in vector Ei can be represented by floating-point numbers, while the dimension of the new OVi vector (i.e., the number of reference vectors P) may be much smaller, and the data can be represented by natural numbers, thus reducing the size and complexity of the data to be searched.

[0042] The applicant has further recognized that storing the dataset of OV vectors in a hierarchical structure and reducing the search to a subset of the records can improve the performance of the search and can provide good response time, high throughput, and low latency.

[0043] Now referring to Figure 2 is a schematic diagram of process 200 for constructing a hierarchical database of OVi vectors from a dataset of Ei vectors, which process 200 is implemented by a system constructed according to an embodiment of the present invention. Process 200 can be performed once on the entire original database storing vectors of higher-dimensional floating-point numbers, and can produce a new hierarchical database storing vectors of smaller-dimensional natural numbers.

[0044] This preprocessing process can reduce the space required to perform the search from the original dimension S to a smaller dimension M (M ≤ S). This process can create a vector of M natural numbers for each original item in the space. Additionally, this process can cluster several such vectors of the original space into bins of lower-dimensional distance vectors, where each bin includes a list of original Ei vectors sharing the same OV. Each bin can be associated with a small descriptor including the bin ID and the OV. The new structure of the bins can be stored in an associative memory array, where an associative tree search can be performed in parallel on multiple bins to find bins similar to the query.

[0045] The input 211 of the process includes the entire original data set with L vectors Ei, each vector having a dimension of S, i.e., including S floating-point numbers. In step 220, the system can be initialized with the number n of levels to be created in the new hierarchical database and the ID of the first level. It should be noted that the number of levels in the hierarchical database can be 1.

[0046] In step 230, the system can be configured to select M Pj (j = 1... M, M <= S) reference vectors of dimension S. The process of selecting the M Pj reference vectors is described below. In step 240, the system can loop through all the bins at this level, and in step 250, the system can be configured to create the bins of the next level. Specifically, in sub-step 252, the system can calculate the order vector OVi for each vector Ei (i = 1... l) by calculating the distance Di-j to each reference vector Pj (j = 1... M); sort the calculated values of Di-j and create a new vector OVi with ID j for each vector Ei, as explained above with respect to Figure 1C and Figure 1D It can be recognized that the size of the order vector OVi can be R (R ≤ M), such that the order vector OVi only contains the R lowest values of Di-j, thus reducing the space to a value lower than the number of reference vectors Pi. This process can support the selection of a large number of reference vectors while maintaining a small search space. In sub-step 254, the system can be configured to cluster all the same vectors OVi into separate bins and add an association between the parent bin and the child bins. Each created bin can include: the bin ID; the level to which the bin belongs; the value of the vector OV common to all the vectors Ei clustered in the bin, and a list of pointers to the vectors Ei included in the bin (e.g., a list of values i).

[0047] In step 260, the system can be configured to check whether the most recently created level of the hierarchical database should be the final level. If the created level is not the last level, the system can proceed to the next level and can return to steps 230, 240, and 250 to create bins for the next level. If the created level is the last level, the system can provide the hierarchical database 281 of the vectors OVi arranged in the bins as output. In one embodiment, the hierarchical database 281 can be stored in an associated processing unit (APU) on a system that can perform parallel search operations on multiple columns, with each OV stored in a column of the memory array of the APU.

[0048] Now referring to Figure 3A is a schematic diagram of an exemplary hierarchical database 28 arranged in three levels 310, 320, and 330. The first level 310 includes: bin 1; bins 2 and 3. The second level 320 includes the bins of the next level. The next level below bin 1 includes bins 1.1 and 1.2. These levels can be connected by lines that indicate the parent-child associations between the bins in different levels. For example, bin 3.3.2 is in the third level of bins and is a child bin of bin 3.3 in the second level, which is also a child bin of bin 3 in the first level. It can be appreciated that using this type of hierarchical database can reduce the number of records to which OVq should be compared to a subset of OVi.

[0049] Now referring to Figure 3B is a schematic diagram of the arrangement of the hierarchical database 281 in the APU 380. The three levels 310, 320, and 330 of the hierarchical database 281 can be stored in different parts of the APU 380, and each part can be activated when a search is performed for that level. Due to all the columns of the parallel search part, this arrangement of data in the APU can achieve a parallel similarity search with a complexity of O(1).

[0050] Now referring to Figure 4 is a schematic diagram of a hierarchical database builder 400 that constructs and operates the implementation process 200 ( Figure 2 ) according to an embodiment of the present invention. The hierarchical database builder 400 includes: a reference vector definer 410; an order vector creator 420, and a bin creator 430. The hierarchical database builder 400 can implement the process 200 on the original database 211 of vectors of dimension S received as input and can create a hierarchical database 281 of vectors of dimension M as output.

[0051] The reference vector definer 410 can define the reference vector Pi to be used in each bin for creating the next level of reference vectors. The reference vector Pi can be defined for each level or each bin. The reference vector definer 410 can select random reference vectors Pi, or can use a clustering method (e.g., K-means) to create the reference vector Pi based on the records Ei associated with the bin. Alternatively, the reference vector definer 410 can use a trained machine learning application to find a set of reference vectors, resulting in a small set of highly accurate search results. After training, the machine learning application can be used on the bins to find the reference vector Pi to be used at that level.

[0052] The order vector creator 420 can implement step 252 of process 200 to calculate the order vector OVi for any given vector Ei, where the order vector OVi includes the IDs of the reference vectors sorted by the distance of the reference vectors (for which the distance to that reference vector has been calculated) from Ei.

[0053] The bin creator 430 can implement step 254 of process 200 to cluster all similar OVs into a single bin, where each bin includes an ID, the OV representing the bin, a list of references to the original Eis, and an indication of the level of the bin in the hierarchy. The bin creator 430 can use several methods to cluster the OVs into a single bin.

[0054] Now referring to Figure 5 is a schematic diagram of process 500 implemented by a system constructed according to an embodiment of the present invention. Process 500 can receive a query Q as input 511 and can use the hierarchical database 281 to find vectors Ei in the database 211 that are similar to the query vector Q.

[0055] In step 520, the system can be initialized with level zero and all bins selected, i.e., potentially starting with all vectors Ei in the database 211. In step 530, the system can use a process similar to the process described with respect to sub-step 252 to create the OVq vector for the query vector Q related to the relevant reference vector Pi, i.e., the system can be configured to calculate the distance Dq-j to each reference vector Pj (j = 1... M), sort the calculated values of Dq-j, and create the vector OVq of the R lowest values of Dq-j, which has the ID j.

[0056] In step 540, the system may loop through all bins in the level, and in step 550, the system may perform a similarity search between the OVq and OVi of each bin in the processed level. In step 560, the similarity score may be compared with a predefined threshold. If the similarity score is higher than the threshold, then in step 564, the selected processed bin may be retained, indicating that the vector Ei associated with the bin is considered similar to the query vector Q; however, if the similarity score is lower than the predefined threshold, then in step 566, the system may remove the bin because the vector Ei associated with the bin is considered different from the query vector Q.

[0057] In step 570, the system may check whether the search has reached the last level of the database. If the search has not reached the last level, the system may increment the level in step 580 and may continue the search. If the search has reached the last level, the search is considered complete, and in step 592, the system may return all vectors Ei pointed to by the retained selected bins. The OV of the returned bins is found to be similar to OVq, and thus the vectors Ei associated with these bins are similar to the query vector Q.

[0058] The similarity threshold may be determined for each bin or each level, and the threshold may be changed (i.e., decreased) when the resulting set of recorded Eis is too large. Process 500 may start at any level (including the last level), which means that a distance similarity search is performed on all lower-level bins (leaf bins) and the tree is not pruned.

[0059] Now referring Figure 6 is a schematic diagram of a similarity searcher 600 implemented and operated in accordance with an embodiment of the present invention to implement process 500 ( Figure 5 ). The similarity searcher 600 includes a rank vector creator 420 (e.g., the rank vector creator 420 used in the hierarchical database builder 400), a similar rank vector finder 610, and a bin converter 620. The similarity searcher 600 may communicate with the hierarchical database 281 and the database 211.

[0060] The rank vector creator 420 may implement step 530 of process 500 to calculate a rank vector OVq that includes the IDs of the reference vectors Pj for which the distance from the query vector Q has been calculated. The relevant reference vectors Pj may be the same reference vectors used to construct the bins.

[0061] A similar order vector finder 610 can perform a similarity search in the hierarchical database 281 stored in the associative memory, and can implement process 500 to find the bin associated with the OV that is most similar to OVq. The similarity search can operate in parallel on all bins of a level, and can find a set of similar order vectors OVi in a single search operation regardless of the number of bins in the level. The similarity search can be based on any similarity algorithm, e.g., Hamming distance algorithm, Euclidean distance algorithm, intersection similarity algorithm, etc.

[0062] The similarity search can be performed in parallel on all bins of a level using any similarity search algorithm. All vectors OVi stored in the columns of the APU 380 can be compared with the vector OVq simultaneously. In the Hamming algorithm, the similarity score can be the number of matching values in the matching positions of the vectors (i.e., vectors having the same value in the same position). In the intersection similarity algorithm, the similarity score can be the number of matching values ignoring the positions (i.e., the order of the values in the OV can be ignored and only the values are considered). In all methods, the similarity scores can be compared with a threshold, and only those similarity scores having a value greater than the threshold can be considered similar.

[0063] The bin converter 620 can deliver all vectors Ei associated with the selected bin. As mentioned above, the bin whose order vector is similar to the order vector of the query vector Q points to the vector Ei that is similar to the query vector Q.

[0064] As already mentioned above, it can be appreciated that storing the hierarchical database 281 in the associative memory array of the APU 380 can enable a parallel similarity search with a complexity of O(1). In addition, the size of the bin descriptor may be small (e.g., 64 bits), and thus, a large number of bins can be stored in a single APU capable of performing a parallel associative tree search on all bins of a level simultaneously.

[0065] Now referring to Figure 7A and Figure 7B are two alternative schematic diagrams of a similarity search system 700 constructed and operated from the components described above according to an embodiment of the present invention. The similarity search system 700 includes: a database 211 that stores the original vectors; a hierarchical database 281 of bins; a hierarchical database builder 400 for constructing the database 281 based on the database 211; and a similarity searcher 600 for receiving a query vector Q, performing a similarity search in the hierarchical database builder 400 to find the bin similar to the order vector representing the query vector Q, and providing a set of the original vectors Ei that are most similar to the query vector Q.

[0066] It will be appreciated that the steps shown in the exemplary processes above are not intended to be limiting, and the processes may be practiced with variations. These variations may include more steps, fewer steps, changing the sequence of steps, skipping steps, and other variations that may be obvious to those skilled in the art.

[0067] Although certain features of the invention have been shown and described herein, many modifications, substitutions, changes, and equivalents will now occur to those of ordinary skill in the art. Accordingly, it is to be understood that the appended claims are intended to cover all such modifications and changes that fall within the true spirit of the invention.

Claims

1. A similarity search system, comprising: an association processing unit (APU) including an associative memory array; a database including a plurality of original vectors; a hierarchical database builder for defining a plurality of different sets of reference vectors in terms of the dimensions of the original vectors, each reference vector having an ID, for calculating the distance from each original vector to each of the plurality of reference vectors, and for creating an order vector including the IDs of the plurality of reference vectors, the IDs of the plurality of reference vectors being sorted within the order vector according to the distance of each of the plurality of reference vectors from the original vector, wherein the dimension of the order vector is less than the dimension of the original vector; a bin hierarchy stored in columns of the associative memory array, each bin being identified by one of the order vectors and each bin further including a reference list of the set of original vectors associated with that bin; and a similarity searcher implemented on the APU for simultaneously searching for at least one similar bin in a plurality of columns of the associative memory array, the order vector of the similar bin being similar to the order vector representing the query vector, and the similarity searcher for providing the at least one original vector associated with the at least one similar bin, the at least one original vector being similar to the query vector, wherein the bin hierarchy includes at least two levels, each level being stored in a different part of the associative memory array and each level being associated with a different set of the plurality of different reference vectors, wherein bins in one level are associated with bins in a lower level, and the similarity searcher for starting the search in the first level of the bin hierarchy and continuing the search for bins in a lower level associated with the bins found in the first level.

2. The similarity search system according to claim 1, wherein the hierarchical database builder is further for clustering order vectors representing different but having the same order vector of original vectors into a single bin.

3. A method for finding a set of vectors in a database of original vectors, the set of vectors being similar to a query vector, the method comprising: defining a plurality of different sets of reference vectors in terms of the dimensions of the plurality of original vectors, each reference vector having an ID; calculating the distance from each original vector to each of the plurality of reference vectors; creating an order vector including the IDs of the plurality of reference vectors, the IDs of the plurality of reference vectors being sorted within the order vector according to the distance of each of the plurality of reference vectors from the original vector, wherein the dimension of the order vector is less than the dimension of the original vector; using the plurality of different sets of reference vectors to create a query order vector associated with the query vector; and storing a bin hierarchy in columns of an associative memory array of an association processing unit (APU), each bin being identified by one of the order vectors and each bin further including a reference list of the set of original vectors associated with that bin; and Performing a simultaneous search for at least one similar bin in multiple columns of the associative memory array, wherein an order vector of the at least one similar bin is similar to the query order vector; and Providing the at least one original vector associated with the at least one similar bin, wherein the at least one original vector is similar to the query vector, wherein the bin hierarchy includes at least two levels, each level is stored in a different part of the associative memory array, and each level is associated with a different set of the multiple different reference vectors, wherein bins in one level are associated with bins in a lower level, and the search is for starting the search in a first level of the bin hierarchy and continuing the search for bins in a lower level that are associated with the bins found in the first level.

4. The method according to claim 3, further comprising clustering order vectors representing original vectors that are different but have the same order vector into a single bin.

Citation Information

Patent Citations

  • SRAM multi-cell operations

    US9558812B2

  • Method for min-max computation in associative memory

    CN109426482A

  • Taxonomic classification system

    US20120278362A1

  • Multiple Message Retrieval for Secure Electronic Communication

    US20180107843A1