Information processing device, information processing method, and program

The described configuration for nearest neighbor search using a hash method addresses inefficiencies by managing reference data in a hash table with bins set to a first distance, enhancing search efficiency and accuracy in classifying determination targets.

JP2025104533APending Publication Date: 2025-07-10OMRON CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023222409
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing methods for nearest neighbor search using the hash method lack an effective way to appropriately delimit the hash, leading to inefficiencies in search efficiency and accuracy.

Method used

A configuration that includes a data storage unit, feature amount calculation unit, search area determination unit, distance calculation unit, and classification unit to manage reference data in a hash table with bins set to a first distance, allowing efficient classification by comparing distances and searching neighboring bins.

Benefits of technology

This approach enables high-speed and accurate classification of determination targets by reducing bias in the hash table, efficiently searching for nearest neighbors within a threshold distance, and accurately determining class membership.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025104533000001_ABST
    Figure 2025104533000001_ABST
Patent Text Reader

Abstract

To construct a hash table with high search efficiency in nearest neighbor search using a hash method.SOLUTION: An information processing device comprises: a data storage unit that stores, in a hash table in which the width of each bin is set to a first distance, reference data represented by a feature vector representing an object to be classified into a first class; a feature amount calculation unit that calculates a feature vector representing a determination object for which classification as the first class is to be determined; a search region determination unit that calculates a hash value by using the feature vector representing the determination object, and determines a bin corresponding to the hash value as a search-object bin in the hash table; a distance calculation unit that calculates the distance between each reference data item stored in the search-object bin and a query represented by the feature vector representing the determination object; and a classification unit that determines whether or not the determination object is classified into the first class by comparing the first distance and the distance between each reference data item and the query.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus, an information processing method, and a program for classifying an object to be determined using nearest neighbor search.

Background Art

[0002] Nearest neighbor search is a method for finding the nearest point in a distance space, and is a method for finding the point that is most similar (closest in distance) to the point indicated by a query from among reference data (points), that is, the nearest neighbor point. As a specific method of nearest neighbor search, a method using a hash method is known. For example, Patent Document 1 describes that in the search for the nearest neighbor point using the hash method, the search area is narrowed down based on the distance between the representative point in the area (bin) corresponding to the hash index and the query, thereby efficiently performing the search for the nearest neighbor point.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to perform high-speed and highly accurate nearest neighbor search using the hash method, it is effective to appropriately set the hash delimiter with respect to the distribution of each reference data to be searched. However, conventionally, a method for appropriately delimiting the hash in hash search has not been studied.

[0005] An object of the present invention is to construct a hash table with high search efficiency in nearest neighbor search using the hash method.

Means for Solving the Problems

[0006] In order to solve the above-described problems, the present invention adopts the following configuration. An information processing apparatus according to one aspect of the present invention includes: a data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to a first distance; a feature amount calculation unit that calculates a feature vector representing a determination target for determining whether or not the determination target is classified into the first class; a search area determination unit that calculates a hash value using the feature vector representing the determination target and determines a bin corresponding to the hash value as a bin to be searched in the hash table; a distance calculation unit that calculates a distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target; and a classification unit that determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance. According to the above configuration, it is possible to provide a hash search method with very high efficiency in a task of classifying a determination target by searching for the presence or absence of reference data whose distance from a query is within a predetermined threshold (first distance).

[0007] Further, the data storage unit may determine an element to be a base based on a statistic of each element of the feature vector representing each reference data. Thereby, the bias of the reference data stored in each bin of the hash table is reduced, and the search within the bin can be made efficient.

[0008] Further, the distance calculation unit may also set bins in the vicinity of the bin corresponding to the hash value as search targets. Thereby, it is possible to efficiently search for the nearest neighbor points (points whose distance from the query is equal to or less than the threshold) that may exist in the bins in the vicinity.

[0009] Further, when the distance calculation unit finds reference data whose distance from the query is equal to or less than the first distance, the distance calculation unit may terminate the calculation of the distance between the query and other reference data, and the classification unit may determine that the determination target is classified into the first class. Thereby, it is possible to efficiently and accurately classify the determination target.

[0010] Further, when no reference data whose distance from the query is equal to or less than the first distance is found, the classification unit may determine that the determination target is not classified into the first class. Thereby, it is possible to efficiently and accurately classify the determination target.

[0011] Further, the distance calculation unit may determine the search order of the neighboring bins based on the distances between the query and the representative points of the bins in each neighborhood. Thereby, since it is possible to search in order from the neighboring bins where there is a higher possibility of the existence of the nearest neighbor point, there is a high possibility of quickly finding reference data whose distance is equal to or less than the threshold value (the first distance), and the classification of the determination target can be efficiently performed.

[0012] Further, a hash table in which the coordinates of the axes or the reference data are transformed may be provided so that the number of bins including one or more reference data is maximized. Thereby, the number of reference data included in one bin can be reduced, enabling efficient hash search.

[0013] Further, the data storage unit may perform thinning of reference data whose mutual distances are equal to or less than the first distance. Thereby, the number of reference data to be searched can be appropriately reduced, enabling efficient hash search.

[0014] Further, the distance calculation unit determines the nearest neighbor point from among the reference data based on the distance between each reference data and the query, and the classification unit may classify the determination target by comparing the distance between the query and the nearest neighbor point with the first distance. Thereby, it is possible to classify the determination target after specifying the nearest neighbor point.

[0015] An information processing apparatus according to an aspect of the present invention includes a data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the minimum width of a bin is set to a first distance, a feature amount calculation unit that calculates a feature vector representing a determination target for determining whether or not the determination target is classified into the first class, a search region determination unit that calculates a hash value using the feature vector representing the determination target and determines a bin corresponding to the hash value as a bin to be searched in the hash table, a distance calculation unit that calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target, and a classification unit that determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance. According to the above configuration, in a task of searching for the presence or absence of reference data whose distance from a query is within a predetermined threshold (first distance), it is possible to provide a hash search method with very high efficiency.

[0016] An inspection apparatus according to an aspect of the present invention is an inspection apparatus that uses the above information processing apparatus to determine whether or not a determination target is classified as a non-defective product. The query is a point represented by a feature vector calculated from image data obtained by imaging the determination target, the reference data is a point represented by a feature vector calculated from image data of an object classified as a non-defective product, and the classification unit classifies the determination target as a non-defective product when there is reference data whose distance from the query is smaller than the first distance. According to the above configuration, it is possible to efficiently classify non-defective products and defective products using the image data of the determination target.

[0017] An information processing method according to an aspect of the present invention is an information processing method executed by a computer. The computer stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to a first distance. The computer calculates a feature vector representing a determination target for determining whether or not the determination target is classified into the first class. The computer calculates a hash value using the feature vector representing the determination target, and determines a bin corresponding to the hash value as a bin to be searched in the hash table. The computer calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target. The computer determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance. According to the above configuration, it is possible to provide a hash search method with very high efficiency in a task of classifying a determination target by searching for the presence or absence of reference data whose distance from a query is within a predetermined threshold (first distance).

[0018] A program according to an aspect of the present invention causes a computer to function as a data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to a first distance, a feature amount calculation unit that calculates a feature vector representing a determination target for determining whether or not the determination target is classified into the first class, a search area determination unit that calculates a hash value using the feature vector representing the determination target and determines a bin corresponding to the hash value as a bin to be searched in the hash table, a distance calculation unit that calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target, and a classification unit that determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance. According to the above configuration, in a task of classifying a determination target by searching for the presence or absence of reference data within a predetermined threshold value (first distance) from a query, a hash search method with very high efficiency can be provided.

Effect of the Invention

[0019] According to the present invention, in nearest neighbor search using a hash method, a hash table with high search efficiency can be constructed.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Mode for Carrying Out the Invention

[0021] Hereinafter, embodiments according to one aspect of the present invention (hereinafter also referred to as "the present embodiment") will be described with reference to the drawings. However, the embodiments described below are merely examples of the present invention in every aspect. Needless to say, various improvements and modifications can be made without departing from the scope of the present invention. That is, in implementing the present invention, a specific configuration according to the embodiment may be appropriately adopted. In the present embodiment, the data that appears is described in natural language, but more specifically, it may be specified by any of a quasi-language, command, parameter, or machine language recognizable by a computer, but is not limited thereto.

[0022] §1 Application Example The present invention can be applied to, for example, an inspection apparatus that classifies a product (object to be determined) into a non-defective product (first class) and a defective product using an image of the product. Specifically, a feature vector is extracted from the image data obtained by imaging the product to be determined using an existing image processing technique, and the distance between the point (query) represented by the feature vector and the points (non-defective product data) represented by the respective feature vectors registered in the non-defective product model is calculated. If there is any whose distance is less than or equal to a predetermined threshold value (first distance), it is determined as a non-defective product; otherwise, it is determined as a defective product. A feature vector is a representation of a plurality of feature amounts obtained from image data as a one-dimensional matrix. The non-defective product model is a set of feature vectors representing non-defective products, and is, for example, a set of feature vectors obtained from images of a plurality of non-defective products. As the image processing technique, SIFT (Scale-Invariant Feature Transform), HOG (Histograms of Oriented Gradients), SURF (Speeded Up Robust Features), LBP (Local Binary Pattern), etc. can be used, and in addition, a learned CNN (convolutional neural network) etc. may be used. In the present invention, in order to efficiently search for non-defective product data in the non-defective product data whose distance from the query is less than or equal to a predetermined threshold value, nearest neighbor search using a hash method is utilized.

[0023] FIG. 1(A) is a diagram illustrating the distribution of points represented by a plurality of feature vectors constituting a good product model. FIG. 1(B) is a diagram showing an example of a hash table in which the points represented by the feature vectors of FIG. 1(A) are hashed by a hash function and stored in corresponding bins. Although the feature vectors are generally high-dimensional vectors (F = {f1, f2, f3, …, fN}), as shown in FIG. 1(B), the hash table is constructed in two dimensions, and the point P (good product data) represented by each feature vector included in the good product model is represented as a point in the two-dimensional space. The present invention efficiently searches for the nearest neighbor point of a point Q (query) represented by a feature vector extracted from an image of a product to be determined among a plurality of good product data P using the constructed hash table.

[0024] §2 Configuration Example (1. Hardware Configuration) FIG. 2 is a diagram showing an example of the hardware configuration of the information processing apparatus 1 according to the present embodiment. The information processing apparatus 1 is a computer including a processor 11, a main memory 12, an input / output interface 13, a communication interface 14, and a storage device 15. The storage device 15 is a computer-readable recording medium such as a semiconductor memory (for example, a volatile memory or a non-volatile memory, but not limited thereto) or a disk medium (for example, a magnetic recording medium or a magneto-optical recording medium, but not limited thereto). A program executed by the processor 11 is stored in the storage device 15. The program is read from the storage device 15 into the main memory 12 and interpreted and executed by the processor 11. Further, a database 2 is implemented in the storage device 15. Note that the database 2 may be implemented in an external storage device. For example, a constructed hash table is stored in the database 2.

[0025] (2. Functional Configuration) FIG. 3 is a diagram showing an example of the functional configuration of the information processing apparatus 1. As shown in FIG. 3, the information processing apparatus 1 includes a data storage unit 101, a feature amount calculation unit 102, a search area determination unit 103, a distance calculation unit 104, and a classification unit 105. The data storage unit 101, the feature amount calculation unit 102, the search area determination unit 103, the distance calculation unit 104, and the classification unit 105 are functional modules executed by the processor 11.

[0026] §3 Operation Example Next, the operation of the information processing apparatus 1 according to the present embodiment will be described. (Construction of Hash Table) First, the construction procedure of the hash table will be described using the flowchart of FIG. 4. Here, as an example, a feature vector constituting a good product model, which is obtained from an image of a product classified as a good product (first class) using an image processing technique such as CNN, is stored in the hash table. Hereinafter, the feature vector constituting the good product model, or the point represented by the feature vector, is referred to as good product data (reference data). Also, the feature vector of the object to be determined as good or bad, or the point represented by the feature vector, is referred to as a query. Here, the hash table is composed of two-dimensional axes (basis). First, the data storage unit 101 determines the axes of the hash table (step S101). The data storage unit 101 may adopt two of the elements of the feature vector as axes. For example, any two elements may be used as axes, such as using the first two elements as axes, but elements with a large variation in the values of each feature vector may also be used as axes. Specifically, the magnitude of the variation can be determined based on statistical quantities such as the variance of the values of the elements of each feature vector. By adopting elements with a large variation as axes, when storing good product data in the hash table, it is highly likely that the good product data will not be biased towards a specific bin but will be scattered and stored in more bins. Also, the two axes do not necessarily have to be elements of the feature vector, and two new axes different from the elements of the feature vector may be set. At this time, a mapping from the space of the feature vector to the two-dimensional space by the two new axes is defined, and a point on the two-dimensional space is uniquely determined corresponding to the point represented by the feature vector. Note that the mapping can be represented by a function.

[0027] After the axes are determined, the data storage unit 101 then determines the width of the bin (step S102). The width of the bin is set to the same value as the threshold T (first distance) for determining whether it is a good product or a bad product based on the distance between the query and the good product data. Note that it is not necessary to set the width of all bins in the hash table to the same value as the threshold T, and it may be set such that the width of the bin with the smallest width among each bin is the same value as the threshold T.

[0028] Once the width of the bin is determined, the data storage unit 101 stores the good product data in the hash table (step S103). Specifically, first, the hash value of each good product data is calculated. The hash value is obtained by a hash function. The hash value may be obtained, for example, as the quotient of dividing the values of two elements determined as the basis (for example, two elements with large dispersion) among the elements of the good product data by the threshold value T. The data storage unit 101 stores each good product data in the corresponding bin based on the hash value. When two new axes different from the elements of the feature vector are set, the hash value may be the quotient of dividing the values of each axis obtained by a function representing the mapping by the threshold value T. The hash value is sometimes called a hash index.

[0029] (Search for the nearest neighbor point of the query and classification of good and bad products) Next, with reference to the flowchart of FIG. 5, the procedure for searching for the nearest neighbor point of the query and classifying good and bad products will be described. First, the feature quantity calculation unit 102 extracts the feature vector of the object to be determined (step S201). The feature vector is extracted from the image of the product to be determined using existing image processing techniques in the same manner as the feature vectors constituting the good product model.

[0030] Next, the search area determination unit 103 calculates the hash value of the query (step S202). The hash value may be obtained, for example, as the quotient of dividing the values of two elements determined as the basis among the elements of the feature vector of the object to be determined by the threshold value T, in the same way as in the case of good product data. When two axes different from the elements of the feature vector are set, the quotient of dividing the values of each axis obtained by a function representing the mapping by the threshold value T may be used as the hash value.

[0031] Next, the search area determination unit 103 determines the bin corresponding to the calculated hash value as the bin to be searched (step S203).

[0032] Next, the distance calculation unit 104 determines whether there is good product data in the bin (bin B0) determined as the search target (step S204). If there is good product data (step S204: YES), the process proceeds to step S205 to identify one piece of good product data.

[0033] The distance calculation unit 104 calculates the distance d1 between the identified good product data (P1) and the query (Q) (step S206). The distance calculation is not the distance in the space of the dimension (2D) of the hash table, but the distance in the space of the dimension of the original feature vector before hashing the two points. That is, if the feature vector is originally an N-dimensional vector, the distance in the N-dimensional space is calculated.

[0034] The distance calculation unit 104 determines whether the calculated distance is less than or equal to the threshold value T (step S207). If the distance d1 is less than or equal to the threshold value T (step S207: YES), the good product data P1 is determined as the nearest neighbor point, and the process proceeds to step S208. The classification unit 105 determines that the determination target represented by the query is a good product because a point less than or equal to the threshold value T has been found among the good product data. Also, the calculated distance value is output.

[0035] In step S207, if the distance d1 exceeds the threshold value T (NO), the distance calculation unit 104 determines whether there is other good product data in the currently searched bin (B0) (step S209). If there is other good product data (YES), other good product data is identified (step S210). Further, the distance calculation unit 104 proceeds to step S206 and calculates the distance d2 between the identified other good product data (P2) and the query (Q).

[0036] Also, in step S209, if it is determined that there is no other non-defective data in the currently searched bin (B0) (NO), the process proceeds to step S211, and it is determined whether or not the search has been performed for all of the 8 bins around (in the vicinity of) the currently searched bin (B0). Note that in step S209, when it is determined that there is no other non-defective data in the currently searched bin (B0), the determination target represented by the query may be determined to be a defective product.

[0037] FIG. 6 is a diagram for explaining the search range of the nearest point by the distance calculation unit 104. As shown in FIG. 6, the bin B0 where the query Q exists and the 8 bins B1 to B8 around it are the search ranges. In the example of FIG. 6, there is no non-defective data P in the bin B0, but there is one or more non-defective data P in the bins B1 to B8, and since the distances between these non-defective data and the query Q may be equal to or less than the threshold value T, the 8 surrounding bins are also search targets. Note that in the present embodiment, the hash table is constructed two-dimensionally, but the hash table may be one-dimensional or may be constructed three-dimensionally or more. In this case, the 26 surrounding bins may be used as the search range.

[0038] In step S211, if there is a surrounding bin that has not been searched (NO), the surrounding bin to be searched is specified (step S212). After specifying the bin to be searched, the process proceeds to step S204, and it is determined whether or not there is non-defective data in the bin.

[0039] In step S211, if the search has been completed for all of the 8 surrounding bins (YES), the process proceeds to step S213, and a selection is made as to whether to end the search for the nearest point or to perform a full search. The selection may be specified by the user each time via an input device or the like, or it may be set in advance which one to select. When "end" is selected, the process proceeds to step S214. Since the classification unit 105 has not found a point equal to or less than the threshold value T among the non-defective data, the determination target represented by the query is determined to be a defective product.

[0040] In step S213, if "exhaustive search" is selected, the process proceeds to step S215. The distance calculation unit 104 calculates the distances between all the non-defective data registered in the hash table and the query. Since the classification unit 105 does not find any points within the threshold T or less among the non-defective data, the query is determined to be defective. Also, the distance to the point (nearest neighbor point) with the shortest distance to the query among the non-defective data for which the exhaustive search has been performed is output (step S216).

[0041] FIG. 7 is a diagram for explaining an example of a method for determining the order in which the surrounding eight bins are searched. In the example of FIG. 7, representative points for the surrounding eight bins are set, and the search is performed in ascending order of the distances between each representative point and the query Q. In the example of FIG. 7, the center of gravity of the non-defective data included in each bin is used as the representative point. The representative point is not limited to such a center of gravity, and for example, the center of the bin may be used as the representative point. In the example of FIG. 7, the search is performed in the order of the numbers shown in the upper right of the surrounding eight bins.

[0042] FIG. 8 is a diagram for explaining another example of a method for constructing a hash table. FIG. 8(A) shows a hash table constructed according to the procedure of the flowchart in FIG. 4. In the state of FIG. 8(A), the non-defective data P exists partially concentrated in some bins, but as shown in FIG. 8(B), by rotating the axis of the hash table by θ, the data points can be made to exist dispersed in all the bins. As a result, no matter in which bin the query exists, the number of non-defective data within the same bin is leveled to some extent, so that the nearest neighbor search can be efficiently performed.

[0043] FIG. 8(C) shows the conversion by rotating the coordinates of the non-defective data by θ instead of the axis of the hash table so that all the non-defective data is distributed and present in all the bins. Note that the conversion method is not limited to rotation and may be translation. That is, any conversion may be used as long as the positional relationship (distance) between each non-defective data and the width of the bin are maintained. Also, in the example of FIG. 8, the axis of the hash table or the coordinates of the non-defective data are converted so that the number of bins containing one or more non-defective data is maximized, but the conversion criterion is not limited to this. For example, the conversion may be performed so that the number of non-defective data contained in one bin is less than or equal to a predetermined number, or so that the difference in the number of non-defective data between the bin with the largest number of non-defective data and the bin with the smallest number of non-defective data is minimized.

[0044] Also, in constructing the hash table, sampling of the data may be performed without registering all the points of the non-defective data. Specifically, after constructing the hash table according to the procedure of the flowchart in FIG. 4, sampling may be performed on non-defective data that is at a distance of threshold T or less from a certain non-defective data P.

[0045] Note that in the above embodiment, the threshold T may be determined using, for example, a machine learning method or a statistical method.

[0046] As described above, according to the present embodiment, when non-defective data whose distance from the query is equal to or less than the threshold T is found, hash search is performed according to the procedure of determining that the determination target represented by the query is non-defective, and the width of the bin of the hash table is set to the threshold T. Thereby, the time until non-defective data whose distance from the query is equal to or less than the threshold T is found in hash search is shortened, and a hash search method with very high efficiency can be provided. In the above-described embodiment, when good product data with a distance from the query within the threshold T is found in the bins to be searched, the search is terminated, and the determination target represented by the query is determined to be a good product. However, it is also possible to calculate the distances from all the good product data included in the bins to be searched, identify the good product data with the shortest distance, and then compare it with the threshold T to make a determination of good or bad. As a result, it is possible to determine whether the determination target represented by the query is a good product or a defective product, and it is also possible to obtain the distance from the nearest neighbor point.

[0047] As described above, in the present invention, the threshold T used in the nearest neighbor search by the hashing method is utilized for the width of the bins when constructing the hash table. The present invention can be applied to any task that utilizes hash search, but in particular, in the task of searching whether there is nearest neighbor data within a predetermined threshold, it exhibits excellent efficiency compared to conventional hash search. For example, it is particularly useful in the task of determining whether a product is good or bad when most of the products to be determined are good products and there are many good product data in the vicinity of the query representing the product.

[0048] Although the embodiments of the present invention have been described in detail above, the above description is merely an exemplification of the present invention in every respect. Needless to say, various improvements and modifications can be made without departing from the scope of the present invention.

[0049] Note that part or all of the above-described embodiments can be described as follows in the appended claims, but are not limited thereto.

[0050] (Appended Claim 1) A data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to a first distance; A feature amount calculation unit that calculates a feature vector representing a determination target for determining whether or not it is classified into the first class; A search area determination unit that calculates a hash value using a feature vector representing the determination target and determines a bin corresponding to the hash value as a bin to be searched in the hash table; A distance calculation unit that calculates a distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target; An information processing apparatus comprising: a classification unit that determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance. (Appendix 2) The data storage unit, The information processing apparatus according to Appendix 1, which determines an element to be used as a basis based on the statistic of each element of the feature vector representing each reference data. (Appendix 3) The distance calculation unit, The information processing apparatus according to Appendix 1 or 2, wherein bins in the vicinity of the bin corresponding to the hash value are also set as search targets. (Appendix 4) The distance calculation unit, When a reference data whose distance from the query is less than or equal to the first distance is found, the calculation of the distance between the other reference data and the query is terminated, The classification unit, The information processing apparatus according to any one of Appendices 1 to 3, which determines that the determination target is classified into the first class. (Appendix 5) The classification unit, The information processing apparatus according to Appendix 4, which determines that the determination target is not classified into the first class when no reference data whose distance from the query is less than or equal to the first distance is found. (Appendix 6) The distance calculation unit, The information processing apparatus according to Appendix 3, which determines the search order of the neighboring bins based on the distances between the representative points of the respective neighboring bins and the query. (Appendix 7) The information processing apparatus according to any one of Appendices 1 to 6, having a hash table in which the coordinates of the axis or reference data are converted so that the number of bins including one or more reference data is maximized. (Appendix 8) The data storage unit The information processing apparatus according to any one of Appendices 1 to 7, which downsamples reference data whose mutual distance is equal to or less than the first distance. (Appendix 9) The distance calculation unit Based on the distance between each reference data and the query, the nearest neighbor point is determined from among the reference data. The classification unit The information processing apparatus according to any one of Appendices 1 to 8, which classifies the determination target by comparing the distance between the query and the nearest neighbor point with the first distance. (Appendix 10) A data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the minimum width of a bin is set to the first distance, A feature amount calculation unit that calculates a feature vector representing a determination target for determining whether or not it is classified into the first class, A search area determination unit that calculates a hash value using the feature vector representing the determination target and determines the bin corresponding to the hash value as the bin to be searched in the hash table, A distance calculation unit that calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target, An information processing apparatus including: a classification unit that determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance. (Appendix 11) An inspection apparatus that determines whether or not the determination target is classified as a non-defective product using the information processing apparatus according to claim 1, The query is a point represented by a feature vector calculated from image data obtained by imaging the determination target. The reference data is represented by a feature vector calculated from image data of an object classified as a non-defective product, The classification unit, An inspection apparatus that classifies the determination target as a non-defective product when there is reference data whose distance from the query is smaller than the first distance. (Appendix 12) An information processing method executed by a computer, A step in which a computer stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to the first distance; A step in which a computer calculates a feature vector representing a determination target for determining whether or not the object is classified into the first class; A step in which a computer calculates a hash value using the feature vector representing the determination target and determines a bin corresponding to the hash value as a bin to be searched in the hash table; A step in which a computer calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target; A step in which a computer determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance. An information processing method including. (Appendix 13) A computer, A data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to the first distance; A feature amount calculation unit that calculates a feature vector representing a determination target for determining whether or not the object is classified into the first class; A search area determination unit that calculates a hash value using the feature vector representing the determination target and determines a bin corresponding to the hash value as a bin to be searched in the hash table; A distance calculation unit that calculates the distance between each reference data stored in the bin to be searched and a query represented by a feature vector representing the object to be determined; A program that functions as a classification unit that determines whether or not the object to be determined is classified into the first class by comparing the distance between each reference data and the query with the first distance.

Explanation of Signs

[0051] 1... Information processing apparatus, 2... Database, 11... Processor, 12... Main memory, 13... Input / output interface, 14... Communication interface, 15... Storage device, 101... Data storage unit, 102... Feature amount calculation unit, 103... Search area determination unit, 104... Distance calculation unit, 105... Classification unit

Claims

1. A data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to a first distance; A feature quantity calculation unit that calculates a feature vector representing a determination target for determining whether or not the determination target is classified into the first class; A search area determination unit that calculates a hash value using the feature vector representing the determination target and determines a bin corresponding to the hash value as a bin to be searched in the hash table; A distance calculation unit that calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the determination target; An information processing apparatus comprising: a classification unit that determines whether or not the determination target is classified into the first class by comparing the distance between each reference data and the query with the first distance.

2. The data storage unit determines an element to be used as a basis based on the statistic of each element of the feature vector representing each reference data, according to the information processing apparatus of claim 1.

3. The distance calculation unit also searches bins in the vicinity of the bin corresponding to the hash value as search targets, according to the information processing apparatus of claim 1.

4. The distance calculation unit ends the calculation of the distance between other reference data and the query when a reference data whose distance from the query is equal to or less than the first distance is found, The classification unit determines that the determination target is classified into the first class, according to the information processing apparatus of claim 1.

5. The classification unit determines that the determination target is not classified into the first class when no reference data whose distance from the query is equal to or less than the first distance is found, according to the information processing apparatus of claim 4.

6. The distance calculation unit determines the search order of the neighboring bins based on the distance between the representative point of each neighboring bin and the query, according to the information processing apparatus of claim 3.

7. having a hash table in which the coordinates of the axis or reference data are transformed so that the number of bins including one or more reference data is maximized, according to the information processing apparatus of claim 1.

8. The data storage unit performs thinning out of reference data whose mutual distance is equal to or less than the first distance, according to the information processing apparatus of claim 1.

9. The distance calculation unit determines the nearest neighbor point from among the reference data based on the distance between each reference data and the query, The classification unit The information processing apparatus according to claim 1, wherein classification of the object to be determined is performed by comparing the distance between the query and the nearest neighbor point with the first distance.

10. A data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the minimum width of a bin is set to the first distance, A feature amount calculation unit that calculates a feature vector representing an object to be determined whether it is classified into the first class, A search area determination unit that calculates a hash value using the feature vector representing the object to be determined and determines a bin corresponding to the hash value as a bin to be searched in the hash table, A distance calculation unit that calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the object to be determined, An information processing apparatus comprising: a classification unit that determines whether the object to be determined is classified into the first class by comparing the distance between each reference data and the query with the first distance.

11. An inspection apparatus that determines whether the object to be determined is classified as a non-defective product by using the information processing apparatus according to claim 1, wherein the query is a point represented by a feature vector calculated from image data obtained by imaging the object to be determined, the reference data is a point represented by a feature vector calculated from image data of an object classified as a non-defective product, the classification unit, An inspection apparatus that classifies the object to be determined as a non-defective product when there is reference data whose distance from the query is smaller than the first distance.

12. An information processing method executed by a computer, a step in which the computer stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to the first distance, a step in which the computer calculates a feature vector representing an object to be determined whether it is classified into the first class, a step in which the computer calculates a hash value using the feature vector representing the object to be determined and determines a bin corresponding to the hash value as a bin to be searched in the hash table, a step in which the computer calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the object to be determined, A step in which a computer determines whether or not the object to be determined is classified into the first class by comparing the distance between each reference data and the query with the first distance, and an information processing method including the same.

13. A computer, A data storage unit that stores reference data represented by a feature vector representing an object classified into a first class in a hash table in which the width of each bin is set to a first distance, A feature amount calculation unit that calculates a feature vector representing an object to be determined whether or not it is classified into the first class, A search area determination unit that calculates a hash value using the feature vector representing the object to be determined and determines a bin corresponding to the hash value as a bin to be searched in the hash table, A distance calculation unit that calculates the distance between each reference data stored in the bin to be searched and a query represented by the feature vector representing the object to be determined, A program that functions as a classification unit that determines whether or not the object to be determined is classified into the first class by comparing the distance between each reference data and the query with the first distance.

Citation Information

Patent Citations

  • Approximate nearest neighbor search device, approximate nearest neighbor search method, and program

    WO2013129580A1