Filtering search method

By clustering and partial division in the data set according to the similarity and distribution density of data points, combined with coarse and fine search methods, the problem of reduced search accuracy in high-dimensional data sets is solved, and efficient and low-cost search results are achieved.

CN119988698APending Publication Date: 2025-05-13MACRONIX INTERNATIONAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410675089.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-06
Filing Date
2024-05-29
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

Existing cluster filtering mechanisms are difficult to achieve uniform distribution when processing high-dimensional data sets, resulting in a decrease in search accuracy, especially in datasets where data points have significant differences in distribution density.

Method used

By distinguishing the data set into multiple clusters according to the similarity of the data points, and distinguishing the cluster into the inner and outer parts according to the distribution density of the data points, a combination of rough search and fine search is used to gradually filter out the data points close to the target point.

Benefits of technology

It improves the search efficiency and accuracy in data sets with significant differences in distribution density, reduces the search cost, and is suitable for efficient search of large-scale data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988698A_ABST
    Figure CN119988698A_ABST
Patent Text Reader

Abstract

The invention provides a filtering search method, which is used for searching in a data set, and the data set comprises a plurality of data points. The filtering search method comprises the following steps. The data set is divided into a plurality of clusters according to the similarity of the data points. Dividing each cluster area into an inner part and a peripheral part according to the distribution density of the data points; performing rough search on all the inner surrounding parts to filter out a first candidate number of inner surrounding parts; and performing fine search on the inner part of the first candidate number to search a second candidate number of data points. And obtaining a search result according to the second candidate number of data points, wherein the second candidate number of data points are close to the target point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a search method, and in particular to a filtering search method for searching according to the distribution density of a data set. Background Art

[0002] The amount of data in artificial intelligence and big data is increasing rapidly. When searching for large data sets, it will require a high search cost. Cluster filtering can be performed based on the distribution characteristics of the data set in an attempt to reduce the search cost. However, the existing cluster filtering mechanism is limited by the structure of the data set. When the data points in the data set have high-dimensional vectors, it is difficult to make each cluster have evenly distributed data points.

[0003] When the number of data points in different clusters is significantly different, the accuracy of the search will be greatly degraded. For example, when the coverage of some clusters is wide, the distance between the representative point in the cluster and other data points will increase, which will reduce the accuracy of the search. In addition, since the distribution density of data points in each cluster is significantly different, it is difficult to achieve a balance between the cluster range and the number of data points.

[0004] In view of the above issues, an improved filtering search method is needed, which can effectively search data sets with different distribution densities and has a lower search cost. Summary of the invention

[0005] According to one aspect of the present invention, a filtering search method is provided for searching in a data set, the data set including a plurality of data points. The filtering search method comprises the following steps. According to the similarity of the data points, the data set is divided into a plurality of clusters. According to the distribution density of the data points, each cluster is divided into an inner part and a peripheral part. A rough search is performed on all the inner parts to filter out the inner parts of the first candidate number. A fine search is performed on the inner parts of the first candidate number to search out the data points of the second candidate number. A search result is obtained based on the data points of the second candidate number, and some data points of the second candidate number are close to the target point. A rough search is selectively performed on the peripheral part to filter out the peripheral part of the third candidate number. A fine search is performed on the peripheral part of the third candidate number to search out the data points of the fourth candidate number. A search result is obtained based on the data points of the second candidate number and the fourth candidate number.

[0006] Other aspects and advantages of the present invention will become apparent from a review of the following drawings, detailed description, and claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0007] Figure 1A FIG. 4 is a schematic diagram of a data set according to an embodiment of the present invention.

[0008] Figure 1B for Figure 1A Schematic diagram of a dataset divided into multiple clusters.

[0009] Figure 1C for Figure 1A Schematic diagram of clusters, representative points and data points after distinguishing the dataset.

[0010] Figure 1D A schematic diagram of the distance between the representative point and the target point.

[0011] Figure 1E Illustration of filtering out multiple clusters for a coarse search.

[0012] Figure 1F A schematic diagram of searching for one or more data points closest to a target point for a refined search.

[0013] Figure 2 FIG. 1 is a flow chart of a filtering and searching method according to a first embodiment of the present invention.

[0014] Figure 3 FIG. 4 is a flow chart of a filtering and searching method according to a second embodiment of the present invention.

[0015] Figure 4A Schematic diagram of an example of dividing a data set into inner and outer parts

[0016] Figure 4B A schematic diagram showing another example of dividing a data set into an inner part and a outer part.

[0017] Figure 4C Schematic diagram of filtering out multiple inliers for a coarse search.

[0018] Figure 4D Schematic diagram of searching for a candidate number of data points from the inner part for a refined search.

[0019] Figure 4E Schematic diagram of a coarse search for the periphery.

[0020] Figure 4F Schematic diagram of filtering out multiple peripheral parts for a coarse search.

[0021] Description of reference numerals:

[0022] DS0: Dataset

[0023] d(1)~d(9): data points

[0024] d(1-1)~d(1-5): data points

[0025] d(2-1)~d(2-11): data points

[0026] d(5-1)~d(5-4): data points

[0027] d(6-1)~d(6-17): data points

[0028] c1~c7: cluster

[0029] c1'~c7': inner part

[0030] o0, o1~o7: peripheral part

[0031] r(1)~r(7): representative points

[0032] tg0: target point

[0033] dst1~dst7: distance

[0034] dst1'~dst4': distance

[0035] S200~S204: Steps

[0036] S300~S310: Steps DETAILED DESCRIPTION

[0037] The technical terms in this specification refer to the customary terms in the technical field. If some terms are explained or defined in this specification, the interpretation of these terms shall be based on the explanation or definition in this specification. Each embodiment of the present invention has one or more technical features. Under the premise of possible implementation, those skilled in the art may selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.

[0038] Figure 1A FIG. 1 is a schematic diagram of a data set DS0 according to an embodiment of the present invention. Figure 1A As shown, the data set DS0 includes multiple data points, for example: data point d(1) to data point d(9) and other data points. These data points of the data set DS0 may be unevenly distributed, that is, the distribution density of these data points is different. For example, the distribution density of data points d(3), d(4) and d(5) is relatively small, and the distribution density of data points d(8) and d(9) is relatively large. The filtering search method of the present invention can be applied to the data set DS0, and one or more data points that meet the predetermined conditions are searched out from these data points of the data set DS0 according to the filtering search method. For example, one or more data points closest to the target point tg0 are searched out from these data points of the data set DS0.

[0039] Figure 2 FIG. 1 is a flow chart of a filtering and searching method according to a first embodiment of the present invention. Figure 2As shown, firstly, step S200 is performed: the data set DS0 is divided into a plurality of clusters. Figure 1C , which is Figure 1A Schematic diagram of dividing a data set DS0 into multiple clusters. For example, the data set DS0 is divided into clusters c1 to c7. The data set DS0 is divided based on the similarity of the data points in the data set DS0. For example, the data points d(1-1) to d(1-4) with high similarity belong to the same cluster c1, and the data points d(5-1) to d(5-4) with high similarity belong to the same cluster c5. In addition, the number and distribution density of data points covered by each of the clusters c1 to c7 may be different. For example, cluster c6 covers a large number of data points, and cluster c4 covers a small number of data points.

[0040] Furthermore, step S200 further includes: selecting or specifying a representative point in each of the distinguished clusters c1-c7, for example: cluster c1 has a representative point r(1), cluster c2 has a representative point r(2), cluster c3 has a representative point r(3), and so on. These representative points r(1)-r(7) can be selected from the existing data points in clusters c1-c7. Alternatively, in addition to the existing data points, virtual points are added as representative points (i.e., the virtual points do not overlap with the existing data points). Figure 1B The differentiated clusters c1 to c7 shown form a Vornonoi Diagram. Figure 1C , which is Figure 1A Schematic diagram of clusters c1~c7, representative points r(1)~r(7) and data points after the data set DS0 is distinguished. The distinguished cluster c1 includes data points d(1-1)~d(1-4), which corresponds to Figure 1A There are 4 existing data points d(1) to d(4). Cluster c1 has a representative point r(1), which is a virtual point other than the existing data points d(1-1) to d(1-4) in cluster c1. That is, the representative point r(1) is added in addition to the data points d(1-1) to d(1-4), and the representative point r(1) does not overlap with the data points d(1-1) to d(1-4).

[0041] Similarly, cluster c5 includes data points d(5-1)~d(5-4), which correspond to Figure 1AThere are 4 existing data points d(5)~d(8). Cluster c5 has a representative point r(5), which is a virtual point other than the existing data points d(5-1)~d(5-4). The representative point r(5) does not overlap with the data points d(5-1)~d(5-4). Similarly, the representative points r(2), r(3), r(4), r(6) and r(7) of the other clusters c2, c3, c4, c6 and c7 are all virtual points other than the existing data points.

[0042] Next, step S202 is executed: a coarse search is performed based on the distinguished clusters c1-c7. The coarse search can also be called the "filter phase", which is a preliminary search based on the representative points r(1)-r(7) of the clusters c1-c7. Step S202 also includes: calculating the distance between each representative point r(1)-r(7) and the target point tg0. See also Figure 1D , which is a schematic diagram of the distance between the representative points r(1)~r(7) and the target point tg0. The representative point r(1) of cluster c1 has a distance dst1 from the target point tg0, the representative point r(2) of cluster c2 has a distance dst2 from the target point tg0, and so on. The representative points r(3), r(4), r(5), r(6) and r(7) of other clusters c3, c4, c5, c6 and c7 have distances dst3, dst4, dst5, dst6 and dst7 from the target point tg0 respectively. Step S202 can perform a rough search based on the above distances dst1~dst7. See also Figure 1E , which is a schematic diagram of a rough search filtering out multiple clusters. Distance dst1, distance dst2, distance dst5, and distance dst6 are shorter, so the rough search filters out clusters c1, c2, c5, and c6 with a candidate number m, where the candidate number m is equal to "4".

[0043] Next, step S204 is executed: a fine search is performed based on the filtered clusters c1, c2, c5 and c6. In the present embodiment, the fine search is performed on all the data points in the filtered clusters c1, c2, c5 and c6 to search for one or more data points closest to the target point tg0. For example, the distances between all the data points d(1-1)~d(1-4) in the filtered cluster c1 and the target point tg0 are calculated to search for one or more data points in the cluster c1 that are closest to the target point tg0. Similarly, the distances between all the data points d(2-1)~d(2-10) in the filtered cluster c2 and the target point tg0 are calculated to search for the closest one or more data points. Similarly, the distances between all the data points d(5-1)~d(5-4) in the filtered cluster c5 and the target point tg0 are calculated to search for the closest one or more data points. According to the distances between all the data points d(6-1)-d(6-17) in the filtered cluster c6 and the target point tg0, one or more data points closest to the target point tg0 are searched.

[0044] See also Figure 1F , which is a schematic diagram of fine search to find one or more data points closest to the target point tg0. The distance dst1' between the data point d(1-2) and the target point tg0, the distance dst2' between the data point d(1-4) and the target point tg0 in the filtered cluster c1, and the distance dst4' between the data point d(2-1) and the target point tg0, and the distance dst3' between the data point d(2-2) and the target point tg0 in the filtered cluster c2 are small, so the fine search finds the four data points d(1-2), d(1-4), d(2-1) and d(2-2) closest to the target point tg0. In other words, a candidate number k of data points are searched out through fine search, where the candidate number k is equal to "4".

[0045] Figure 3 FIG. 1 is a flow chart of a filtering and searching method according to a second embodiment of the present invention. Figure 3 As shown, firstly, step S300 is performed: the data set DS0 is divided into an inlier part and an outlier part according to the distribution density of the data points. Figure 4A , which is a schematic diagram of an example of dividing the data set DS0 into an inner part and a peripheral part, and is compared with the data set DS0 in Figures 1B and 1C. Figure 1B The data set DS0 has not yet been divided into inner and outer parts, but according to Figure 1C The distribution density of data points d(1-1), d(1-2), d(1-3) and d(1-4) of cluster c1 can be divided into cluster c1 Figure 4AThe inner part c1' and the outer part o1. The distribution density of the inner part c1' is relatively large, and the inner part c1' covers all the data points d(1-1), d(1-2), d(1-3) and d(1-4) in the cluster c1). On the other hand, the distribution density of the outer part o1 is relatively small, and the outer part o1 does not cover any data points. Figure 4A An example is to differentiate only cluster c1 into Figure 4A The inner part c1' and the outer part o1 of the cluster are changed, and the other clusters c2~c7 keep the original range.

[0046] See also Figure 4B , which is a schematic diagram of another example of dividing the data set DS0 into an inner part and a peripheral part. Each of the clusters c1 to c7 is divided into an inner part. For example, the clusters c1 to c7 have inner parts c1' to c7'. The inner parts c1' to c7' can also be called "constrained clusters".

[0047] The parts other than the inner parts c1'~c7' are collectively referred to as the outer parts o0. The inner parts c1'~c7' have a larger distribution density, which may cover all or most of the data points. The outer parts o0 have a smaller distribution density, which may cover a small number of data points (or no data points).

[0048] In addition, step S300 further includes: performing a rough search on the inner part c1'-c7'. Figure 4C , which is a schematic diagram of filtering out multiple inner parts by rough search. The four inner parts c1', c2', c5' and c6' closest to the target point tg0 are filtered out from the inner parts c1'~c7' by rough search. That is, all the inner parts c1'~c7' are narrowed down to the candidate number m1 equal to "4" by rough search to filter out the four inner parts c1', c2', c5' and c6'. In this embodiment, the four inner parts c1', c2', c5' and c6' closest to the target point tg0 are filtered out based on the distance between the representative points r(1)~r(7) of the inner parts c1'~c7' and the target point tg0.

[0049] Next, step S302 is executed: a detailed search is performed based on the filtered four inner parts c1', c2', c5' and c6' to search for one or more data points closest to the target point tg0 from all the data points in the inner parts c1', c2', c5' and c6'. Figure 4D As shown, through fine search, candidate number k1 of data points d(1-2), d(1-3), d(2-1), d(2-2) and d(2-3) are searched from the inner parts c1', c2', c5' and c6', where the candidate number k1 is equal to "5".

[0050] Next, step S304 is executed: the shortest distance d_min between the representative points r(1)-r(7) of the inner part c1'-c7' and the target point tg0 is selected. Figure 4D As shown, the representative point r(1) of the inner part c1' is closest to the target point tg0, and the distance dst1 between the representative point r(1) and the target point tg0 is the shortest, and the distance dst1 is the short distance d_min. In addition, it is determined whether the shortest distance d_min is less than or equal to the predetermined distance d_th. If the determination result is no, step S306 is executed: a rough search is performed on the outer part o0.

[0051] The parts other than the inner parts c1'-c7' are collectively referred to as the outer part o0. Therefore, in step S306, the outer part o0 is further divided into a plurality of outer parts. Figure 4E , which is a schematic diagram of a rough search of the peripheral part o0. The peripheral part o0 is further divided into multiple clusters to form 7 peripheral parts o1~o7. That is, each of the peripheral parts o1~o7 is a cluster after division. The divided peripheral parts o1~o7 can correspond one-to-one to the inner parts c1'~c7', or the peripheral parts o1~o7 have no direct correspondence with the inner parts c1'~c7'.

[0052] Furthermore, step S306 further includes: performing a rough search on the differentiated peripheral parts o1 to o7 to filter out one or more parts closest to the target point tg0. Figure 4F , which is a schematic diagram of filtering out a plurality of peripheral parts through rough search. Through rough search, peripheral parts o1, o2, o5 and o6 of candidate number m2 are filtered out, wherein the candidate number m2 is equal to 4.

[0053] Then, step S308 is executed: a fine search is performed on the filtered peripheral parts o1, o2, o5 and o6, and one or more data points closest to the target point tg0 are searched from all the data points covered by the peripheral parts o1, o2, o5 and o6. Figure 4F , the data points d(1-5) covered by the outer part o1 and the data points d(2-11) covered by the outer part o2 are closest to the target point tg0. Therefore, through fine search, the candidate number k2 of data points d(1-5) and d(2-11) are selected, where the candidate number k2 is equal to 2.

[0054] Then, step S310 is executed: the candidate number k1 data points d(1-2), d(1-3), d(2-1), d(2-2) and d(2-3) obtained in step S302 and the candidate number k2 data points d(1-5) and d(2-11) obtained in step S308 are analyzed. For example, the distances between the candidate number (k1+k2) data points d(1-2), d(1-3), d(2-1), d(2-2), d(2-3), d(1-5) and d(2-11) and the target point tg0 are compared. From the above candidate number (k1+k2) data points, the candidate number k data points are selected as the final search results. For example, the candidate number k data points d(1-2), d(1-3) and d(2-2) are selected as the final search results, where the candidate number k is equal to 3.

[0055] On the other hand, if the judgment result of step S304 is yes (i.e., the shortest distance d_min is less than or equal to the predetermined distance d_th), step S310 is directly executed: the data points d(1-2), d(1-3), d(2-1), d(2-2) and d(2-3) of the candidate number k1 obtained in step S302 are analyzed, and the distances between the above data points and the target point tg0 are compared to select the data points of the candidate number k as the final search results. For example, data points d(1-2) and d(2-1) are selected as the final search results, where the candidate number k is equal to 2. In other words, if the judgment result of step S304 is yes, the processing of the outer part o0 is skipped, and only the inner part c1'~c7' and the data points covered by it are filtered and searched.

[0056] Figure 3 In the filtering search method of the second embodiment shown, the candidate numbers ml, m2, k1 and k2 can be determined based on the shortest distance d_min between the representative points r(1)~r(7) of the inner part c1'~c7' and the target point tg0. The candidate numbers ml, m2, k1 and k2 can also be determined by referring to the outlier rate O_R and the utilization rate U_R.

[0057] More specifically, the peripheral ratio O_R is defined as the ratio of the number of data points covered by the peripheral part o1~o7 to the number of all data points in the data set DS0. The usage rate U_R is defined as the probability of actually utilizing the peripheral part o1~o7 when executing the filtering search method (i.e., the probability that the basic truth falls into the peripheral part o1~o7). The usage rate U_R may be positively correlated with the peripheral ratio O_R. As shown in Table 1, the so-called "BigANN" type data set DS0 is distinguished to obtain the inner and outer parts of multiple clusters. When the peripheral ratio O_R of the peripheral part of BigANN is 8.5%, the usage rate U_R of the peripheral part is 3.2%. When the peripheral ratio O_R of the peripheral part of BigANN is increased to 15%, the usage rate U_R of the peripheral part is correspondingly increased to 7.4%. Similarly, as shown in Table 2, the so-called "DEEP" type data set DS0 is distinguished to obtain the inner and outer parts of multiple clusters. When the peripheral ratio O_R of the peripheral part of DEEP is 8.5%, the utilization rate U_R of the peripheral part is 3.492%. When the peripheral ratio O_R of the peripheral part of DEEP increases to 15%, the utilization rate U_R of the peripheral part increases to 7.043% accordingly.

[0058] According to the data in Table 1 and Table 2, the usage rate U_R of the peripheral part is obviously lower than the peripheral ratio O_R. Therefore, the filtering search method of the present invention can search the peripheral part based on a lower search cost, which can effectively reduce the overall search cost.

[0059] Table 1 (BigANN)

[0060]

[0061] Table 2 (Deep)

[0062]

[0063] As mentioned above, when the usage rate U_R of the peripheral part is extremely low, the peripheral part can be processed based on the lower search cost. In one example, the peripheral part can be skipped without processing (ie, the peripheral part is not searched). Figure 3 In step S304, if it is determined that the shortest distance d_min is less than or equal to the predetermined distance d_th, step S306 (i.e., the rough search of the peripheral part) and step S308 (i.e., the fine search of the peripheral part) are skipped, and step S310 is directly executed. In other words, the number of candidates m2 for the rough search of the peripheral part is set to "0", as shown in formula (1). Skipping the peripheral part without processing can greatly reduce the search cost, but may affect the accuracy of the search.

[0064] m2=0 (1)

[0065] In another example, if the shortest distance d_min is determined to be less than or equal to the predetermined distance d_th, the number of candidates m2 for the rough search of the peripheral part may be set to be much smaller than the number of candidates m1 for the rough search of the inner part, as shown in formula (2).

[0066] (2)

[0067] In yet another example, the search cost of the peripheral portion may be determined based on the difficulty of the filtered search. Figure 1D , the difficulty level D_L is defined according to the distances dst1~dst7 between the representative points r(1)~r(7) of clusters c1~c7 and the target point tg0. The difficulty level D_L is a value that is positively correlated with the distances dst1~dst7. When the value of the distances dst1~dst7 is small (the shorter the distance), the value of the difficulty level D_L is small (indicating that the difficulty of the filtering search is lower). Conversely, when the value of the distances dst1~dst7 is large, the value of the difficulty level D_L is large. As shown in Table 3, when the peripheral ratio O_R of the peripheral part is 8.5%, different values ​​of the difficulty level D_L correspond to different values ​​of the usage rate U_R. For example, when the difficulty level D_L is 0~10%, the usage rate U_R of the peripheral part of the BigANN data set DS0 is 0.00002. When the difficulty level D_L is 70%~80%, the usage rate U_R of the peripheral part of the BigANN dataset DS0 is 0.032.

[0068] Table 3

[0069]

[0070] Furthermore, the candidate numbers m1, m2, k1 and k2 are adjusted according to the value of the difficulty level D_L. For example, if the difficulty level D_L has a larger value, indicating that the difficulty of the filtering search is higher, the candidate number m2 can be set to a larger value.

[0071] In summary, the filtering search method of the present invention can be adapted to large-scale data sets and improve the search efficiency. The filtering search method of the present invention has many advantages and effects, for example: based on the existing cluster structure, some clusters are further divided into inner and outer parts. Selective searches for the inner and outer parts can improve the search efficiency. In addition, the search cost of the outer part is reduced by adjusting the candidate numbers m1, m2, k1 and k2. Furthermore, no additional search cost is required during the filtering search process.

[0072] Although the present invention has been disclosed in detail with preferred embodiments and examples, it is to be understood that these examples are intended to be illustrative rather than restrictive. It is contemplated that a variety of modifications and combinations may be conceived by a person skilled in the art, and the various modifications and combinations thereof fall within the spirit of the present invention and the scope of the appended claims.

Claims

1. A filtering search method for searching in a data set, the data set comprising a plurality of data points, the filtering search method comprising: Based on the similarity of these data points, the data set is divided into multiple clusters; According to the distribution density of these data points, each cluster is divided into an inner part and an outer part; Performing a rough search on all of the inner parts to filter out a first candidate number of the inner parts; Performing a fine search on the inner portions of the first candidate quantity to search for a second candidate quantity of these data points; as well as Obtain a search result based on the second candidate number of these data points; The second candidate number of these data points are close to a target point.

2. The filtering search method according to claim 1, wherein the step of performing the rough search on all of the inner parts comprises: specifying a representative point in each of the clusters; Calculating a first distance between each of the representative points and the target point; as well as The inner parts of the first candidate number are selected according to the first distance.

3. The filtering search method according to claim 2, wherein the step of performing the fine search on the inner parts of the first candidate quantity comprises: Calculating a second distance between each of the data points covered by the inner portions of the first candidate number and the target point; as well as The data points of the second candidate number are selected from the coverage of the inner parts of the first candidate number according to the second distance. 4 . The filtering search method according to claim 2 , wherein each of the representative points is one of the data points or a virtual point other than the data points.

5. The filtering and searching method according to claim 2, further comprising: Obtaining a difficulty level of each cluster according to each first distance; as well as The first candidate number and the second candidate number are set according to the difficulty levels. 6 . The filtering search method according to claim 5 , wherein the peripheral parts have a utilization rate and a peripheral ratio, and the difficulty levels are related to the utilization rate and the peripheral ratio.

7. The filtering and searching method according to claim 2, further comprising: A rough search is selectively performed on the peripheral portions to filter out a third candidate number of the peripheral portions.

8. The filtering and searching method according to claim 7, further comprising: performing a fine search on the peripheral portions of the third candidate number to search for a fourth candidate number of these data points; as well as The search result is obtained according to these data points of the second candidate number and the fourth candidate number.

9. The filtering search method according to claim 7, wherein the first distances include a shortest distance, when the shortest distance is less than or equal to a predetermined distance, the third candidate number is set to "0", and the search result is obtained based on the data points of the second candidate number. 10 . The filtering search method according to claim 9 , wherein when the shortest distance is less than or equal to the predetermined distance, the third candidate number is set to be smaller than the first candidate number.