Information processing device, information processing method, and information processing program
The information processing device employs a group connection graph with a search query acquisition and processing unit to address the limitations of conventional search technologies, enabling flexible and efficient retrieval of objects within classified groups.
Patent Information
- Application Number
- JP2021200957
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-10-01
- Estimated Expiration
- 2041-12-10
AI Technical Summary
Conventional information search technologies struggle with flexible search processing using graphs, particularly when dealing with groups of classified objects, limiting their ability to support diverse search queries.
An information processing device that utilizes a group connection graph with a search query acquisition unit and a search processing unit to perform flexible search processing by employing a first and second search range, allowing for efficient extraction of nearby objects using a graph structure.
Enables flexible and efficient search processing by utilizing a group connection graph to handle diverse search queries, enhancing the capability to find relevant objects within classified groups.
Smart Images

Figure 0007747506000002 
Figure 0007747506000003 
Figure 0007747506000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device, an information processing method, and an information processing program. [Background technology]
[0002] Conventionally, various techniques for searching (retrieving) information have been proposed. For example, a technique has been proposed in which a graph is generated in which nodes corresponding to the search target are connected by edges, and the generated graph is used for searching. Such a technique is also used for image retrieval, for example. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6293335 [Patent Document 2] Patent No. 6300982 [Patent Document 3] Patent No. 6311000 [Non-patent literature]
[0004] [Non-Patent Document 1] Masajiro Iwasaki, "Neighborhood Search Using Approximate k-Nearest Neighbor Graphs with Tree-Structured Indexes," Transactions of Information Processing Society of Japan, February 2011, Vol. 52, No. 2, pp.817-828. Summary of the Invention [Problem to be solved by the invention]
[0005] However, there is room for improvement in the above-mentioned conventional technology. For example, the above-mentioned conventional technology performs search processing using a search range from one perspective, and it is difficult to support search processing using a graph related to groups into which objects are classified, and there is room for improvement in terms of flexible search processing. Therefore, there is a demand for flexible search processing using graphs.
[0006] The present application has been made in view of the above, and aims to provide an information processing device, an information processing method, and an information processing program that perform flexible search processing using graphs. [Means for solving the problem]
[0007] The information processing device according to the present application is characterized by comprising: a group connection graph in which a plurality of groups into which a plurality of objects to be the subject of a data search are classified are connected by edges; an acquisition unit that acquires a search query for the plurality of objects; and a search processing unit that executes a search process to extract nearby objects corresponding to the search query from the plurality of objects by searching the group connection graph using a first search range that indicates a range corresponding to a search of the plurality of groups and a second search range that indicates a range corresponding to a search of the plurality of objects. [Effects of the Invention]
[0008] According to one aspect of the embodiment, it is possible to provide an effect that flexible search processing can be performed using graphs. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of information processing according to the first embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of data according to the first embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of data according to the first embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of the configuration of an information processing system according to the first embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of the configuration of the information processing device according to the first embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of an object information storage unit according to the first embodiment. [Figure 7]FIG. 7 is a diagram illustrating an example of a graph information storage unit according to the first embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of a quantization information storage unit according to the first embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of a codebook information storage unit according to the first embodiment. [Figure 10] FIG. 10 is a flowchart illustrating an example of information processing according to the first embodiment. [Figure 11] FIG. 11 is a flowchart showing an example of the search process according to the first embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of information processing according to the second embodiment. [Figure 13] FIG. 13 is a diagram illustrating an example of the configuration of an information processing device according to the second embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of a quantization information storage unit according to the second embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of a blob information storage unit according to the second embodiment. [Figure 16] FIG. 16 is a flowchart showing an example of the search process according to the second embodiment. [Figure 17] FIG. 17 is a diagram illustrating an example of information processing according to the modified example. [Figure 18] FIG. 18 is a diagram illustrating an example of a graph information storage unit according to a modified example. [Figure 19] FIG. 19 is a flowchart showing an example of a search process according to a modified example. [Figure 20] FIG. 20 is a diagram showing another example of information processing according to the modified example. [Figure 21] FIG. 21 is a diagram illustrating an example of a process for generating blob information. [Figure 22] FIG. 22 is a flowchart showing an example of information processing according to the modified example. [Figure 23] FIG. 23 is a diagram illustrating an example of information processing according to the third embodiment. [Figure 24]FIG. 24 is a diagram illustrating an example of the configuration of an information processing device according to the third embodiment. [Figure 25] FIG. 25 is a diagram illustrating an example of a blob connection graph information storage unit according to the third embodiment. [Figure 26] FIG. 26 is a flowchart illustrating an example of information processing according to the third embodiment. [Figure 27] FIG. 27 is a flowchart showing an example of search processing according to the third embodiment. [Figure 28] FIG. 28 is a flowchart showing an example of search processing according to the third embodiment. [Figure 29] FIG. 29 is a diagram showing an example of blob connection graph information according to a modified example. [Figure 30] FIG. 30 is a diagram illustrating an example of a blob information storage unit according to a modified example. [Figure 31] FIG. 31 is a diagram illustrating an example of a blob connection graph information storage unit according to a modified example. [Figure 32] FIG. 32 is a diagram illustrating an example of information processing according to the fourth embodiment. [Figure 33] FIG. 33 is a diagram illustrating an example of the configuration of an information processing device according to the fourth embodiment. [Figure 34] FIG. 34 is a conceptual diagram regarding quantization according to the fourth embodiment. [Figure 35] FIG. 35 is a diagram illustrating an example of an inverted index according to the fourth embodiment. [Figure 36] FIG. 36 is a diagram illustrating a specific example of an inverted index according to the fourth embodiment. [Figure 37] FIG. 37 is a flowchart showing an example of information processing according to the fourth embodiment. [Figure 38] FIG. 38 is a flowchart showing an example of information processing according to the fourth embodiment. [Figure 39] FIG. 39 is a diagram illustrating an example of information processing according to the fifth embodiment. [Figure 40] FIG. 40 is a diagram illustrating an example of information processing according to the fifth embodiment. [Figure 41] FIG. 41 is a diagram illustrating an example of the configuration of an information processing device according to the fifth embodiment. [Figure 42] FIG. 42 is a flowchart showing an example of information processing according to the fifth embodiment. [Figure 43] FIG. 43 is a flowchart showing an example of search processing according to the fifth embodiment. [Figure 44] FIG. 44 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device. DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, a detailed description will be given of an information processing device, an information processing method, and an information processing program (hereinafter referred to as an "embodiment") according to the present application, with reference to the drawings. Note that the information processing device, the information processing method, and the information processing program according to the present application are not limited to the embodiment. Furthermore, the same components in the following embodiments are denoted by the same reference numerals, and duplicated descriptions will be omitted.
[0011] (Embodiment) 1. First Embodiment [1-1. Information Processing] An example of information processing according to the first embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram illustrating an example of information processing according to the first embodiment. An information processing device 100 executes search processing using a graph index (also simply referred to as a "graph") in which multiple objects to be the target of data search are graph-structured. FIG. 1 illustrates a portion of search processing in which the information processing device 100 performs a neighborhood search using a graph in which nodes corresponding to vectors obtained by vectorizing objects to be the target of data search are connected by edges. The information processing device 100 searches for nodes near a given search query (vector) by traversing the graph, with each node of the graph being the object to be searched. Through the search processing, the information processing device 100 extracts a predetermined number of nodes (hereinafter also referred to as the "search count") designated as the number of neighboring nodes to be extracted as nodes near the search query. While the following describes an example in which image information is the target of data search, the target of data search may be various objects such as video information or audio information.
[0012] The information processing device 100 performs search processing on nodes corresponding to a huge amount of image information, such as millions to hundreds of millions, but only a portion of them (several dozens of nodes such as node N1 in FIG. 1) are shown in the drawings. For example, the information processing device 100 acquires information on a plurality of nodes (vectors) such as nodes N1, N7, N9, etc., as shown in spatial information SP1 in FIG. 1. When "node N* (* is an arbitrary numerical value)" is written in this way, it indicates that the node is identified by the node ID "N*." For example, when "node N1" is written, the node is identified by the node ID "N1." Each white circle (◯) in the spatial information SP1 indicates a node.
[0013] In the spatial information SP1 in FIG. 1, nodes related to the explanation are primarily labeled, but each of the unlabeled white circles (◯) is also a node, and many other nodes are included in addition to the nodes shown in the figure. Each node corresponds to an object (search target). For example, each of multiple local features extracted from an image may be an object. Also, for example, various data in which the distance between objects is defined may be an object.
[0014] For example, the spatial information SP1 in Fig. 1 may be a Euclidean space. For example, the spatial information SP1 corresponds to the number of dimensions of the object vector, and is a multidimensional space of 100 dimensions, 1000 dimensions, etc. Note that Fig. 1 conceptually illustrates a state in which the vector (space) is divided into four partial regions (subspaces) by Cartesian product quantization as shown in the spatial information SP1, but the Cartesian product quantization will be described later.
[0015] Dotted lines connecting white circles (◯), which are nodes in the spatial information SP1, indicate edges connecting the nodes. For simplicity's sake, the example in FIG. 1 shows a case where nodes are connected by undirected (bidirectional) edges (hereinafter simply referred to as "edges"). Note that an undirected edge here refers to an edge that allows data to be traced in both directions between the connected nodes. For example, a dotted line connects the white circle (◯) representing node N1 and the white circle (◯) representing node N4, indicating that nodes N1 and N4 are connected by an edge. That is, in the graph shown in the spatial information SP1, it is possible to trace between nodes N1 and N4 in both directions. Specifically, in the spatial information SP1, it is possible to trace from node N1 to node N4, and from node N4 to node N1.
[0016] In the example of Figure 1, for convenience of illustration, only the edges connecting the illustrated nodes are shown, but many other edges are included in addition to the illustrated edges. In this way, the spatial information SP1 in Figure 1 shows only a portion of the edges, but it is assumed to be, for example, a k-nearest neighbor graph. Note that the spatial information SP1 may be various graphs.
[0017] Furthermore, the edges of a graph are not limited to undirected edges, but may also be directed edges. In the case of a directed edge, it is possible to trace only from the node that is the reference source of the directed edge to the referenced node. For example, when two nodes are connected by two directed edges, a first edge with one being the reference source and the other being the reference destination, and a second edge with one being the reference destination and the other being the reference source, the state is the same as when the two nodes are connected by an undirected (bidirectional) edge.
[0018] The search process will now be described with reference to Fig. 1. In the example of Fig. 1, it is assumed that the information processing device 100 has already acquired a graph GR1 as shown in spatial information SP1. Note that the information processing device 100 may generate the graph GR1 by appropriately using various conventional techniques. First, prior to describing the search process, direct product quantization of vectors (space) will be described.
[0019] Fig. 1 conceptually illustrates a state in which a vector (space) is divided into four subspaces AR11 to AR14 by product quantization as shown in spatial information SP1. The subspaces AR11 to AR14 shown in Fig. 1 are shown as a conceptual diagram for explaining the distance between vectors of each node (object), and each of the subspaces AR11 to AR14 is a multidimensional space. For example, the subspace AR11 shown in Fig. 1 is shown in a two-dimensional form in order to be illustrated on a plane, but it is assumed to be a multidimensional space with, for example, 100 or 1000 dimensions.
[0020] Each of the subspaces AR11 to AR14 indicates a space corresponding to each division (subvector) of a vector divided into four. Note that, for example, a 100-dimensional vector may be divided into 100 parts. For example, the subspace AR11 indicates a dimensional space (also referred to as a "first subspace") corresponding to the first subvector (also referred to as a "first subvector") among the four subvectors (also referred to as "subvectors") obtained by dividing a vector into four parts. The subspace AR11 indicates a subspace (first subspace) corresponding to the first division position of the vector. In the case of the search query QE1, the subspace AR11 is a space corresponding to the first subquery QE1-1, which is the first subvector (first subvector).
[0021] Furthermore, subspace AR12 indicates a dimensional space (also referred to as a "second subspace") corresponding to the second-first subvector (also referred to as a "second subvector") among the four subvectors obtained by dividing a vector into four. Subspace AR12 indicates a subspace (second subspace) corresponding to the second-first division position of the vector. In the case of search query QE1, subspace AR12 is a space corresponding to second subquery QE1-2, which is the second-first subvector (second subvector).
[0022] Furthermore, subspace AR13 indicates a dimensional space (also referred to as a "third subspace") corresponding to the third partial vector (also referred to as a "third subvector") from the top out of the four partial vectors obtained by dividing a vector into four. Subspace AR13 indicates a subspace (third subspace) corresponding to the third division position from the top of the vector. In the case of search query QE1, subspace AR13 is a space corresponding to the third partial query QE1-3, which is the third partial vector (third subvector) from the top.
[0023] Furthermore, subspace AR14 indicates a dimensional space (also referred to as the "fourth subspace") corresponding to the fourth (i.e., last) partial vector (also referred to as the "fourth subvector") from the beginning of the four partial vectors obtained by dividing a vector into four. Subspace AR14 indicates the subspace (fourth subspace) corresponding to the fourth division position from the beginning of the vector. In the case of search query QE1, subspace AR14 is a space corresponding to the fourth partial query QE1-4, which is the fourth partial vector (fourth subvector) from the beginning.
[0024] 1, the subspaces AR11 to AR14 are shown with similar shapes, but the shapes of the subspaces AR11 to AR14 may be different, and the division of the areas in the subspaces AR11 to AR14 may also be different. Also, the connection relationships between nodes via edges are assumed to be common to the subspaces AR11 to AR14. That is, the graph GR1 is a graph corresponding to the spatial information SP1, and is common to the subspaces AR11 to AR14.
[0025] 1, a lookup table is generated for each divided partial vector (subvector). For example, the first subvector, the second subvector, the third subvector, and the fourth subvector are clustered, and a lookup table (codebook information) is generated for each of them.
[0026] 1, the first subvectors are clustered into nine groups corresponding to the codebooks CD11 to CD19 as shown in the subspace AR11, and the vectors corresponding to each of the codebooks CD11 to CD19 are calculated as representative vectors (centroids). For example, the first subvectors of nodes N7, N9, etc. are clustered into the group corresponding to the codebook CD11. In this case, the first subvectors of nodes N7, N9, etc. are vector-quantized into vectors of codebook CD11 using the codebook information (also referred to as "first codebook information") corresponding to the first subvectors in the distance calculation.
[0027] Furthermore, the second subvectors are clustered into a plurality of groups corresponding to the codebooks CD21 to CD24 (see FIGS. 2 and 9), and a vector corresponding to each of the codebooks CD21 to CD24 is calculated as a representative vector (centroid). In this case, the second subvectors of each node are vector-quantized into vectors of the corresponding codebooks using codebook information (also referred to as "second codebook information") corresponding to the second subvectors in distance calculation (calculation).
[0028] The third subvectors are clustered into a plurality of groups corresponding to the codebooks CD31 to CD34 (see FIGS. 2 and 9), and a vector corresponding to each of the codebooks CD31 to CD34 is calculated as a representative vector (centroid). In this case, the third subvectors of each node are vector-quantized into vectors of the corresponding codebooks using codebook information (also referred to as "third codebook information") corresponding to the third subvectors in distance calculation (calculation).
[0029] The fourth subvectors are clustered into a plurality of groups corresponding to codebooks CD41 to CD44, etc. (see FIGS. 2 and 9), and a vector corresponding to each of codebooks CD24 to CD44, etc. is calculated as a representative vector (centroid). In this case, the fourth subvectors of each node are vector-quantized into vectors of the corresponding codebooks using codebook information (also referred to as "fourth codebook information") corresponding to the fourth subvectors in distance calculation (calculation).
[0030] The codebook information, such as the process of determining the representative vector, is generated by appropriately using various techniques. Since the generation of the codebook information is a conventional technique, a detailed description thereof will be omitted. The lookup table is not limited to the above, and for example, one lookup table (codebook information) may be used for all the divided partial vectors (sub-vectors), which will be described later.
[0031] Next, a search process for the search query QE1 will be described. First, the information processing device 100 acquires the search query QE1 (step S11). For example, the information processing device 100 acquires the search query QE1 from the terminal device 10 (see FIG. 4) used by the user.
[0032] Then, the information processing device 100 divides the vector that is the search query QE1 into four. That is, the information processing device 100 divides the search query QE1 into four partial queries. In FIG. 1, the information processing device 100 divides the search query QE1 into four sub-vectors: a first partial query QE1-1 that is a first sub-vector, a second partial query QE1-2 that is a second sub-vector, a third partial query QE1-3 that is a third sub-vector, and a fourth partial query QE1-4 that is a fourth sub-vector. Specifically, the information processing device 100 defines "45, 23, 2..." as the first partial query QE1-1, "127, 34, 5..." as the second partial query QE1-2, "20, 98, 110..." as the third partial query QE1-3, and "12, 45, 4..." as the fourth partial query QE1-4.
[0033] Then, the information processing device 100 calculates the distance between each sub-vector of the search query QE1 and the vector in the codebook corresponding to that sub-vector. The information processing device 100 calculates the distance (difference) between each sub-vector of the search query QE1 and the vector in the codebook, as shown in codebook information TB1 to TB4 in Fig. 2. Fig. 2 is a diagram showing an example of data according to the first embodiment.
[0034] Codebook information TB1 indicates first codebook information corresponding to the first subvector, and the information processing device 100 calculates the distance between each vector in codebooks CD11 to CD19 corresponding to the first subvector and the first partial query QE1-1. In Fig. 2, the information processing device 100 calculates the distance between codebook CD11 and the first partial query QE1-1 as distance DS11, as shown in codebook information TB1. Similarly, the information processing device 100 calculates the distance between each of codebooks CD12 to CD14, etc., and the first partial query QE1-1 as distances DS12 to DS14, etc.
[0035] Similarly, codebook information TB2 to TB4 indicate second to fourth codebook information corresponding to the second to fourth subvectors, respectively. As shown in codebook information TB2, the information processing device 100 calculates the distances DS21 to DS24, etc. between each of the codebooks CD21 to CD24, etc. and the second partial query QE1-2. As shown in codebook information TB3, the information processing device 100 calculates the distances DS31 to DS34, etc. between each of the codebooks CD31 to CD34, etc. and the third partial query QE1-3. As shown in codebook information TB4, the information processing device 100 calculates the distances DS41 to DS44, etc. between each of the codebooks CD41 to CD44, etc. and the fourth partial query QE1-4.
[0036] Note that, although FIG. 2 shows abstractions for the sake of explanation, distances DS11 to DS44 and the like are assumed to be concrete values. The distances are values expressed as floating-point numbers, but may be compressed to 1-byte integers using scale-offset-compression (see, for example, "https: / / www.unidata.ucar.edu / blogs / developer / entry / compression_by_scaling_and_offfset") to increase speed using SIMD (Single Instruction, Multiple Data), which will be described later. The information processing device 100 may calculate the distance between each codebook and the query when the search query QE1 is acquired, or when that information becomes necessary. In the search process, the information processing device 100 uses the codebook information TB1 to TB4 as a lookup table to calculate the distance between each node and the search query QE1, i.e., the approximate distance (also referred to as the "first distance"). In this way, when searching a graph in a search process, the information processing device 100 calculates and processes the approximate distance (first distance) rather than the true distance (also called the "second distance") between each node and the search query, thereby enabling efficient search processing.
[0037] The information processing device 100 executes a search process targeting the search query QE1 (step S12). The information processing device 100 performs a search process as shown in Fig. 11 using the graph GR1 targeting the search query QE1, thereby obtaining search results for the search query QE1. The search process shown in Fig. 11 will be described in detail later. By performing the search process targeting the search query QE1, the information processing device 100 extracts nodes of the search number as nodes in the vicinity of the search query QE1.
[0038] In a search process targeting the search query QE1, the information processing device 100 selects a predetermined node from among the nodes as a node (hereinafter also referred to as the "starting node") that will be the starting point (origin) of the search of the graph GR1. For example, the information processing device 100 selects the starting node using a predetermined index such as a tree structure index. In the example of FIG. 1, the information processing device 100 selects node N7 as the starting node. Note that, for simplicity of explanation, FIG. 1 shows a case where only node N7 is selected as the starting node using a predetermined index, but the information processing device 100 may select multiple nodes as the starting node, or may select the starting node using various methods such as randomly.
[0039] Here, in the search processing targeting the search query QE1, the information processing device 100 calculates in parallel the distance between the search query and nodes connected by edges from one node (hereinafter also referred to as "connecting nodes") (step S13). In Fig. 1, as shown in the batch processing information LT1, the information processing device 100 calculates in parallel the distance between the search query QE1 and nodes N9, N12, N54, N85, etc., which are nodes connected by edges from node N7 (connecting nodes).
[0040] For example, as shown in the node information INF1, node N9 indicates that the first subvector corresponds to codebook CD12, the second subvector corresponds to codebook CD23, the third subvector corresponds to codebook CD35, and the fourth subvector corresponds to codebook CD47. Note that the size of the variable CD is determined by the size of the codebook. For example, if the codebook size is 16, the variable CD can be 4 bits, enabling significant data compression. Therefore, the information processing device 100 calculates the distance between node N9 and the search query QE1 using the distance DS12 of codebook CD12, the distance DS23 of codebook CD23, the distance DS35 of codebook CD35, and the distance DS47 of codebook CD47. For example, the information processing device 100 calculates the sum of the distance DS12 of codebook CD12, the distance DS23 of codebook CD23, the distance DS35 of codebook CD35, and the distance DS47 of codebook CD47 as the distance between node N9 and the search query QE1.
[0041] For example, as indicated in node information INF2, node N12 indicates that its first subvector corresponds to codebook CD14, its second subvector corresponds to codebook CD29, its third subvector corresponds to codebook CD31, and its fourth subvector corresponds to codebook CD45. For example, the information processing device 100 calculates the sum of the distance DS14 of codebook CD14, the distance DS29 of codebook CD29, the distance DS31 of codebook CD31, and the distance DS45 of codebook CD45 as the distance between node N12 and search query QE1. Similarly, the information processing device 100 calculates the distance between node N54 and search query QE1 using node information INF3 of node N54, and calculates the distance between node N85 and search query QE1 using node information INF4 of node N85.
[0042] For example, the information processing device 100 calculates the distance between each of nodes N9, N12, N54, and N85 and the search query QE1 in a batch by parallelizing SIMD operations. In this way, the information processing device 100 can speed up the distance calculation by parallelizing the distance calculation, thereby enabling efficient search processing.
[0043] For ease of explanation, FIG. 1 illustrates a case in which distances are calculated in parallel for four nodes, but the number of nodes to be parallelized is determined based on the specifications of the information processing device 100. For example, when the number of nodes that can be collectively processed by the information processing device 100 using SIMD (also referred to as the "batch processable unit") is "4," the same processing as in FIG. 1 is performed. However, when the number of nodes that can be collectively processed by the SIMD (batch processable unit) is "16," the information processing device 100 calculates distances in parallel for 16 nodes. When the batch processable unit is "32," the information processing device 100 calculates distances in parallel for 32 nodes. In this way, the information processing device 100 can perform efficient search processing by collectively calculating distances for the number of nodes corresponding to the number of nodes that can be collectively processed by SIMD (batch processable unit).
[0044] The information processing device 100 performs distance calculations for multiple nodes in parallel as described above, and then performs a search process using the graph GR1 as shown in Figure 11, targeting the search query QE1, thereby extracting the number of searched nodes as nodes near the search query QE1.
[0045] As described above, the information processing device 100 can perform efficient search processing by calculating the approximate distance (first distance) between the vector of each vector-quantized node and the search query and performing search processing using the first distance. Furthermore, the information processing device 100 can perform efficient search processing by simultaneously calculating the distances of as many nodes as can be parallelized. As described above, the information processing device 100 can perform high-speed search processing by repeatedly using a limited memory area (lookup table) (stored in a memory cache) by simultaneously calculating the approximate distances of a graph node to multiple connecting nodes.
[0046] Furthermore, in the example of FIG. 1 , the information processing device 100 performs search processing using vectors divided by Cartesian product quantization, thereby enabling more efficient search processing compared to performing search processing without dividing the vectors. For example, the information processing device 100 calculates the distance between short vectors using Cartesian product quantization, thereby reducing the size of the vectors to be subjected to distance calculation. Furthermore, the information processing device 100 can limit the memory space accessed by the lookup table used for distance calculation to the lookup table for the divided subvectors. The information processing device 100 can reduce the vector size while suppressing a decrease in search accuracy through Cartesian product quantization. Note that, although FIG. 1 describes processing when Cartesian product quantization is performed, this is merely an example, and the information processing device 100 may perform search processing using vectors that have not been subjected to Cartesian product quantization.
[0047] [1-1-1. Other] The above-described process is merely an example, and the information processing device 100 may perform search processing using various information and methods for efficient search processing. Each aspect of this will be described in detail below.
[0048] (Storage of edges (connection nodes)) Regarding the information on edges (connecting nodes) connected to each node, there are two possible storage modes (storage methods) for the edges (connecting nodes) at the node, for example, the first and second storage modes below, and the storage mode may be selected depending on the usage pattern.
[0049] For example, in a first storage mode, each node may be stored in association with the ID of the connecting node (object), and information on the object that has been subjected to Cartesian product quantization may be stored in a separate table. For example, the data storage modes shown in Figures 7 and 8 correspond to the first storage mode. In the first storage mode, memory usage can be reduced.
[0050] Also, for example, in the second storage mode, each node has a quantized object directly. In the second storage mode, a connection node (object) of a certain node stores quantized information in association with each node. For example, the second storage mode corresponds to a data storage mode in which the quantized information (node information INF1 to INF4 in FIG. 2) of each of nodes N9, N12, N54, and N85, which are connection nodes of node N7, is stored in association with node N7. In the second storage mode, objects are accessed sequentially during a search, which can prevent a decrease in speed.
[0051] (Lookup table) In the above example, clustering is performed for each divided sub-vector, and a lookup table (codebook information) is generated for each of them. However, this is not limiting, and the lookup table may take any form.
[0052] For example, all sub-vectors may be clustered to generate one lookup table. In the above example, all four sub-vectors, i.e., the first sub-vector, the second sub-vector, the third sub-vector, and the fourth sub-vector, may be clustered to generate one lookup table.
[0053] Alternatively, individual subvectors may be partially merged to generate a lookup table for each of a plurality of classes. For example, subvectors with similar variances may be merged to form a plurality of classes, and a lookup table may be generated for each of the plurality of classes. In the above example, for example, if the variances of two subvectors, the first subvector and the third subvector, are similar, and the variances of two subvectors, the second subvector and the fourth subvector, are similar, the first subvector and the third subvector may be merged to form a first class, and the second subvector and the fourth subvector may be merged to form a second class. In this case, two lookup tables may be generated: a lookup table (codebook information) for the first class and a lookup table (codebook information) for the second class.
[0054] (Transpose) The above-described data storage is merely an example, and the information processing device 100 may store data in various modes to enable efficient data reference, etc. For example, the information processing device 100 may store transposed data. This point will be described with reference to Fig. 3. Fig. 3 is a diagram showing an example of data according to the first embodiment.
[0055] For example, when data is held as shown in the node information INF1 to INF4 and the node information INF1 to INF4 is processed sequentially, the codebook information TB1 to TB4, which are lookup tables, are repeatedly referenced as shown in step S14 of FIG.
[0056] Therefore, as shown in step S15, the node information INF1 to INF4 is transposed to generate transposed data TR1 to TR4. For example, the information processing device 100 transposes the node information INF1 to INF4 to generate transposed data TR1 to TR4.
[0057] 3, first data TR1 is generated, which is a list of codebooks corresponding to the first subvectors of nodes N9, N12, N54, and N85 among the node information INF1 to INF4. Specifically, first data TR1 is generated, which is a list of codebook CD12 corresponding to the first subvector of node N9, codebook CD14 corresponding to the first subvector of node N12, codebook CD13 corresponding to the first subvector of node N54, and codebook CD18 corresponding to the first subvector of node N85.
[0058] Similarly, second data TR2 is generated which is a list of codebooks corresponding to the second sub-vectors of nodes N9, N12, N54, and N85 among the node information INF1 to INF4. Also, third data TR3 is generated which is a list of codebooks corresponding to the third sub-vectors of nodes N9, N12, N54, and N85 among the node information INF1 to INF4. Fourth data TR4 is generated which is a list of codebooks corresponding to the fourth sub-vectors of nodes N9, N12, N54, and N85 among the node information INF1 to INF4.
[0059] Correspondence information is generated that indicates which node corresponds to which data item in the list of each of the transposed data TR1 to TR4. In the example of Fig. 3, correspondence information is generated that indicates that, in the list of each of the transposed data TR1 to TR4, the first (initial) data corresponds to node N9, the second data corresponds to node N12, the third data corresponds to node N54, and the fourth (last) data corresponds to node N85. The information processing device 100 can identify which node corresponds to each data item in the list of each of the transposed data TR1 to TR4 by referring to the correspondence information. For example, the information processing device 100 generates the correspondence information.
[0060] The information processing device 100 performs processing using transposed data TR1 to TR4 obtained by transposing the node information INF1 to INF4. For example, when the information processing device 100 refers to a lookup table using the transposed data TR1, it refers only to the codebook information TB1, i.e., it refers to only one piece of codebook information TB1. Similarly, when the information processing device 100 refers to a lookup table using the transposed data TR2, it refers only to the codebook information TB2, i.e., it refers to only one piece of codebook information TB2. Similarly, for the transposed data TR3 and TR4, it refers to only one piece of codebook information. This allows the information processing device 100 to refer to data efficiently, thereby enabling efficient search processing.
[0061] In this way, by transposing the data of the node's neighboring objects (connecting nodes) and storing the data in the node in the order in which it was compiled for each subclass, just as when referencing a lookup table for each subclass (subvector, etc.) during bulk distance calculation, the locality of reference is further improved, making it possible to speed up the process.
[0062] For simplicity of explanation, FIG. 3 shows a case where transposed data is generated for four nodes, but the unit (number of nodes) for generating transposed data is determined based on the specifications of the information processing device 100. For example, when the batch processable unit of the information processing device 100 is "4," the same processing as in FIG. 3 is performed. However, when the batch processable unit is "16," transposed data is generated with 16 nodes as one unit. When the batch processable unit is "32," transposed data is generated with 32 nodes as one unit. In this way, by using transposed data generated according to the number that the information processing device 100 can process collectively by SIMD (batch processable unit), the information processing device 100 can efficiently refer to data, thereby enabling efficient search processing.
[0063] For example, much of the time spent in the above-mentioned search process is due to distance calculation, but distance calculation can be sped up by parallel calculation using SIMD as described above. When distance calculation is parallelized in this way, the time spent fetching object data takes up the majority of the search process. If this fetch time can be reduced, further speed-up can be achieved. Note that prefetching can be used to reduce the time spent fetching data, but there are limitations to this.
[0064] On the other hand, the information processing device 100 can suppress data fetching and achieve further speedup by increasing the locality of reference as described above and enabling efficient data reference.
[0065] (Search results) When returning search results based on the second distance (true distance) described above, the following two methods, a first method and a second method, can be considered.
[0066] For example, a first method may involve searching for a number of searches greater than a specified search number, and upon completion of the search, calculating the true distance and sorting the objects by distance to obtain the specified number of search results. In this case, the information processing device 100 sets the expanded search number (also referred to as the "second number") as the search number so as to extract a number of nodes (also referred to as the "expanded search number") greater than the specified search number (also referred to as the "first number"), and performs search processing as shown in FIG. 11 to extract the nodes of the expanded search number, i.e., the number of nodes greater than the specified search number, as neighbor candidate nodes. The information processing device 100 then calculates a second distance (true distance) for the neighbor candidate nodes, and extracts a first number of nodes from the neighbor candidate nodes with the shortest second distance as neighbor nodes of the search query.
[0067] As another example, a second method may involve calculating a second distance (true distance) based on the first distance (approximate distance) only when a node (object) falls within the search range or the search area during a search, and then replacing the approximate distance with the true distance to return search results based on the true distance.
[0068] The search range here is the range defined by "r" in FIG. 11, and the search range is the range defined by "r(1+ε)" using the search range coefficient "ε" in FIG. 11. For example, when calculating the second distance (true distance) when a node falls within the search range, the information processing device 100 calculates the second distance (true distance) of the node when the first distance (approximate distance) of the node is equal to or less than "r". Furthermore, when calculating the second distance (true distance) when a node falls within the search range, the information processing device 100 calculates the second distance (true distance) of the node when the first distance (approximate distance) of the node is equal to or less than "r(1+ε)".
[0069] For example, if accuracy is to be improved, the condition may be whether the object falls within the search range. Alternatively, if speeding up processing is desired, the condition may be whether the object falls within the search range. Note that the above is merely an example, and the condition to be used may be set appropriately depending on the value of the search range coefficient "ε" and the purpose of processing.
[0070] (Vector quantization) The information processing device 100 may perform two-stage vector quantization as disclosed in Patent Document 3. Then, the information processing device 100 may calculate the distance between the search query and each node by the following formula (1).
[0071]
number
[0072] Here, the value on the left side of the above formula (1) indicates, for example, the squared distance between the search query and the node. Also, for example, "x" in the above formula (1) corresponds to the query. Also, for example, "y" in the above formula (1) corresponds to the node. Also, for example, "q" in the right side of the above formula (1) c"(y)" indicates the representative vector (centroid) of "y". For example, when the information processing device 100 does not have vector data of the node for "y" in the above formula (1), it may use the numerical value of the centroid of the partial area to which each node belongs. c (y)" indicates the residual vector. For example, "q" in the right-hand side of the above formula (1) p " indicates a predetermined quantizer (function).
[0073] Also, for example, "j" on the right side of the above formula (1) may be the number of divided spaces. For example, in the example of FIG. 1, "j" on the right side of the above formula (1) may be the number of divided spaces, "4". Also, for example, "u" on the right side of the above formula (1) j ()" indicates a partial residual vector between the vectors in the parentheses. For example, the information processing device 100 may calculate the distance between the query and the node by calculating and adding up the squared distance between the query and the node in each subspace using the above formula (1).
[0074] [1-2. Information Processing System Configuration] As shown in Fig. 4, the information processing system 1 includes a terminal device 10, an information providing device 50, and an information processing device 100. The terminal device 10, the information providing device 50, and the information processing device 100 are connected to each other via a predetermined network N so as to be able to communicate with each other via wired or wireless communication. Fig. 4 is a diagram showing an example of the configuration of the information processing system according to the first embodiment. Note that the information processing system 1 shown in Fig. 4 may include a plurality of terminal devices 10, a plurality of information providing devices 50, and a plurality of information processing devices 100.
[0075] The terminal device 10 is an information processing device used by a user. The terminal device 10 accepts various operations by the user. In the following, the terminal device 10 may be referred to as a user. In other words, in the following, the user may also be read as the terminal device 10. The above-mentioned terminal device 10 may be realized, for example, by a smartphone, a tablet terminal, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), or the like.
[0076] The information providing device 50 is an information processing device that stores information for providing various information to users, etc. For example, the information providing device 50 stores object IDs based on character information, etc. collected from various external devices, such as web servers. For example, the information providing device 50 is an information processing device that provides an image search service to users, etc. For example, the information providing device 50 stores various pieces of information for providing the image search service. For example, the information providing device 50 provides the information processing device 100 with vector information corresponding to an image that is a target of the image search service. Furthermore, the information providing device 50 transmits a query to the information processing device 100, thereby receiving from the information processing device 100 an object ID, etc., indicating an image corresponding to the query.
[0077] The information processing device 100 is a computer that provides a search service. The information processing device 100 provides a specified number of objects corresponding to a search query as search results. The information processing device 100 provides nodes (objects) near the search query as search results, using a graph in which nodes (objects) to be searched are connected by edges. In a search process that searches for nodes near the search query using a graph in which nodes corresponding to each of a plurality of objects are connected by edges, the information processing device 100 calculates the distance between the search query and multiple nodes selected according to a predetermined criterion, using vector information of the multiple nodes that have been vector-quantized.
[0078] For example, when the information processing device 100 receives a query (search query) from the terminal device 10, it searches for a target (object) similar to the search query and provides the terminal device with the search results. Furthermore, for example, the data provided by the information processing device 100 to the terminal device may be the data itself, such as image information, or may be information for referencing corresponding data, such as a uniform resource locator (URL). Furthermore, the search query and the search target (object) may be any type of data, such as image, audio, or text data.
[0079] 1-3. Configuration of information processing device Next, the configuration of the information processing device 100 according to the first embodiment will be described with reference to Fig. 5. Fig. 5 is a diagram showing an example of the configuration of the information processing device 100 according to the first embodiment. As shown in Fig. 5, the information processing device 100 has a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing device 100 may also have an input unit (e.g., a keyboard, a mouse, etc.) that accepts various operations from an administrator of the information processing device 100, and a display unit (e.g., a liquid crystal display, etc.) that displays various information.
[0080] (Communication unit 110) The communication unit 110 is realized by, for example, a network interface card (NIC) etc. The communication unit 110 is connected to a network (for example, network N in FIG. 4) by wire or wirelessly, and transmits and receives information to and from the terminal device 10 and the information providing device 50.
[0081] (Storage unit 120) The storage unit 120 is realized by, for example, a semiconductor memory element such as a random access memory (RAM) or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in FIG. 5 , the storage unit 120 according to the first embodiment includes an object information storage unit 121, a graph information storage unit 122, a quantization information storage unit 123, and a codebook information storage unit 124.
[0082] (Object information storage unit 121) The object information storage unit 121 according to the first embodiment stores various information related to objects. For example, the object information storage unit 121 stores object IDs and vector data. FIG. 6 is a diagram showing an example of the object information storage unit according to the first embodiment. The object information storage unit 121 shown in FIG. 6 includes items such as "object ID" and "vector information."
[0083] "Object ID" indicates identification information for identifying an object. "Vector information" indicates vector information corresponding to an object identified by the object ID. That is, in the example of FIG. 6, vector data (vector information) corresponding to an object is registered in association with the object ID that identifies the object.
[0084] For example, in the example of FIG. 6, an object (target) identified by the object ID "OB1" is associated with multidimensional vector information of "10, 24, 51, 2...".
[0085] The object information storage unit 121 is not limited to the above, and may store various types of information depending on the purpose.
[0086] (Graph information storage unit 122) The graph information storage unit 122 according to the first embodiment stores various types of information related to graphs. For example, the graph information storage unit 122 stores a generated graph. FIG. 7 is a diagram illustrating an example of the graph information storage unit according to the first embodiment. The graph information storage unit 122 shown in FIG. 7 has items such as "node ID," "object ID," and "connection node information."
[0087] "Node ID" indicates identification information for identifying each node (object) in the graph. Also, "object ID" indicates identification information for identifying an object. When the node ID and object ID are the same, the object ID is stored in "node ID", and the graph information storage unit 122 does not need to include an item for "object ID". For example, when used as an object ID and a node ID, the object ID is stored in "node ID", and the graph information storage unit 122 does not need to include an item for "object ID".
[0088] Furthermore, "connection node information" indicates information about nodes (referenced nodes) that can be traced from the corresponding node. For example, "connection node information" includes information such as "reference destination." "Reference destination" indicates information for identifying a reference destination (node) that is connected by an edge and can be traced from that node. That is, in the example of FIG. 7, a node ID (object ID) that identifies a node is associated with a reference destination (node) that can be traced from that node by an edge and is registered. Note that "connection node information" may also include information (edge ID) for identifying an edge connected to a reference destination.
[0089] 7, the node (node N1) identified by the node ID "N1" corresponds to the object (target) identified by the object ID "OB1." Also, an edge connects node N1 to a node (node N4) identified by the node ID "N4," indicating that it is possible to trace from node N1 to node N4.
[0090] It also indicates that the node (node N2) identified by the node ID "N2" corresponds to the object (target) identified by the object ID "OB2." An edge is connected from node N2 to the node (node N6) identified by the node ID "N6," indicating that it is possible to trace from node N2 to node N6.
[0091] It also shows that the node identified by the node ID "N7" (node N7) corresponds to the object (target) identified by the object ID "OB7." Edges are connected from node N7 to nodes N9, N12, N54, and N85, indicating that it is possible to trace from node N2 to each of nodes N9, N12, N54, and N85.
[0092] The graph information storage unit 122 is not limited to the above, and may store various types of information depending on the purpose. For example, the graph information storage unit 122 may store the lengths of edges connecting nodes (vectors). That is, the graph information storage unit 122 may store information indicating the distances between nodes (vectors). The graph information storage unit 122 is not limited to the above, and may store graph information using various data structures.
[0093] The graph may also include a program module that receives a query as input, searches for nodes by tracing edges in the graph, and extracts and outputs nodes similar to the query. That is, the graph may be intended for use as a program module that performs search processing using the graph. For example, the graph GR1 may be a program that, when vector data is input as a query, extracts and outputs nodes corresponding to vector data similar to the vector data from the graph. For example, the graph GR1 may be data used as a program module that searches for images similar to a query image. For example, the graph GR1 causes a computer to function to extract and output nodes similar to the query in the graph based on an input query.
[0094] (Quantization information storage unit 123) The quantization information storage unit 123 according to the first embodiment stores various information related to the allocation process. Fig. 8 is a diagram illustrating an example of the quantization information storage unit according to the first embodiment. In the example of Fig. 8, the quantization information storage unit 123 has items such as "node ID", "object ID", and "quantization information".
[0095] "Node ID" indicates identification information for identifying each node (object) in the graph. Also, "object ID" indicates identification information for identifying an object. Note that if the node ID and object ID are the same, the object ID is stored in "node ID," and the quantization information storage unit 123 does not need to include an "object ID" item.
[0096] Furthermore, "quantization information" indicates information about the quantized vector of each node (object). For example, "quantization information" includes information such as "element" and "codebook ID." "Element" indicates the arrangement of the vector of the corresponding object. In the example of FIG. 8, "element" indicates a case where "#1," "#2," "#3," and "#4" are included. In this case, the vector of each node (object) is divided into four, and each divided partial vector is quantized by the codebook. Note that the number of divisions is not limited to four. For example, if the number of divisions is six, "element" includes "#1," "#2," "#3," "#4," "#5," and "#6." "Codebook ID" indicates information for identifying the codebook corresponding to each element (partial vector).
[0097] 8, the vector of node N9 (object OB9) indicates that the first partial vector of the four divided partial vectors is to be quantized by a codebook (codebook CD12) identified by a codebook ID "CD12". Also, the vector of node N9 (object OB9) indicates that the second partial vector from the top of the four divided partial vectors is to be quantized by a codebook (codebook CD23) identified by a codebook ID "CD23".
[0098] The vector of node N9 (object OB9) indicates that the third partial vector from the top of the four divided partial vectors is to be quantized by the codebook (codebook CD35) identified by the codebook ID "CD35". The vector of node N9 (object OB9) also indicates that the fourth partial vector from the top (i.e., the last partial vector) of the four divided partial vectors is to be quantized by the codebook (codebook CD47) identified by the codebook ID "CD47".
[0099] The quantization information storage unit 123 is not limited to the above, and may store various types of information depending on the purpose.
[0100] (Codebook information storage unit 124) The codebook information storage unit 124 according to the first embodiment stores various pieces of information related to codebooks. For example, the codebook information storage unit 124 stores codebook IDs and vector information of each codebook. FIG. 9 is a diagram illustrating an example of the codebook information storage unit according to the first embodiment. In the example of FIG. 9, the codebook information storage unit 124 stores a lookup table indicating the correspondence between each codebook and a vector.
[0101] The codebook information storage unit 124 stores codebook information TB1 used to quantize the first partial vector among the four divided partial vectors, and codebook information TB2 used to quantize the second partial vector from the top among the four divided partial vectors. The codebook information storage unit 124 also stores codebook information TB3 used to quantize the third partial vector from the top among the four divided partial vectors, and codebook information TB4 used to quantize the fourth partial vector from the top (i.e., the last) among the four divided partial vectors.
[0102] In the example of Fig. 9, codebook information TB1 stores codebook information such as a codebook (codebook CD11) identified by a codebook ID "CD11" and a codebook (codebook CD12) identified by a codebook ID "CD12". For example, codebook CD11 indicates that it is associated with multidimensional vector information of "5, 13...". Also, codebook CD12 indicates that it is associated with multidimensional vector information of "27, 51...".
[0103] The codebook information storage unit 124 may store various types of information depending on the purpose, without being limited to the above. For example, the codebook information storage unit 124 may store information indicating the difference (distance) between each codebook and the search query.
[0104] (control unit 130) 5, the control unit 130 is a controller, and is realized by, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), or the like, executing various programs (corresponding to examples of information processing programs) stored in a storage device inside the information processing device 100 using a RAM as a work area. The control unit 130 is also a controller, and is realized by, for example, an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0105] 5, the control unit 130 has an acquisition unit 131, a generation unit 132, a search processing unit 133, and a provision unit 134, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 5, and may be any other configuration as long as it performs the information processing described below.
[0106] (Acquisition part 131) The acquisition unit 131 acquires various types of information. For example, the acquisition unit 131 acquires various types of information from the storage unit 120. For example, the acquisition unit 131 acquires various types of information from the object information storage unit 121, the graph information storage unit 122, the quantization information storage unit 123, the codebook information storage unit 124, etc. The acquisition unit 131 also acquires various types of information from an external information processing device. The acquisition unit 131 acquires various types of information from the terminal device 10 and the information providing device 50.
[0107] The acquiring unit 131 acquires a graph. For example, the information processing device 100 may acquire spatial information SP1 in Fig. 1. For example, the information processing device 100 may acquire a graph from an external device such as the information providing device 50.
[0108] The acquisition unit 131 acquires a search query for multiple objects that are the target of a data search. For example, the acquisition unit 131 acquires information about a search query QE1. For example, the acquisition unit 131 acquires a search query related to an image search. For example, the acquisition unit 131 acquires a query from the terminal device 10 used. For example, the acquisition unit 131 acquires a query from the information providing device 50 that has accepted the query from the terminal device 10 used.
[0109] (Generation unit 132) The generation unit 132 generates various types of information. For example, the generation unit 132 generates various types of information (data) from information (data) stored in the storage unit 120. For example, the generation unit 132 generates various types of information from information (data) stored in the object information storage unit 121, the graph information storage unit 122, the quantization information storage unit 123, the codebook information storage unit 124, etc.
[0110] The generation unit 132 may generate a graph as shown in the graph information storage unit 122. For example, the generation unit 132 generates spatial information SP1. The generation unit 132 may also generate information related to vector quantization as shown in the quantization information storage unit 123. For example, the generation unit 132 generates information obtained by vector quantizing each object such as node N1 (object OB1). The generation unit 132 may also generate information related to a codebook as shown in the codebook information storage unit 124. The generation unit 132 may also generate a lookup table for the codebook. For example, the generation unit 132 generates information related to a codebook such as codebook information TB1 to TB4. Note that when the information processing device 100 acquires the information shown in the graph information storage unit 122, the quantization information storage unit 123, and the codebook information storage unit 124 from an external device such as the information providing device 50, the information processing device 100 does not need to include the generation unit 132.
[0111] (Search processing unit 133) The search processing unit 133 provides a search service related to objects. The search processing unit 133 searches for various types of information. The search processing unit 133 searches for various types of information. For example, the search processing unit 133 searches for objects by searching a graph. For example, when a query is acquired by the acquisition unit 131, the search processing unit 133 searches a graph to search for objects similar to the query. For example, the search processing unit 133 extracts objects similar to the query by searching a graph. For example, the search processing unit 133 extracts objects similar to the query by searching a graph based on the processing procedure shown in FIG. 11. Note that if the information processing device 100 does not provide a search service, it may not have the search processing unit 133.
[0112] The search processing unit 133 selects various pieces of information in the search process. The search processing unit 133 extracts various pieces of information in the search process. The search processing unit 133 determines various pieces of information in the search process. The search processing unit 133 determines various pieces of information in the search process. The search processing unit 133 changes various pieces of information in the search process. The search processing unit 133 updates various pieces of information in the search process.
[0113] In a search process that searches for nodes near a search query using a graph in which nodes corresponding to each of a plurality of objects are connected by edges, the search processing unit 133 calculates the distance between the search query and a plurality of nodes selected according to a predetermined criterion using vector information of the plurality of nodes that have been vector-quantized. The search processing unit 133 calculates the distances between the plurality of nodes and the search query in parallel and performs the calculations in a batch. The search processing unit 133 performs parallel calculations of the distances between the search query and a plurality of nodes for a batch processing number determined based on the specifications of the information processing device 100.
[0114] The search processing unit 133 calculates the distance between the search query and multiple nodes using a codebook associated with the representative vector corresponding to each of the multiple nodes. The search processing unit 133 calculates the distance between the search query and multiple nodes connected by edges from one node. The search processing unit 133 calculates the distance between the search query and multiple nodes using node information in which reference information indicating the multiple nodes and the representative vectors associated with each of the multiple nodes is stored in association with one node.
[0115] The search processing unit 133 calculates the distance between the multiple nodes and the search query using vector information of the multiple nodes, each divided into multiple partial vectors by Cartesian product quantization. The search processing unit 133 calculates the distance between the multiple nodes and the search query using multiple codebooks associated with representative vectors corresponding to vectors at each division position of the multiple partial vectors into which each of the multiple nodes is divided. The search processing unit 133 calculates the distance between the multiple nodes and the search query using node information in which reference information indicating the multiple nodes and the representative vectors to which each of the multiple partial vectors of each of the multiple nodes is associated is stored in association with one node.
[0116] The search processing unit 133 calculates the distances between the multiple nodes and the search query using reference information in which a list of representative vectors, each of which is associated with each of the multiple partial vectors of each node, is associated with each of the multiple nodes.The search processing unit 133 calculates the distances between the multiple nodes and the search query using reference information including transposition information in which representative vectors corresponding to each of the multiple partial vectors of each of the multiple nodes are listed for each division position, and association information indicating the corresponding positions of each of the multiple nodes in the list.The search processing unit 133 calculates the distances between the multiple nodes and the search query using a codebook in which each of the multiple partial vectors of each of the multiple nodes is associated with a corresponding representative vector.
[0117] The search processing unit 133 calculates a second distance that is not vector quantized, unlike the first distance, which is a distance that is vector quantized, for a predetermined node among the nodes that are processed in the search process. The search processing unit 133 extracts a second number of nodes that is greater than the first number of nodes extracted as neighboring nodes of the search query in the search process as neighboring candidate nodes, and calculates the second distance for the neighboring candidate nodes. The search processing unit 133 extracts the first number of nodes from the neighboring candidate nodes with the shortest second distance as neighboring nodes of the search query.
[0118] The search processing unit 133 calculates the second distance for nodes whose first distance is within a predetermined threshold. The search processing unit 133 calculates the second distance for nodes within a search range that indicates the target range to be extracted as nearby nodes. The search processing unit 133 calculates the second distance for nodes within a search range that indicates the target range of the search process.
[0119] (Provider 134) The providing unit 134 provides various types of information. For example, the providing unit 134 transmits various types of information to the terminal device 10 or the information providing device 50. For example, the providing unit 134 provides an object ID corresponding to a search query as a search result. The providing unit 134 transmits the search result to the terminal device 10. The providing unit 134 provides the object ID searched for by the search processing unit 133 to the terminal device 10 as the search result corresponding to the search query.
[0120] Furthermore, the providing unit 134 may provide the object ID searched for by the search processing unit 133 to the information providing device 50. For example, the providing unit 134 provides the object ID extracted by the search processing unit 133 through a search to the information providing device 50. The providing unit 134 provides the object ID extracted by the search processing unit 133 to the information providing device 50 as information indicating a vector corresponding to the query.
[0121] [1-4. Information processing flow] Next, the procedure of information processing by the information processing system 1 according to the first embodiment will be described with reference to Fig. 10. Fig. 10 is a flowchart showing an example of information processing according to the first embodiment.
[0122] 10, the information processing device 100 acquires a search query for a plurality of objects to be subjected to a data search (step S101). In the example of FIG. 1, the information processing device 100 acquires a search query QE1.
[0123] Then, in a search process in which nodes corresponding to each of a plurality of objects are connected by edges to search for nodes near the search query, the information processing device 100 calculates the distance between the search query and a plurality of nodes selected according to a predetermined criterion using vector information of the plurality of vector-quantized nodes (step S102). In the example of Fig. 1, the information processing device 100 calculates the distance between the search query QE1 and a plurality of nodes N9, N12, N54, and N85 using vector information of the plurality of vector-quantized nodes N9, N12, N54, and N85.
[0124] [1-5. Search processing example] Here, an example of search processing according to the first embodiment will be described using FIG. 11 as an example. FIG. 11 is a flowchart showing an example of search processing according to the first embodiment. The search processing described below is performed by the search processing unit 133 of the information processing device 100. Furthermore, the term "object" below may be read as "node." Note that, below, the information processing device 100 (search processing unit 133) performs the search processing. Note that, if a search service is not provided, the information processing device 100 does not need to have the search processing unit 133. The search query in the processing described below may be an additional node, a target node, an object specified by the user, or the like.
[0125] Here, the neighborhood object set N(G, y) is a set of neighborhood objects associated with the node y by an edge attached thereto. "G" may be predetermined graph data (for example, graph GR1 shown in the spatial information SP1). For example, the information processing device 100 executes a k-nearest neighbor search process.
[0126] For example, the information processing device 100 sets the radius r of the hypersphere to ∞ (infinity) (step S300) and extracts a subset S from an existing object set (step S301). For example, the information processing device 100 may extract an object (node) selected as a root node as the subset S. Furthermore, for example, the hypersphere is a virtual sphere indicating the search range. Note that the objects included in the object set S extracted in step S301 are also included in the initial set of object set R of the search results.
[0127] Next, when the search query object is y, the information processing device 100 extracts the object having the shortest distance from the search query object y from among the objects included in the object set S, and sets the extracted object as object s (step S302). For example, if the object (node) selected as the root node is the only element of S, the information processing device 100 extracts the root node as object s as a result. Next, the information processing device 100 excludes object s from the object set S (step S303).
[0128] Next, the information processing device 100 determines whether the distance d(s, y) between object s and object y exceeds r(1+ε) (step S304). Here, ε is an extension factor, and r(1+ε) is a value indicating the radius of the search range (only nodes within this range are searched. Precision can be improved by making it larger than the search range). If the distance d(s, y) between object s and object y exceeds r(1+ε) (step S304: Yes), the information processing device 100 outputs object set R as a neighborhood object set of object y (step S305), and ends the process.
[0129] If the distance d(s, y) between object s and search query object y does not exceed r(1+ε) (step S304: No), the information processing device 100 selects one object that is not included in object set C from among the objects that are elements of the neighborhood object set N(G, s) of object s, and stores the selected object u in object set C (step S306). Object set C is provided for convenience to avoid duplicate searches, and is set to an empty set at the start of processing.
[0130] Next, the information processing device 100 determines whether the distance d(u, y) between the object u and the object y is r(1+ε) or less (step S307). If the distance d(u, y) between the object u and the object y is r(1+ε) or less (step S307: Yes), the information processing device 100 adds the object u to the object set S (step S308). If the distance d(u, y) between the object u and the object y is not r(1+ε) or less (step S307: No), the information processing device 100 performs the determination (processing) of step S309.
[0131] Next, the information processing device 100 determines whether the distance d(u, y) between the object u and the object y is equal to or less than r (step S309). If the distance d(u, y) between the object u and the object y exceeds r (step S309: No), the information processing device 100 performs the determination (processing) of step S315. That is, if the distance d(u, y) between the object u and the object y is not equal to or less than r, the information processing device 100 performs the determination (processing) of step S315.
[0132] If the distance d(u, y) between object u and object y is less than or equal to r (step S309: Yes), the information processing device 100 adds object u to object set R (step S310). Then, the information processing device 100 determines whether the number of objects included in object set R exceeds ks (step S311). The predetermined number ks is a natural number that is determined arbitrarily. For example, ks may be the number of searches or the number of extraction targets. Furthermore, for example, when no upper limit is set on the number of objects to be extracted in a range search or the like, ks may be set to infinity. For example, ks may be 4. If the number of objects included in object set R does not exceed ks (step S311: No), the information processing device 100 performs the determination (processing) of step S313.
[0133] If the number of objects included in object set R exceeds ks (step S311: Yes), information processing device 100 excludes from object set R the object that is the longest (farthest) distance from object y among the objects included in object set R (step S312).
[0134] Next, the information processing device 100 determines whether the number of objects included in the object set R matches ks (step S313). If the number of objects included in the object set R does not match ks (step S313: No), the information processing device 100 performs the determination (processing) of step S315. If the number of objects included in the object set R matches ks (step S313: Yes), the information processing device 100 sets the distance between the object y and the object that is the farthest (farthest) from object y among the objects included in the object set R as a new r (step S314).
[0135] Then, the information processing device 100 determines whether or not all objects that are elements of the neighborhood object set N(G, s) of object s have been selected and stored in the object set C (step S315). If all objects that are elements of the neighborhood object set N(G, s) of object s have not been selected and stored in the object set C (step S315: No), the information processing device 100 returns to step S306 and repeats the process.
[0136] When all objects that are elements of the neighborhood object set N(G, s) of object s have been selected and stored in the object set C (step S315: Yes), the information processing device 100 determines whether the object set S is an empty set (step S316). If the object set S is not an empty set (step S316: No), the information processing device 100 returns to step S302 and repeats the process. If the object set S is an empty set (step S316: Yes), the information processing device 100 outputs the object set R and ends the process (step S317). For example, the information processing device 100 may select an object (node) included in the object set R as a neighborhood node corresponding to the added node (input object y). For example, the information processing device 100 may extract (select) an object (node) included in the object set R as a neighborhood node corresponding to the target node (input object y). Furthermore, for example, the information processing device 100 may provide the objects (nodes) included in the object set R as a search result corresponding to the search query (input object y) to the terminal device or the like that performed the search.
[0137] 2. Second Embodiment From here, a second embodiment will be described. In the second embodiment, nearby objects are grouped by clustering or the like, and distance calculation is performed on each group (hereinafter also referred to as "blob") collectively. That is, in the second embodiment, distance calculation is performed collectively on a blob-by-blob basis. Note that explanations of points similar to those in the first embodiment will be omitted as appropriate. In the second embodiment, the information processing system 1 has an information processing device 100A instead of the information processing device 100.
[0138] [2-1. Information Processing] First, an overview of information processing according to the second embodiment will be described with reference to Fig. 12. Fig. 12 is a diagram showing an example of information processing according to the second embodiment.
[0139] In the example of Fig. 12, it is assumed that the information processing device 100A has already acquired a graph GR21 as shown in spatial information SP21. For example, the spatial information SP21 in Fig. 1 may be a Euclidean space. For example, the spatial information SP21 corresponds to the number of dimensions of the vector of the object, and is assumed to be a multidimensional space of 100 dimensions, 1000 dimensions, etc. Note that, for the sake of simplicity, the example of Fig. 12 will be described based on a diagram in which direct product quantization has not been performed, but direct product quantization of vectors (space) may also be performed as in Fig. 1.
[0140] First, each piece of information shown in Fig. 12 will be described. The dotted lines connecting the white circles (◯) that are nodes in the spatial information SP21 in Fig. 12 indicate edges that connect the nodes. In Fig. 12, undirected edges are shown as an example, as in Fig. 1, but the edges of the graph GR21 are not limited to undirected edges and may be directed edges. Note that the nodes and edges are the same as in Fig. 1, so a detailed description will be omitted.
[0141] In the spatial information SP21 in FIG. 12, the areas surrounded by straight lines are obtained by grouping nearby objects together through clustering or the like, and these areas are called blobs. In the example shown below, the blob is the processing unit for the collective distance calculation. FIG. 12 shows a case where the objects are clustered into 10 areas (blobs) of blobs BL1 to BL10. For example, blob BL1 indicates that it is a blob to which multiple nodes such as nodes N7, N9, N85, and N126 belong. The black dots in blobs BL1 to BL10 indicate the centroids (representative vectors) of each blob.
[0142] The blobs BL1 to BL10 are generated using various techniques related to clustering, etc., as appropriate. For example, the clustering for classifying into blobs may be k-means clustering. Furthermore, for example, the clustering for classifying into blobs may not use the centroids obtained from the center coordinates (average) of each cluster in one iteration (assignment) of k-means in the next iteration, but may instead replace the centroids with the objects closest to each centroid before performing the next iteration. In this case, the centroids of the final clusters obtained by k-means are existing objects, i.e., nodes. This increases the tendency for clusters (blobs) and the edges (neighboring nodes) of each node to match, further improving search performance.
[0143] In the example of FIG. 12, the information processing device 100A performs a search process using information on a plurality of blobs BL1 to BL10 classified by clustering a plurality of nodes (objects) as shown in space information SP21.
[0144] Next, a search process for the search query QE2 will be described. First, the information processing device 100A acquires the search query QE2 (step S21). For example, the information processing device 100A acquires the search query QE2 from the terminal device 10 (see FIG. 4) used by the user.
[0145] Then, the information processing device 100A calculates the distance between the vector of the search query QE2 and the vector of the codebook. Although not shown, if product quantization is not being performed, the information processing device 100A calculates the distance (difference) between the vector of the search query QE2 and the vector of the codebook for quantizing the entire vector. Then, the information processing device 100A stores the distance (difference) between the vector of the search query QE2 and the vector of each codebook in association with information for identifying each codebook as codebook information TB21. Note that if product quantization is being performed, the information processing device 100A calculates the distance (difference) between each sub-vector of the search query QE2 and the vector of each codebook, as shown in codebook information TB1 to TB4 in FIG. 2, similarly to FIG. 1. Note that the calculation of the distance between the vector of the search query and the vector of the codebook is similar to FIG. 1, FIG. 2, etc., and therefore detailed description thereof will be omitted.
[0146] The information processing device 100A executes a search process targeting the search query QE2 (step S22). The information processing device 100A performs a search process as shown in FIG. 16 using the graph GR21 targeting the search query QE2, thereby obtaining search results for the search query QE2. The search process shown in FIG. 16 will be described in detail later. By performing the search process targeting the search query QE2, the information processing device 100A extracts nodes of the search number as nodes in the vicinity of the search query QE2.
[0147] In a search process targeting the search query QE2, the information processing device 100A selects a predetermined node from among the nodes as a node (starting node) that will be the starting point of a search of the graph GR21. In the example of Fig. 12, the information processing device 100A selects node N9 as the starting node. In the example of Fig. 12, the information processing device 100A starts a search process targeting, for example, node N9.
[0148] Here, in the search processing for the search query QE2, the information processing device 100A calculates in parallel the distances between the search query and multiple nodes belonging to the blob to which the node to be processed belongs (step S23). In Fig. 12, as shown in the batch processing information LT2, the information processing device 100A calculates in parallel the distances between the search query QE2 and nodes N7, N9, N85, N126, etc., which are nodes belonging to the blob BL1 to which the node N9 to be processed belongs.
[0149] The information processing device 100A calculates the distance between each of nodes N7, N9, N85, and N126 and the search query QE2 using the distance between each codebook indicated in codebook information TB21 and the search query QE2. For example, the information processing device 100A refers to the codebook information TB21 and calculates the distance of the codebook corresponding to the vector of node N7 as the distance between node N7 and the search query QE2. Note that distance calculation when direct product quantization is performed is similar to that in FIGS. 1 and 2, and therefore detailed description thereof will be omitted.
[0150] The batch processable unit is determined based on the specifications of the information processing device 100A, as in the case of the information processing device 100. For example, if the number of nodes that the information processing device 100A can process in a batch using SIMD (the batch processable unit) is "4," the information processing device 100A calculates the distance between four nodes in parallel, as in FIG. 1. For example, the information processing device 100A calculates the distance between each of nodes N7, N9, N85, and N126 and the search query QE2 in a batch by parallelizing SIMD operations. By parallelizing the distance calculation, the information processing device 100A can speed up the distance calculation and enable efficient search processing. Furthermore, the information processing device 100A uses the graph GR21 to trace nodes (e.g., nodes N12 and N64) connected by edges from node N7, and performs a batch distance calculation for blobs (e.g., blobs BL2 and BL3) to which these nodes belong, thereby performing search processing. In the search process, the information processing device 100A prevents the same blob from being repeatedly subjected to batch distance calculation, as will be described in detail later.
[0151] 12 shows a case where distances are calculated in parallel for four nodes for simplicity of explanation, but the number of nodes to be parallelized is determined based on the specifications of the information processing device 100A. This point is also the same as in FIGS. 1 and 2, and therefore a detailed explanation will be omitted. The information processing device 100A may also use transposed data, as in the example shown in FIG. 3.
[0152] As described above, the information processing device 100A can perform efficient search processing by calculating the approximate distance (second distance) between the vector of each vector-quantized node and the search query and performing search processing using the second distance. Furthermore, the information processing device 100A can perform efficient search processing by simultaneously performing distance calculations for as many nodes as can be parallelized. As described above, the information processing device 100A can perform efficient search processing by simultaneously performing calculations for blobs as a unit of batch processing.
[0153] In the search process, the information processing device 100A traverses the graph GR21 and performs the above-described process on the node to be processed, thereby extracting the number of searched nodes as nodes near the search query QE2. For example, like the information processing device 100, the information processing device 100A obtains search results by either the first method or the second method described above.
[0154] The above-described process is merely an example, and the information processing device 100A may perform search processing using various information and methods for efficient search processing. Each item in this regard will be described in detail. For example, the information processing device 100A may perform clustering to generate blobs so that the number of nodes belonging to each blob is a number that the information processing device 100 can process collectively using SIMD (a batch-processable unit).
[0155] Furthermore, the information processing device 100A may use information about blobs to prevent duplicate processing. The information processing device 100A may have a blob distance calculation flag (hereinafter simply referred to as a "flag") indicating whether distance calculation has been performed for each blob, and a table indicating to which blob each object belongs. In this case, the information processing device 100A may use the flag to manage whether each blob has been processed. For example, the information processing device 100A may identify the blob to which the node belongs using the table immediately before processing nearby nodes one by one, determine whether the blob has already been subjected to a collective distance calculation using the flag, and if not, perform collective calculation and process each node. This point will also be described in FIGS. 15 and 16. This allows the information processing device 100A to prevent duplicate processing.
[0156] 2-2. Configuration of information processing device Next, the configuration of an information processing device 100A according to the second embodiment will be described with reference to Fig. 13. Fig. 13 is a diagram showing an example of the configuration of the information processing device according to the second embodiment. As shown in Fig. 13, the information processing device 100A has a communication unit 110, a storage unit 120A, and a control unit 130A. Note that, in the information processing device 100A, descriptions of the same aspects as those of the information processing device 100 will be omitted as appropriate.
[0157] (Storage unit 120A) The storage unit 120A is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in FIG. 13 , the storage unit 120A according to the second embodiment includes an object information storage unit 121, a graph information storage unit 122, a quantization information storage unit 123A, a codebook information storage unit 124, and a blob information storage unit 125.
[0158] (Quantization information storage unit 123A) The quantization information storage unit 123A according to the second embodiment stores various information related to the allocation process. Fig. 14 is a diagram illustrating an example of the quantization information storage unit according to the second embodiment. In the example of Fig. 14, the quantization information storage unit 123A has items such as "node ID," "object ID," "blob ID," and "quantization information."
[0159] "Node ID" indicates identification information for identifying each node (object) in the graph. Also, "object ID" indicates identification information for identifying an object. Note that if the node ID and object ID are the same, the object ID is stored in "node ID," and the quantization information storage unit 123A does not need to include an "object ID" item.
[0160] "Blob ID" indicates information for identifying the blob to which the node (object) belongs.
[0161] Furthermore, "quantization information" indicates information about the quantized vector of each node (object). For example, "quantization information" includes information such as "element" and "codebook ID." "Element" indicates the arrangement of the vector of the corresponding object. In the example of FIG. 14, "element" indicates the case where "#1," "#2," "#3," and "#4" are included. In this case, the vector of each node (object) is divided into four parts, and each divided partial vector is quantized by the codebook.
[0162] In the example of FIG. 14, node N1 (object OB1) belongs to a blob (blob BL9) identified by a blob ID "BL9." For example, when a vector is divided into four parts by product quantization, the quantization information of node N1 stores information similar to the four codebooks indicated by the quantization information of node N1 shown in FIG. 8. Also, for example, when product quantization is not performed, the quantization information of node N1 stores information indicating one codebook for quantizing the vector of node N1.
[0163] The quantization information storage unit 123A is not limited to the above, and may store various types of information depending on the purpose.
[0164] (Blob information storage unit 125) The blob information storage unit 125 according to the second embodiment stores information about blobs. For example, the blob information storage unit 125 stores various information for identifying objects associated with each blob. FIG. 15 is a diagram illustrating an example of the blob information storage unit according to the second embodiment. In the example of FIG. 15, the blob information storage unit 125 includes items such as "blob ID," "node ID," and "vector information."
[0165] "Blob ID" indicates identification information for identifying a blob. "Node ID" indicates a node (object) associated with the blob identified by the blob ID. "Vector information" indicates vector information of the blob. For example, "vector information" indicates a vector corresponding to the centroid of the blob.
[0166] 15, the nodes (objects) associated with the blob (blob BL1) identified by the blob ID "BL1" are nodes N7, N9, N85, N126, etc. The example shows that the blob BL1 is associated with multidimensional vector information of "51, 4, 102, 33...".
[0167] It also indicates that the nodes (objects) associated with the blob (blob BL9) identified by the blob ID "BL9" are nodes N1, N4, N5, etc. It indicates that multidimensional vector information of "12, 55, 12, 6..." is associated with blob BL9.
[0168] Note that the blob information storage unit 125 may store various information depending on the purpose, without being limited to the above. The blob information storage unit 125 may store a flag indicating whether each blob has been processed for distance calculation. For example, the blob information storage unit 125 sets the flag value of each blob to either a value indicating that the blob has not been processed as a target for distance calculation (also referred to as an "unprocessed flag value") or a value indicating that the blob has been processed as a target for distance calculation (also referred to as a "processed flag value"). For example, at the start of the search process, the blob information storage unit 125 sets the flag value of each blob to a value indicating that distance calculation for that blob has not been processed (e.g., 0), and changes the flag value of a blob that has been processed for distance calculation to a value indicating that distance calculation has been processed (e.g., 1).
[0169] For example, the information processing device 100A refers to the flag value of the blob to be processed and determines whether or not to perform a batch distance calculation process for that blob. The information processing device 100A refers to the flag value of each blob, and if the flag value of that blob is an unprocessed flag value, it determines that the blob has not yet been subjected to batch distance calculation and performs batch distance calculation for the nodes belonging to that blob. On the other hand, if the flag value of that blob is a processed flag value, the information processing device 100A determines that the blob has already been subjected to batch distance calculation and does not perform batch distance calculation for the nodes belonging to that blob.
[0170] (control unit 130A) 13, the control unit 130A is a controller, and is realized by, for example, a CPU, an MPU, a GPU, etc., executing various programs (corresponding to examples of information processing programs) stored in a storage device inside the information processing device 100A using a RAM as a work area. The control unit 130A is also a controller, and is realized by, for example, an integrated circuit such as an ASIC or an FPGA.
[0171] 13, control unit 130A has an acquisition unit 131, a generation unit 132A, a search processing unit 133A, and a provision unit 134, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of control unit 130A is not limited to the configuration shown in FIG. 13, and may be any other configuration as long as it performs the information processing described below.
[0172] (Generation unit 132A) The generating unit 132A generates various types of information in the same manner as the generating unit 132.
[0173] The generation unit 132A may generate information related to vector quantization as shown in the quantization information storage unit 123. For example, the generation unit 132 generates information indicating that node N1 (object OB1) belongs to blob BL9. The generation unit 132 may also generate information related to blobs as shown in the blob information storage unit 125. For example, the generation unit 132 generates information indicating that nodes N7, N9, N85, N126, etc. belong to blob BL1. Note that when the information processing device 100 acquires the information shown in the graph information storage unit 122, the quantization information storage unit 123A, the codebook information storage unit 124, and the blob information storage unit 125 from an external device such as the information providing device 50, the information processing device 100 does not need to include the generation unit 132A.
[0174] (Search processing unit 133A) Similar to the search processing unit 133, the search processing unit 133A performs various processes related to the search process.
[0175] In a search process for searching for nodes near a search query among a group of nodes corresponding to each of a plurality of objects, the search processing unit 133A uses information about a plurality of blobs into which the plurality of objects are classified to calculate the distance between the search query and a plurality of nodes belonging to one blob, using vector information of the plurality of vector-quantized nodes. The search processing unit 133A calculates the distances between the plurality of nodes and the search query in parallel and performs the calculations in a batch. The search processing unit 133A performs parallel calculations of the distances between the search query and a plurality of nodes for a batch processing number determined based on the specifications of the information processing device 100A.
[0176] The search processing unit 133A calculates the distance between the search query and multiple nodes belonging to a blob adjacent to the blob to which the search query applies.The search processing unit 133A calculates the distance between the search query and multiple nodes in a search process that searches for nodes near the search query using a graph in which multiple nodes are connected by edges.The search processing unit 133A calculates the distance between the search query and multiple nodes that belong to a blob to which a connecting node connected by an edge from a target node that is the processing target in the search process belongs.
[0177] The search processing unit 133A calculates the distance between the search query and multiple nodes that belong to a blob other than the blob to which the target node belongs. The search processing unit 133A calculates the distance between the search query and multiple nodes in a search process that searches for nodes near the search query using a blob graph that connects each of the multiple nodes to blobs other than the blob to which each of the multiple nodes belongs. The search processing unit 133A calculates the distance between the search query and multiple nodes in a search process that searches for nodes near the search query using a blob graph that is a converted graph generated by connecting edges from one of the multiple nodes in a pre-conversion graph in which multiple nodes are connected by edges to a blob to which a node connected by an edge belongs, other than the blob to which the one node belongs.
[0178] The search processing unit 133A calculates the distance between the search query and multiple nodes belonging to a blob, which is a blob connected by edges from a target node to be processed in the search process. The search processing unit 133A determines whether to process the blob using information indicating a processed blob, which is a blob that has already been processed. If the blob is a processed blob, the search processing unit 133A determines that the blob is not to be processed.
[0179] The search processing unit 133A calculates a second distance that is not vector quantized, unlike the first distance, which is a distance that is vector quantized, for a predetermined node among the nodes that are processed in the search process. The search processing unit 133A extracts a second number of nodes that is greater than the first number of nodes extracted as neighboring nodes of the search query in the search process as neighboring candidate nodes, and calculates the second distance for the neighboring candidate nodes. The search processing unit 133A extracts the first number of nodes from the neighboring candidate nodes with the shortest second distance as neighboring nodes of the search query.
[0180] The search processing unit 133A calculates the second distance for nodes whose first distance is within a predetermined threshold. The search processing unit 133A calculates the second distance for nodes within a search range that indicates the target range to be extracted as nearby nodes. The search processing unit 133A calculates the second distance for nodes within a search range that indicates the target range of the search process.
[0181] [2-3. Search processing example] Here, an example of the search processing according to the second embodiment will be described with reference to FIG. 16. FIG. 16 is a flowchart showing an example of the search processing according to the second embodiment. The search processing described below is performed by the search processing unit 133A of the information processing device 100A. Note that explanations of the same points as those in the first embodiment, such as those in FIG. 11, will be omitted as appropriate. For example, in FIG. 16, the same points as those in FIG. 11 will be assigned the same step numbers, and explanations thereof will be omitted as appropriate.
[0182] Here, the neighborhood object set N(G, y) is a set of neighborhood objects associated with the node y by an edge attached thereto. "G" may be predetermined graph data (for example, a graph GR21 shown in the spatial information SP21). For example, the information processing device 100A executes a k-nearest neighbor search process.
[0183] For example, the search process shown in Fig. 16 differs from the search process shown in Fig. 11 in that the process shown in step S304a is performed after step S304. In this way, the search process shown in Fig. 16 calculates the distance between the query and the objects (nodes) in the blob in a normal graph search when an unaccessed blob is reached, and processes each node.
[0184] Specifically, in the search process of FIG. 16, if the information processing device 100A has not yet performed distance calculation for the blob to which object s belongs, it executes all of the following (four processes shown after "-" in step S304a in FIG. 16; hereinafter referred to as "first process" to "fourth process") (step S304a). In step S304a, the information processing device 100A executes a process of batch-calculating the distances of objects in the blob (first process). Also, in step S304a, the information processing device 100A executes a process of setting a distance calculation flag for the blob (second process). Also, in step S304a, the information processing device 100A executes a process of storing all objects in the blob for which batch distance calculation has been performed in C (third process). Also, in step S304a, the information processing device 100A executes a process of performing the processes from step S307 to step S314 one by one, with all objects in the blob as u (fourth process). Thereafter, the information processing device 100A performs the processes from step S306 onwards.
[0185] [2-4. Modifications] From here, a modified example will be described. In the modified example of the second embodiment, a graph including the concept of blobs may be used. For example, in the modified example of the second embodiment, a graph in which the reference destination of an edge from a node is a blob may be used. Note that explanations of points similar to those of the first and second embodiments will be omitted as appropriate. The information processing device 100A according to the modified example has a graph information storage unit 122A instead of the graph information storage unit 122.
[0186] [2-4-1. Information Processing] First, an overview of information processing according to the modified example will be described with reference to Fig. 17. Fig. 17 is a diagram showing an example of information processing according to the modified example.
[0187] The information processing device 100A converts the graph GR21 in which nodes are connected by edges into a graph GR31 in which the reference destinations of the edges from the nodes are blobs (step S31). That is, the information processing device 100A generates the graph GR31, which is a graph into which the concept of blobs has been introduced (a graph for blobs), from the graph GR21, which is a normal graph, as shown in FIG.
[0188] In the example of FIG. 17, the information processing device 100A generates a graph GR31 including directed edges from nodes to blobs, as shown in the spatial information SP31. Note that in the graph GR31, edges between nodes are deleted, and the graph GR31 includes directed edges from nodes to blobs, but does not include edges between nodes. The arrowed line shown in FIG. 17 indicates a directed edge from the node at the origin of the arrow to the blob at the end of the arrow. That is, the arrowed line shown in FIG. 17 indicates a directed edge with the node at the origin of the arrow as the reference source and the blob at the end of the arrow as the reference destination. For example, FIG. 17 shows that edges to two blobs, blob BL8 and blob BL10, are connected from node N1.
[0189] For example, the information processing device 100A converts the edges of each node from edges to connecting nodes (neighboring nodes) to edges to blobs to which the connecting nodes belong. Note that the information processing device 100A does not generate edges to blobs to which each node itself belongs. In the example of FIG. 17, the information processing device 100A does not generate an edge from node N9 to blob BL1 or an edge from node N126 to blob BL1 because node N9 and node N126 belong to the same blob BL1. On the other hand, in the example of FIG. 17, the information processing device 100A generates an edge from node N9 to blob BL4 and an edge from node N18 to blob BL1 because node N9 and node N18 belong to different blobs BL1 and BL4.
[0190] In the example of FIG. 17, the information processing device 100A performs search processing using a graph GR31 including directed edges from nodes (objects) to blobs, as shown in spatial information SP31. For example, the information processing device 100A performs search processing as shown in FIG. 19 using graph GR21 for search query QE2, thereby obtaining search results for search query QE2. Details of the search processing shown in FIG. 19 will be described later. Note that the search processing in FIG. 17 is common to FIG. 12 in that search processing is performed using the concept of blobs, and is similar to the search processing shown in FIG. 12 except for the processing related to edges.
[0191] [2-4-2. Graph] Next, an overview of a graph according to a modified example will be described with reference to FIG. 18. FIG. 18 is a diagram illustrating an example of a graph information storage unit according to a modified example. For example, graph information storage unit 122A shown in FIG. 18 stores a graph into which the concept of blobs has been introduced (a graph for blobs). Graph information storage unit 122A shown in FIG. 18 has items such as "node ID," "object ID," and "blob information." Thus, graph information storage unit 122A according to a modified example differs from graph information storage unit 122 of FIG. 7 in that it has "blob information" instead of "connection node information." Note that descriptions of the same points as those in graph information storage unit 122 of FIG. 7 will be omitted where appropriate.
[0192] Furthermore, "blob information" indicates information about blobs (referenced blobs) that can be traced from the corresponding node. For example, "blob information" includes information such as "reference destination." "Reference destination" indicates information for identifying a reference destination (blob) that is connected by an edge and can be traced from the node. That is, in the example of FIG. 18, a node ID (object ID) that identifies a node is associated with a reference destination (blob) that can be traced from the node by an edge and is registered. Note that "blob information" may also include information (edge ID) for identifying an edge connected to the reference destination.
[0193] In the example of FIG. 18, a node (node N1) identified by a node ID "N1" corresponds to an object (target) identified by an object ID "OB1." An edge is also connected from node N1 to a blob (blob BL8) identified by a blob ID "BL8," indicating that blob BL8 can be traced from node N1. An edge is also connected from node N1 to a blob (blob BL10) identified by a blob ID "BL10," indicating that blob BL10 can be traced from node N1.
[0194] It also shows that the node (node N2) identified by node ID "N2" corresponds to the object (target) identified by object ID "OB2." An edge connects node N2 to a blob (blob BL5) identified by blob ID "BL5," indicating that blob BL5 can be traced from node N1.
[0195] The graph information storage unit 122A is not limited to the above, and may store various types of information depending on the purpose.
[0196] The graph may also include a program module that receives a query as input, searches for nodes by tracing edges in the graph, and extracts and outputs nodes similar to the query. That is, the graph may be intended for use as a program module that performs search processing using the graph. For example, the graph GR31 may be a program that, when vector data is input as a query, extracts and outputs nodes corresponding to vector data similar to the vector data from the graph. For example, the graph GR31 may be data used as a program module that searches for images similar to a query image. For example, the graph GR31 causes a computer to function to extract and output nodes similar to the query in the graph based on an input query.
[0197] [2-4-3. Search processing example] From here, an example of search processing according to the modified example will be described using FIG. 19 as an example. FIG. 19 is a flowchart showing an example of search processing according to the modified example. The search processing described below is performed by the search processing unit 133A of the information processing device 100A. Note that explanations of the same points as those in the first or second embodiment, such as FIG. 11 and FIG. 16, will be omitted as appropriate. For example, in FIG. 19, the same points as those in FIG. 11 will be assigned the same step numbers, and explanations thereof will be omitted as appropriate.
[0198] 19 differs from the search processes of FIGS. 11 and 16 in that set N(G, y) is a set of blobs associated with node y by an edge. In FIG. 19, N(G, s) and C are blob sets. That is, the "neighborhood object set N" in FIG. 11 is replaced with "blob set N" in FIG. 19, and the "object set C" in FIG. 11 is replaced with "blob set C" in FIG. 19. Furthermore, the word "object" in the process related to the above change to a blob is replaced with "blob" as appropriate. "G" may be graph data (blob graph) incorporating the concept of a blob (for example, graph GR31 shown in spatial information SP31). For example, the information processing device 100A executes a k-nearest neighbor search process.
[0199] For example, in the search process shown in Fig. 19, the process of step S306 is changed from Fig. 11 as follows: In the search process in Fig. 19, the information processing device 100A selects one blob from N(G, s) that is not included in blob set C, stores the selected blob in blob set C, calculates the objects in the blob collectively, and sets the set to B (also referred to as "object set B in target blob") (step S306).
[0200] 19 differs from the search process shown in Fig. 11 in that the process shown in step S306a is performed after step S306. Specifically, in the search process shown in Fig. 19, information processing device 100A selects one object u from object set B in the target blob (step S306a). Thereafter, information processing device 100A performs the processes from step S307 onwards.
[0201] 19 differs from the search process shown in Fig. 11 in that the process shown in step S314a is performed before step S315. Specifically, in the search process of Fig. 19, the information processing device 100A determines whether all objects have been selected from the object set B within the target blob (step S314a).
[0202] If all objects have been selected from the set B of objects in the target blob (step S314a: Yes), the information processing device 100A performs the process of step S315. On the other hand, if all objects have not been selected from the set B of objects in the target blob (step S314a: No), the information processing device 100A returns to step S306a and repeats the process.
[0203] [2-4-4. Other examples of generation processing] In the above example, a case where a blob graph is generated by converting the edges of each node from edges to connecting nodes to edges to blobs to which the connecting nodes belong is shown, but the information processing device 100A may generate a blob graph using various information as appropriate. For example, the information processing device 100A may generate a blob graph using a proximity search index (also simply referred to as an "index") that searches multiple objects. Note that explanations of points similar to those described above will be omitted as appropriate.
[0204] An example of generating a graph for blobs will be described with reference to FIG. 20. FIG. 20 is a diagram showing another example of information processing according to a modified example. Note that, although the following description will be given of a case where a graph is used as an example of an index, the index may be any type that searches multiple objects. For example, the index may be an index (hash index) that uses a description related to hashing such as a hash table, an index having a tree structure (tree index), or the like.
[0205] Also, information indicating an object or a node corresponding to that object (also referred to as an "object node") may be referred to as first information, and information used to classify multiple objects may be referred to as second information. Also, information indicating a blob may be referred to as third information, and a graph in which object nodes are connected by edges may be referred to as fourth information. For example, graph GR21 in FIG. 20 corresponds to fourth information, in which object nodes such as node N1 are connected by edges. Also, graph GR31 in FIG. 17 and graph GR32 in FIG. 20 are blob graphs (also referred to as "group graphs") in which object nodes and blobs are connected by edges, unlike the fourth information. In other words, a group graph is a graph in which edges connect object nodes to blobs, and in which object nodes are reference sources and blobs (groups) are reference destinations.
[0206] The information processing device 100A uses the graph GR21 in which object nodes are connected by edges as an index (fourth information) to generate a graph GR32, which is a group graph in which the reference destinations of edges from object nodes are blobs (step S41). In Fig. 20, the information processing device 100A generates a graph GR32 including directed edges from nodes to blobs, as shown in spatial information SP32.
[0207] For example, the information processing device 100A selects one object from a plurality of objects, searches for nearby objects of the selected object using the graph GR21, and generates the graph GR32 by performing a connection process that connects edges from one object node corresponding to the selected object to other groups to which the nearby objects belong. For example, the information processing device 100A randomly selects one object from a plurality of objects, performs a connection process on the selected object, and generates the graph GR32.
[0208] For example, the information processing device 100A performs the following processes (1-1) to (1-4) to generate the graph GR32. Note that the object may be read as a node corresponding to the object.
[0209] (1-1): Randomly get object NX (select) (1-2): Search for k neighboring objects of object NX (1-3): Get the blob that each neighboring object belongs to (1-4): Generate edges to the blob obtained from Object NX
[0210] For example, in the process (1-1), the information processing device 100A randomly acquires one object NX from a plurality of objects stored in the object information storage unit 121.
[0211] For example, in process (1-2), the information processing device 100A performs a search process as shown in FIG. 11 using the graph GR21 with the object NX as the target (query), thereby extracting k objects as neighboring objects of the object NX.
[0212] For example, in the process (1-3), the information processing device 100A refers to the information on the blob stored in the blob information storage unit 125, and acquires the blob associated with the neighboring object (node) of the object NX.
[0213] For example, in process (1-4), the information processing device 100A generates an edge from the object NX to the acquired blob. For example, the information processing device 100A generates a graph GR32 by associating information indicating the acquired blob with the object NX in the quantization information storage unit 123A.
[0214] For example, the information processing device 100A repeats the above processes (1-1) to (1-4) until there are no more objects to be processed, thereby generating the graph GR32. Note that, if the blobs indicated in the spatial information SP21 are generated by search, the search results when generating the blobs may be used to generate the graph GR32. Note that an example of generating blobs by search will be described later.
[0215] For example, the information processing device 100A performs a search process using a graph GR32 including directed edges from nodes (objects) to blobs as shown in the spatial information SP32. For example, the information processing device 100A performs a search process using the graph GR32 as shown in Fig. 19 for the search query QE2, thereby obtaining search results for the search query QE2.
[0216] In the above example, the blob is generated by clustering, but the blob may be generated by various methods other than clustering. For example, the information processing device 100A may generate information indicating the blob (third information) by using second information such as an index.
[0217] The information processing device 100A generates third information indicating a blob using a graph GR21 as an index, as shown in Fig. 21. Fig. 21 is a diagram showing an example of a process for generating blob information. Spatial information SP20 in Fig. 21 indicates a state before blob information in the spatial information SP21 is generated, and the graph GR21 in Fig. 21 is the same as the graph GR21 in Fig. 20. Note that the process shown in Fig. 21 shows part of a first method, which will be described later in detail.
[0218] For example, the information processing device 100A generates third information that classifies multiple objects into multiple blobs by a search using the graph GR21. For example, the information processing device 100A selects one node (also referred to as a "processing target node") from multiple nodes in the graph GR21, and generates the third information by a classification process that classifies into one group the classification target nodes, which are at least some of the neighboring nodes that are nodes connected to the processing target node by edges, and the object groups corresponding to each of the processing target nodes. For example, the information processing device 100A selects the processing target nodes from multiple nodes by a predetermined process, and performs classification processing on the selected processing target nodes, thereby generating the third information.
[0219] For example, the information processing device 100A may select a processing target node based on the number of neighboring nodes (connecting nodes) connected to the node by an edge. For example, the information processing device 100A may select processing target nodes in descending order of the number of neighboring nodes. A method of selecting processing target nodes in descending order of the number of neighboring nodes and generating third information indicating a blob is also referred to as a "first method." Furthermore, for example, the information processing device 100A may select processing target nodes in descending order of the number of neighboring nodes (connecting nodes). A method of selecting processing target nodes in descending order of the number of neighboring nodes and generating a blob is also referred to as a "second method." Furthermore, for example, the information processing device 100A may randomly select a processing target node. A method of randomly selecting a processing target node and generating third information indicating a blob is also referred to as a "third method." Below, an overview of the processing procedures of each of the first method, second method, and third method will be described.
[0220] [2-4-4-1. 1st method] First, the first method will be described. For example, in the first method, the information processing device 100A performs the following processes (2-1) to (2-3) to generate third information indicating a blob.
[0221] (2-1): Select the node ND with the largest number of neighboring nodes (number of edges) for each node. (2-2): The neighboring nodes of node ND and node ND are considered as one blob. (2-3): Delete the blob node and its edges.
[0222] For example, the information processing device 100A performs the processes (2-1) to (2-3) in all nodes.
[0223] For example, in the process (2-1), the information processing device 100A refers to the graph GR21 and sequentially acquires one node ND. The information processing device 100A may acquire one node ND using list information of nodes arranged in descending order of the number of neighboring nodes.
[0224] For example, in process (2-2), the information processing device 100A refers to the graph to be processed (graph GR21, etc.), and generates information that treats the neighboring nodes of node ND and node ND as one blob. Note that the information processing device 100A may treat neighboring nodes equal to or less than a constant K as blobs, rather than treating all neighboring nodes of node ND as blobs together with node ND. For example, the information processing device 100A may treat K neighboring nodes among the neighboring nodes and node ND as one blob, rather than treating all neighboring nodes of node ND together with node ND as blobs.
[0225] For example, in process (2-3), the information processing device 100A deletes the node designated as a blob in process (2-2) and the edges of that node so that duplicate nodes are not selected. For example, if the node designated as a blob is a neighboring node of another node, the information processing device 100A deletes the node (edge and) that node. Note that the above method is merely an example, and process (2-3) is not limited to a method of deleting edges, etc., and any method capable of generating a blob may be used. For example, the information processing device 100A may generate a blob by maintaining a list of nodes already belonging to the blob and managing nodes to be excluded. For example, the information processing device 100A may manage nodes already belonging to the blob by setting a predetermined flag (such as a deletion flag) on the node designated as a blob. In this way, when using flags, the information processing device 100A refers to the flag of each node during processing to determine whether to process that node.
[0226] A specific example of the above-mentioned first method will be described with reference to Fig. 21. In Fig. 21, the information processing device 100A acquires node N7, which has the maximum number of neighboring nodes (four), from graph GR21 (step S51). In Fig. 21, node N1 from graph GR21 also has four neighboring nodes, but the information processing device 100A acquires node N7. Note that when there are multiple nodes with the same number of neighboring nodes, the information processing device 100A may randomly acquire a node from the multiple nodes.
[0227] Then, the information processing device 100A classifies the four nodes N9, N12, N54, and N126, which are neighboring nodes of node N7, and the five nodes surrounding node N7, into one blob BL11 (step S52). The information processing device 100A generates third information BLT1 indicating that the five nodes N7, N9, N12, N54, and N126 belong to one blob BL11. The information processing device 100A then deletes the five nodes N7, N9, N12, N54, and N126 and the edges of each node from the graph GR21. As a result, the graph GR21 is updated to graph GR21-1, as shown in the spatial information SP20-1.
[0228] The information processing device 100A then acquires node N1, which has four neighboring nodes, from the graph GR21 (step S53). The information processing device 100A then sets four neighboring nodes of node N1, namely nodes N4, N5, N88, and N99, and the five neighboring nodes of node N1, as one blob BL12 (step S54). The information processing device 100A generates third information BLT2 indicating that the five neighboring nodes of node N1, N4, N5, N88, and N99 belong to one blob BL12. The information processing device 100A then deletes the five neighboring nodes of node N1, N4, N5, N88, and N99 and the edges of each node from the graph GR21-1.
[0229] The information processing device 100A repeats the above-described process until there are no more nodes to be processed, and generates third information indicating a plurality of blobs.
[0230] [2-4-4-2. Second method] Next, the second method will be described. For example, in the case of the second method, the information processing device 100A performs the following processes (3-1) to (3-3) to generate third information indicating a blob.
[0231] (3-1): Select the node ND with the smallest number of neighboring nodes (number of edges) for each node. (3-2): The neighboring nodes of node ND and node ND are considered as one blob. (3-3): Delete the blob node and its edges.
[0232] For example, the information processing device 100A performs the processes (3-1) to (3-3) in all nodes.
[0233] For example, in process (3-1), the information processing device 100A refers to the graph GR21 and sequentially acquires one node ND. The information processing device 100A may acquire one node ND using list information of nodes arranged in ascending order of the number of neighboring nodes.
[0234] For example, in process (3-2), the information processing device 100A refers to the graph to be processed (graph GR21, etc.), and generates information that treats the neighboring nodes of node ND and node ND as one blob. Note that the information processing device 100A may treat neighboring nodes equal to or less than a constant K as blobs, rather than treating all neighboring nodes of node ND as blobs together with node ND. For example, the information processing device 100A may treat K neighboring nodes of the neighboring nodes and node ND as one blob, rather than treating all neighboring nodes of node ND together with node ND as blobs.
[0235] For example, in process (3-3), the information processing device 100A deletes the node that was set as a blob in process (3-2) and the edges of that node so that the node is not selected more than once. For example, if the node that was set as a blob is a neighboring node of another node, the information processing device 100A deletes the node (edge and) that node. Note that the second method is similar to the first method except that the order in which nodes are acquired is from the node with the smallest number of neighboring nodes (number of edges), and therefore a description of specific examples will be omitted. For example, process (3-3) is not limited to a method of deleting edges, etc., as in the first method, but may be any method that can generate a blob. For example, the information processing device 100A may generate a blob by a method of retaining a list of nodes already belonging to the blob and managing nodes to be excluded, as in the first method.
[0236] [2-4-4-3. Third method] Next, the third method will be described. For example, in the third method, the information processing device 100A performs the following processes (4-1) to (4-3) to generate third information indicating a blob.
[0237] (4-1): Randomly acquire nodes ND sequentially. (4-2): The neighboring nodes of node ND and node ND are considered as one blob. (4-3): Delete the blob node and its edges.
[0238] For example, the information processing device 100A performs processes (4-1) to (4-3) on all nodes. Note that the third method is similar to the first and second methods except that the order in which nodes are acquired is random, and detailed description thereof will be omitted. For example, the information processing device 100A may generate the third information in a random order (third method) if all the nodes have the same number of edges.
[0239] [2-4-4-4. 4th method] The above-described first to third methods are merely examples, and the information processing device 100A may generate blobs through various processes. For example, the information processing device 100A may generate blobs through a search using a proximity search index (index) that searches multiple objects. The method of generating blobs through a search using an index is also referred to as a "fourth method."
[0240] The fourth method will be described below. Note that explanations of points similar to those described above will be omitted where appropriate. In the following, an example in which a graph is used as an index will be described as an example, but as described above, the index may be any type that searches multiple objects. For example, the index may be any type of index, such as a hash index or a tree index.
[0241] Hereinafter, a case will be described in which the information processing device 100A generates third information indicating a blob by a search using a graph GR21 as an index as shown in Fig. 21. For example, the information processing device 100A performs the following processes (5-1) to (5-4) to generate a graph GR32.
[0242] (5-1): Randomly get object NX (select) (5-2): Search for k neighboring objects of object NX (5-3): Neighboring objects of object NX and object NX are treated as one blob (5-4): Execute the process to avoid duplicate processing.
[0243] For example, in the process (5-1), the information processing apparatus 100A randomly acquires one object NX from the graph GR21.
[0244] For example, in process (5-2), the information processing device 100A performs a search process as shown in Fig. 11 using the graph GR21 with the object NX as the target (query) to extract k objects as neighboring objects of the object NX. Note that the information processing device 100A may also scan (search) all objects to extract neighboring objects of the object NX.
[0245] For example, in the process (5-3), the information processing device 100A generates information that treats the neighboring nodes of the neighboring objects of the object NX and the object NX as one blob.
[0246] For example, in process (5-4), the information processing device 100A executes a process for avoiding duplication so that one object does not belong to multiple blobs due to duplicate searches, etc. For example, the information processing device 100A deletes the object that has been set as a blob from the index. For example, the information processing device 100A deletes the object that has been set as a blob from the graph GR21.
[0247] The information processing device 100A is not limited to the above, and may use any method to prevent duplicate processing as long as one object does not belong to multiple blobs. For example, the information processing device 100A may set a predetermined flag (such as a deletion flag) on an object that has been made a blob, thereby preventing objects that already belong to the blob from being processed. When using flags in this way, the information processing device 100A refers to the flag of each object during processing to determine whether the object should be processed. For example, when using search results to generate the above-mentioned group graph, the information processing device 100A may use a method using flags to prevent duplicate processing.
[0248] 2-4-5. Information processing device according to a modified example In the information processing device 100A according to the modified example of the second embodiment, the acquisition unit 131 also performs the following processes. The acquisition unit 131 acquires first information indicating multiple objects to be targets of data search and second information used to classify the multiple objects. The acquisition unit 131 acquires the second information which is an index for searching multiple objects. The acquisition unit 131 acquires the second information which is a graph in which multiple nodes corresponding to each of the multiple objects are connected by edges. The acquisition unit 131 acquires fourth information which is an index for searching multiple objects. The acquisition unit 131 acquires the fourth information which is a graph in which multiple object nodes corresponding to each of the multiple objects are connected by edges.
[0249] Furthermore, in the information processing device 100A according to the modified example of the second embodiment, the generation unit 132A also performs the following process. The generation unit 132A uses the second information acquired by the acquisition unit 131 to classify the multiple objects indicated by the first information and generate third information indicating multiple groups used for performing batch processing in search processing targeting multiple objects. The generation unit 132A uses the second information to generate third information that classifies the multiple objects into multiple groups. The generation unit 132A generates third information that classifies the multiple objects into multiple groups through a search using the second information.
[0250] The generation unit 132A selects one node from multiple nodes in the graph, and generates third information by a classification process that classifies into one group the classification target nodes, which are at least some of the neighboring nodes that are nodes connected to the one node by edges, and the object group corresponding to each of the one node. The generation unit 132A generates the third information by a classification process that classifies into one group the classification target nodes, which are a predetermined number of the neighboring nodes of the one node, and the object group corresponding to each of the one node. The generation unit 132A generates the third information by a classification process that classifies into one group the classification target nodes, which are all of the neighboring nodes of the one node, and the object group corresponding to each of the one node.
[0251] The generation unit 132A generates third information by a classification process that selects one node from the plurality of nodes based on the number of neighboring nodes, which are nodes connected to each of the plurality of nodes by edges, and classifies the classification target node of the one node and a group of objects corresponding to each of the one node into one group. The generation unit 132A generates third information by a classification process that selects one node in descending order of the number of neighboring nodes, and classifies the classification target node of the one node and a group of objects corresponding to each of the one nodes into one group. The generation unit 132A generates third information by a classification process that selects one node in descending order of the number of neighboring nodes, and classifies the classification target node of the one node and a group of objects corresponding to each of the one nodes into one group.
[0252] The generation unit 132A generates third information by performing a classification process that randomly selects one node from multiple nodes and classifies the classification target node of the one node and the object group corresponding to each of the one node into one group. The generation unit 132A removes the classification target node of the one node classified into one group and the processed node group that is the one node from the graph, and generates the third information by repeating the classification process using the graph after removing the processed node group. The generation unit 132A generates the third information by repeating the classification process using the graph after removing the processed node group until multiple objects are classified into one of the groups. The generation unit 132A generates third information indicating multiple groups whose objects belong to each group and are mutually exclusive.
[0253] The generation unit 132A uses the fourth information to generate a group graph in which edges are connected from object nodes, which are nodes corresponding to at least some of the multiple objects, to groups other than the group to which the object node belongs, among the multiple groups. The generation unit 132A selects one object from the multiple objects, searches for neighboring objects of the selected object using the fourth information, and generates a group graph by connecting edges from the one object node corresponding to the selected object to other groups to which the neighboring objects belong. The generation unit 132A randomly selects one object from the multiple objects, searches for neighboring objects of the selected object, and generates a group graph by connecting edges from the one object node corresponding to the selected object to other groups to which the neighboring objects belong. The generation unit 132A uses the fourth information to generate a group graph.
[0254] [2-4-6. Information processing flow] Next, the procedure of information processing according to the modified example will be described with reference to Fig. 22. Fig. 22 is a flowchart showing an example of information processing according to the modified example.
[0255] 22, the information processing device 100A acquires first information indicating a plurality of objects to be subjected to data search (step S401). For example, the information processing device 100A acquires the first information indicating a plurality of objects from the object information storage unit 121.
[0256] Furthermore, the information processing device 100A acquires second information used to classify a plurality of objects (step S402). For example, the information processing device 100A acquires a graph from the graph information storage unit 122A as the second information.
[0257] Then, the information processing device 100A uses the second information to classify the multiple objects indicated by the first information and generates third information indicating multiple groups used for performing batch processing in search processing targeting multiple objects (step S403). For example, the information processing device 100A generates third information indicating multiple blobs, each of which has mutually exclusive objects.
[0258] 3. Third Embodiment In the above example, two types of graphs are used: a graph in which edges connect nodes corresponding to objects (object nodes) (also called an "object graph"), and a graph in which edges connect object nodes and groups (blobs) (blob graph), but the information processing system 1 may use various types of graphs. For example, the information processing system 1 may generate a graph in which groups (blobs) are connected (hereinafter also called a "group connection graph" or "blob connection graph"), different from the object graph and the blob graph, and use the blob connection graph for searches.
[0259] This point will be described below as a third embodiment. In the third embodiment, the information processing system 1 has an information processing device 100B instead of the information processing device 100 or the information processing device 100A. Note that the description of the same points as those described above in the first embodiment, the second embodiment, etc. will be omitted as appropriate.
[0260] The information processing device 100B uses an index (index information) to generate a group connection graph in which edges are connected from a first group (first blob) to which one object belongs to a second group (second blob) to which neighboring objects of the one object belong. In the following, a case where a graph is used as an example of an index will be described as an example, but as described above, the index may be any type of index as long as it searches multiple objects. For example, the index may be any type of index, such as a hash index or a tree index.
[0261] [3-1. Information Processing] First, an overview of information processing according to the third embodiment will be described with reference to Fig. 23. Fig. 23 is a diagram showing an example of information processing according to the third embodiment.
[0262] 23, it is assumed that the information processing device 100B has already acquired a graph GR21 as shown in the spatial information SP21. Note that the information processing device 100B may generate the graph GR21 by appropriately using various techniques related to graph generation. The graph GR21 is similar to the graph GR21 shown in FIG. 12 and the like, and therefore description thereof will be omitted.
[0263] The information processing device 100B generates information indicating blobs BL1 to BL10, etc., by an arbitrary method. The information processing device 100B may classify each node into one of blobs BL1 to BL10 by an arbitrary clustering method such as k-means. Note that the information processing device 100B may generate blobs by any method as long as it is possible to generate blobs. For example, the information processing device 100B may generate information indicating blobs BL1 to BL10, etc., by any of the first to fourth methods described above.
[0264] The information processing device 100B generates a blob connection graph using the graph GR21 and blob information as shown in the spatial information SP21 (step S61). In the example of FIG. 23, the information processing device 100B generates a graph GR41, which is a blob connection graph in which a blob (first blob) is connected to another blob (second blob) by an edge as shown in the spatial information SP41. The arrow line shown in FIG. 23 indicates a directed edge from the blob at the source of the arrow (first blob) to the blob at the tip of the arrow (second blob). In other words, the arrow line shown in FIG. 23 indicates a directed edge in which the blob at the source of the arrow is the reference source and the blob at the tip of the arrow is the reference destination. For example, FIG. 23 shows that edges connect blob BL1 to five blobs, namely, blob BL2, blob BL3, blob BL4, blob BL5, and blob BL8.
[0265] For example, the information processing device 100B searches the graph GR21 and generates a graph GR41 based on the search results. For example, the information processing device 100B performs the following processes (6-1) to (6-4) to generate the graph GR41, which is a blob-connected graph. Note that the following objects may be read as nodes corresponding to the objects.
[0266] (6-1): Randomly get object NX (select) (6-2): Search for k neighboring objects of object NX (6-3): Get the blob that each neighboring object belongs to (6-4): Generate an edge from the blob to which the object NX belongs to the acquired blob.
[0267] For example, in the process (6-1), the information processing device 100B randomly acquires one object NX from the plurality of objects stored in the object information storage unit 121.
[0268] For example, in process (6-2), the information processing device 100B performs a search process as shown in FIG. 11 using the graph GR21 with the object NX as the target (query), thereby extracting k objects as neighboring objects of the object NX.
[0269] For example, in the process (6-3), the information processing device 100B refers to the blob information stored in the blob information storage unit 125, and acquires the blob associated with the neighboring object (node) of the object NX.
[0270] For example, in process (6-4), the information processing device 100B generates an edge from the blob to which the object NX belongs to the acquired blob. The information processing device 100B generates edges from the blob to which the object NX belongs to the blobs to which all of the neighboring objects belong. For example, the information processing device 100B generates a graph GR41 by registering information (blob ID) indicating the blob acquired as a reference destination of the blob to which the object NX belongs in association with information (blob ID) indicating the blob to which the object NX belongs in the blob connection graph information storage unit 126.
[0271] For example, the information processing device 100B repeats the above processes (6-1) to (6-4) until there are no more objects to be processed, and generates the graph GR41. Note that when blob Y has already been registered as a reference destination of blob X, even if an object belonging to blob Y is searched (extracted) again as a neighboring object of an object belonging to blob X, the information processing device 100B may skip the registration and prevent the same blob from being registered as a reference destination of each blob more than once.
[0272] For example, the information processing device 100B performs search processing using a graph GR41 that connects blobs with edges, as shown in the spatial information SP41. For example, the information processing device 100B performs search processing as shown in Fig. 27 and Fig. 28 using the graph GR41, which is a blob-connected graph, for the search query QE2, thereby obtaining search results for the search query QE2.
[0273] The above-described method for generating a blob connection graph is merely an example, and the information processing device 100B may generate a blob connection graph using various methods. For example, the information processing device 100B may generate a blob connection graph without searching a graph. If the blobs indicated in the spatial information SP21 are generated by searching, the search results from the generation of the blobs may be used to generate a blob connection graph such as graph GR41.
[0274] For example, the information processing device 100B may convert the connection relationships of object nodes in graph GR21 into connection relationships of blobs to which the object nodes belong, to generate a blob connection graph. For example, an edge connects node N1 belonging to blob BL9 to node N88 belonging to blob BL8 and a node belonging to blob BL10, and an edge connects node N4 belonging to blob BL9 to a node belonging to blob BL7. Therefore, the information processing device 100B generates a blob connection graph in which edges connect blob BL9 to the three blobs, blob BL7, blob BL8, and blob BL10. In this way, the information processing device 100B may generate a blob connection graph in which edges connect blobs based on edges connecting object nodes.
[0275] 3-2. Configuration of information processing device Next, the configuration of an information processing device 100B according to the third embodiment will be described with reference to Fig. 24. Fig. 24 is a diagram showing an example of the configuration of an information processing device according to the third embodiment. As shown in Fig. 24, the information processing device 100B has a communication unit 110, a storage unit 120B, and a control unit 130B. Note that, in the information processing device 100B, descriptions of the same aspects as those of the information processing device 100 or the information processing device 100A will be omitted as appropriate.
[0276] (Storage unit 120B) The storage unit 120B is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in Fig. 24 , the storage unit 120B according to the third embodiment includes an object information storage unit 121, a graph information storage unit 122, a quantization information storage unit 123, a codebook information storage unit 124, a blob information storage unit 125, and a blob connection graph information storage unit 126.
[0277] (Blob connection graph information storage unit 126) The blob connection graph information storage unit 126 according to the third embodiment stores various information related to the blob connection graph. For example, the blob connection graph information storage unit 126 stores the generated blob connection graph. FIG. 25 is a diagram illustrating an example of the blob connection graph information storage unit according to the third embodiment. The blob connection graph information storage unit 126 shown in FIG. 25 has items such as "blob ID" and "connected blob information."
[0278] "Blob ID" indicates identification information for identifying each blob (group) in the graph. "Connected blob information" indicates information about blobs (referenced blobs) that can be traced from the corresponding blob. For example, "connected blob information" includes information such as "reference destination." "Reference destination" indicates information for identifying a reference destination (blob) that is connected by an edge and can be traced from that blob. That is, in the example of Figure 25, a blob ID that identifies a blob is registered in association with a reference destination (blob) that can be traced from that blob by an edge. Note that "connected blob information" may also include information (edge ID) for identifying an edge connected to a reference destination.
[0279] 25 shows that edges connect a blob (blob BL1) identified by blob ID "BL1" to five blobs identified by blob IDs "BL2," "BL3," "BL4," "BL5," and "BL8." In other words, it shows that it is possible to trace from blob BL1 to each of the five blobs BL2, BL3, BL4, BL5, and BL8.
[0280] The blob connection graph information storage unit 126 may store various types of information depending on the purpose, without being limited to the above. The blob connection graph may include a program module that receives a query as input, searches for objects (object nodes) by tracing edges in the blob connection graph, and extracts and outputs objects similar to the query. That is, the blob connection graph may be intended for use as a program module that performs search processing using the blob connection graph. For example, the graph GR41, which is a blob connection graph, may be a program that, when vector data is input as a query, extracts and outputs objects corresponding to vector data similar to the vector data using the blob connection graph. For example, the graph GR41 may be data used as a program module that searches for similar images corresponding to a query image. For example, the graph GR41 causes a computer to function to extract and output objects similar to the query in the blob connection graph based on an input query.
[0281] (control unit 130B) 24, the control unit 130B is a controller, and is realized by, for example, a CPU, an MPU, a GPU, etc. executing various programs (corresponding to examples of information processing programs) stored in a storage device inside the information processing device 100B using a RAM as a work area. The control unit 130B is also a controller, and is realized by, for example, an integrated circuit such as an ASIC or an FPGA.
[0282] 24, control unit 130B has an acquisition unit 131, a generation unit 132B, a search processing unit 133B, and a provision unit 134, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of control unit 130B is not limited to the configuration shown in FIG. 24, and may be any other configuration as long as it performs the information processing described below.
[0283] An acquisition unit 131 according to the third embodiment acquires blob information indicating a plurality of blobs into which a plurality of objects to be the target of data search are classified, and index information indicating an index for searching the plurality of objects. The acquisition unit 131 acquires index information indicating a graph in which a plurality of nodes corresponding to each of the plurality of objects are connected by edges. The acquisition unit 131 acquires blob information indicating a plurality of blobs into which a plurality of objects are classified by a clustering process. The acquisition unit 131 acquires blob information indicating a plurality of blobs in which objects belonging to each blob are mutually exclusive. The acquisition unit 131 acquires a query.
[0284] (Generation unit 132B) The generating unit 132B generates various types of information in the same manner as the generating unit 132 or the generating unit 132A.
[0285] Generator 132B generates a blob connection graph in which edges are connected from a first blob to which one of the multiple objects belongs to a second blob, which is a blob to which a neighboring object of the selected object belongs, using the index information acquired by acquirer 131. Generator 132B selects one object from the multiple objects, searches for a previous neighboring object of the selected object among the multiple objects using the index information, and generates a blob connection graph by connecting edges from the first blob to which the selected object belongs to the second blob to which the neighboring object belongs.
[0286] The generation unit 132B randomly selects one object from the plurality of objects, searches for neighboring objects of the selected object, and generates a blob connection graph by connecting edges from a first blob to which the selected object belongs to a second blob to which the neighboring objects belong. The generation unit 132B uses the graph to generate a blob connection graph in which edges are connected from the first blob to the second blob.
[0287] The generation unit 132B uses the graph to search for neighboring objects of one object, and generates a blob connection graph by connecting an edge from a first blob to which the one object belongs to a second blob to which the neighboring object belongs. The generation unit 132B generates a blob connection graph by connecting an edge from the first blob to a second blob to which a neighboring object belongs that corresponds to a neighboring node that is a node connected by an edge to one node corresponding to the one object in the graph.
[0288] The generation unit 132B generates a blob connection graph by connecting, with an edge, a first representative point that is a representative point of the first blob and a second representative point that is a representative point of the second blob from the first blob. The generation unit 132B generates a blob connection graph by connecting, with an edge, the first representative point that is a centroid of the first blob and a second representative point that is a centroid of the second blob. The generation unit 132B generates a blob connection graph by connecting, with an edge, the first representative point that is a center point calculated based on objects belonging to the first blob and the second representative point that is a center point calculated based on objects belonging to the second blob.
[0289] (Search processing unit 133B) The search processing unit 133B performs various processes related to search processing in the same manner as the search processing unit 133 or the search processing unit 133A.
[0290] The search processing unit 133B performs a search process using the blob connection graph generated by the generation unit 132B. The search processing unit 133B performs a search process to search for nearby objects of a search query using the blob connection graph. The search processing unit 133B performs a search process to search for nearby objects of a search query by tracing edges that connect blobs in the blob connection graph.
[0291] In the search process, the search processing unit 133B uses blob information indicating multiple blobs to calculate the distance between a target object, which is an object belonging to one blob, and the search query using vector information of the vector-quantized target object. The search processing unit 133B calculates the distance between the target object and the search query in parallel and in a batch. The search processing unit 133B calculates the distance between the search query and the target object for a batch processing number determined based on the specifications of the information processing device in parallel.
[0292] [3-3. Information processing flow] Next, the procedure of information processing according to the third embodiment will be described with reference to Fig. 26. Fig. 26 is a flowchart showing an example of information processing according to the third embodiment.
[0293] 26, the information processing device 100B acquires group information indicating a plurality of groups into which a plurality of objects to be the target of data search are classified (step S501). For example, the information processing device 100B acquires blob information indicating a plurality of blobs into which nodes (object nodes) corresponding to a plurality of objects to be the target of data search are classified.
[0294] The information processing device 100B acquires index information indicating an index for searching multiple objects (step S502). For example, the information processing device 100B acquires, as an index, a graph in which nodes (object nodes) corresponding to multiple objects are connected by edges.
[0295] The information processing device 100B uses the index information to generate a group connection graph in which edges are connected from a first group to which one of the multiple objects belongs to to a second group, which is a group to which a neighboring object of the one object belongs (step S503). For example, the information processing device 100B uses the index information to generate a blob connection graph in which edges are connected from a first blob to which one of the multiple objects belongs to to a second blob, which is a blob to which a neighboring object of the one object belongs.
[0296] [3-4. Search processing example] Here, an example of the search process according to the third embodiment will be described using Fig. 27 and Fig. 28 as an example. Fig. 27 and Fig. 28 are flowcharts showing an example of the search process according to the third embodiment. The search process described below is performed by the search processing unit 133B of the information processing device 100B. Note that the description of the same points as those in the search processes described in the first and second embodiments will be omitted as appropriate.
[0297] 27 and 28, N(G, s) and C are a set of blobs (blob set). Also, "G" may be graph data (blob connection graph) in which blobs are connected by edges (for example, graph GR41 shown in spatial information SP41). For example, the information processing device 100B executes a k-nearest neighbor search process.
[0298] First, the processing (main processing) shown in Fig. 27 will be described. For example, the information processing device 100B sets the radius r of the hypersphere to ∞ (infinity) and sets the blob set C to an empty set (Φ) (step S601). Then, the information processing device 100B extracts a partial blob set B from the existing blob set (all blobs) (step S602). For example, the information processing device 100B may extract a blob to which an object (node) selected as the root node belongs as the partial blob set B. Alternatively, for example, the information processing device 100B may randomly extract blobs as the partial blob set B.
[0299] Then, the information processing device 100B executes the determination process (step S603). The information processing device 100B executes the determination process of steps S701 to S717 shown in Fig. 28. The details of the determination process shown in Fig. 28 will be described later.
[0300] Then, the information processing device 100B sets a partial blob set B in the blob set S (step S604).
[0301] The information processing device 100B selects a blob s having the smallest query distance d from among the blobs included in the blob set S (step S605). Note that the "query distance" here refers to the shortest distance between all objects in the blob and the query object (query). For example, if an object that is a query (search query) is y, the information processing device 100B selects a blob s from the blob set S having the shortest query distance to object y. Note that the query distance is not limited to the above, and may be, for example, the distance between the query and a representative point (such as a centroid) of the blob. Then, the information processing device 100B excludes the blob s from the blob set S (step S606).
[0302] Then, the information processing device 100B determines whether the query distance d of the blob s exceeds r(1+ε) (step S607). If the query distance d of the blob s exceeds r(1+ε) (step S607: Yes), the information processing device 100B outputs the object set R as a neighborhood object set of the object y (step S608), and ends the process.
[0303] If the query distance d of the blob s does not exceed r(1+ε) (step S607: No), the information processing device 100B determines whether the blob s is included in the blob set C (step S609).
[0304] If the blob s is included in the blob set C (step S609: Yes), the information processing device 100B returns to step S605 and repeats the process.
[0305] If the blob s is not included in the blob set C (step S609: No), the information processing device 100B adds the blob s to the blob set C (step S610).
[0306] Then, the information processing device 100B sets a neighborhood blob set N(G, s) of blob s in the partial blob set B (step S611). The neighborhood blob set N(G, s) is, for example, a set of blobs (neighboring blobs) associated with blob s. For example, the neighborhood blob set N(G, s) is a set of blobs (neighboring blobs) connected by edges from blob s. Then, the information processing device 100B executes a determination process (step S612). Details will be described later, but for example, the information processing device 100B executes the determination process of steps S701 to S717 shown in FIG. 28, similar to step S603 described above.
[0307] Then, the information processing device 100B determines whether the blob set S is an empty set (Φ) (step S613). If the blob set S is not an empty set (step S613: No), the information processing device 100B returns to step S605 and repeats the process. If the blob set S is an empty set (step S613: Yes), the information processing device 100B outputs the object set R and ends the process (step S614). For example, the information processing device 100B may select objects (nodes) included in the object set R as neighboring nodes corresponding to the added node (input object y). For example, the information processing device 100B may extract (select) objects (nodes) included in the object set R as neighboring nodes corresponding to the target node (input object y). Furthermore, for example, the information processing device 100B may provide the objects (nodes) included in the object set R as search results corresponding to the search query (input object y) to the terminal device or the like that performed the search.
[0308] Next, a description will be given of the process (determination process) shown in Fig. 28. First, the information processing device 100B sets a partial blob set B in the blob set T (step S701).
[0309] Then, the information processing device 100B acquires one blob b from the blob set T and deletes the blob b from the blob set T (step S702). Then, the information processing device 100B sets min, which is a variable used in the determination process, to ∞ (infinity) (step S703).
[0310] Then, the information processing device 100B acquires one object u from the blob b, and deletes the object u from the blob b (step S704).
[0311] Then, the information processing device 100B determines whether the distance d(u, y) between the object u and the object y is less than min (step S705).
[0312] If the distance d(u, y) between the object u and the object y is less than min (step S705: Yes), the information processing device 100B sets (updates) the value of min to the distance d(u, y) (step S706), and performs the process of step S707.
[0313] If the distance d(u, y) between the object u and the object y is not less than min (step S705: No), the information processing device 100B does not perform the process of step S706, but performs the process of step S707.
[0314] The information processing device 100B determines whether the distance d(u, y) between the object u and the object y is equal to or less than r(1+ε) (step S707).
[0315] If the distance d(u, y) between object u and object y is equal to or less than r(1+ε) (step S707: Yes), information processing device 100B adds blob b to blob set S (step S708) and performs the process of step S709.
[0316] If the distance d(u, y) between the object u and the object y is not equal to or less than r(1+ε) (step S707: No), the information processing device 100B does not perform the process of step S708, but performs the process of step S709.
[0317] The information processing device 100B determines whether the distance d(u, y) between the object u and the object y is equal to or less than r (step S709).
[0318] If the distance d(u, y) between object u and object y is equal to or less than r (step S709: Yes), information processing device 100B adds object u to object set R (step S710), and performs the process of step S711.
[0319] If the distance d(u, y) between the object u and the object y is not equal to or less than r (step S709: No), the information processing device 100B does not perform the process of step S710, but performs the process of step S711.
[0320] The information processing device 100B determines whether the number of objects included in the object set R exceeds ks (step S711). The predetermined number ks is a natural number that is determined arbitrarily. For example, ks may be the number of searches or the number of objects to be extracted. Furthermore, for example, when no upper limit is set on the number of objects to be extracted in a range search or the like, ks may be set to infinity. For example, ks may be 4.
[0321] If the number of objects included in object set R exceeds ks (step S711: Yes), information processing device 100B excludes from object set R the object that is farthest from object y among the objects included in object set R (step S712), and performs processing of step S713.
[0322] If the number of objects included in object set R does not exceed ks (step S711: No), information processing device 100B does not perform the process of step S712, but performs the process of step S713.
[0323] Information processing device 100B determines whether the number of objects included in object set R matches ks (step S713).
[0324] If the number of objects included in object set R matches ks (step S713: Yes), information processing device 100B sets the distance between object y and the object farthest from object y among the objects included in object set R to r (step S714), and performs processing of step S715.
[0325] If the number of objects included in object set R does not match ks (step S713: No), information processing device 100B does not perform the process of step S714, but performs the process of step S715.
[0326] The information processing device 100B determines whether or not the blob b is an empty set (Φ) (step S715). If the blob b is not an empty set (step S715: No), the information processing device 100B returns to step S704 and repeats the process.
[0327] Also, if blob b is an empty set (step S715: Yes), the information processing device 100B sets min as the query distance of blob b in partial blob set B (step S716). For example, the information processing device 100B sets min as the query distance, which is the shortest distance among the distances between all objects of blob b and the query.
[0328] The information processing device 100B determines whether the blob set T is an empty set (Φ) (step S717). If the blob set T is not an empty set (step S717: No), the information processing device 100B returns to step S702 and repeats the process.
[0329] If the blob set T is an empty set (step S717: Yes), the determination process ends.
[0330] [3-5. Modifications] The blob connection graph may be configured in various ways, not limited to the above-mentioned examples. For example, the blob connection graph may be a graph in which representative points of each group (blob) are connected by edges. This point will be described below as a modified example of the third embodiment. Note that descriptions of similar aspects to the first, second, and third embodiments will be omitted as appropriate. An information processing device 100B according to a modified example of the third embodiment has a blob information storage unit 125A and a blob connection graph information storage unit 126A instead of the blob information storage unit 125 and the blob connection graph information storage unit 126.
[0331] [3-5-1. Information Processing] First, an overview of information processing according to the modified example will be described with reference to Fig. 29. Fig. 29 is a diagram showing an example of blob connection graph information according to the modified example. The example of Fig. 29 shows a case where blobs are generated by clustering.
[0332] 29, the information processing device 100B generates a blob connection graph that connects the centroids of each blob. The information processing device 100B generates a graph GR42, which is a blob connection graph that connects the centroids C1 to C10 corresponding to each of the blobs BL1 to BL10 with edges, as shown in the spatial information SP41. The information processing device 100B generates the graph GR42 by appropriately using various conventional techniques related to graph generation.
[0333] For example, the information processing device 100B generates a graph GR42 in which the centroid C1 of the blob BL1 is connected by edges to five centroids: the centroid C2 of the blob BL2, the centroid C3 of the blob BL3, the centroid C4 of the blob BL4, the centroid C7 of the blob BL7, and the centroid C8 of the blob BL8. Also, for example, the information processing device 100B generates a graph GR42 in which the centroid C2 of the blob BL2 is connected by edges to three centroids: the centroid C1 of the blob BL1, the centroid C3 of the blob BL3, and the centroid C4 of the blob BL4.
[0334] The information processing device 100B may generate the graph GR42 by any method. The information processing device 100B may generate the graph GR42 by using various indexes. For example, the information processing device 100B may generate the graph GR42 by using the graph GR21 in which object nodes are linked. In this case, the information processing device 100B can generate the graph GR42 by the same processing as for the graph GR41 in FIG. 23, so a detailed description will be omitted. For example, the information processing device 100B may generate the graph GR42 by the same processing as the example shown in FIG. 23. Although the graph GR42 shows undirected (bidirectional) edges as an example, directed edges may also be used.
[0335] In this way, the information processing device 100B according to the modified example generates a centroid graph as a blob-linked graph. Then, the information processing device 100B performs a normal search using the centroids by using the generated blob-linked graph, and performs distance calculations on a blob-by-blob basis. In this way, the information processing device 100B can simplify processing by representing blobs with centroids. Note that, if the blobs are not clustered, the information processing device 100B may calculate a center point based on the objects belonging to each blob and use the calculated center point as the representative point.
[0336] In the example of Fig. 29, the information processing device 100B performs search processing using a graph GR42 in which centroids, which are representative points of blobs, are connected by edges, as shown in spatial information SP42. For example, the information processing device 100B performs search processing as shown in Figs. 27 and 28 using the graph GR21 for the search query QE2, thereby obtaining search results for the search query QE2. This is similar to the example described above, and detailed description thereof will be omitted.
[0337] [3-5-2. Information] Next, an overview of information according to the modified example will be described with reference to Fig. 30 and Fig. 31. Fig. 30 is a diagram showing an example of a blob information storage unit according to the modified example. Fig. 31 is a diagram showing an example of a blob connection graph information storage unit according to the modified example.
[0338] First, a description will be given of blob information storage unit 125A according to a modified example shown in Fig. 30. Blob information storage unit 125A according to the modified example shown in Fig. 30 includes items such as "blob ID," "node ID," "vector information," and "centroid ID." In this way, blob information storage unit 125A according to the modified example differs from blob information storage unit 125 of Fig. 15 in that it includes a "centroid ID." Note that descriptions of similar points to blob information storage unit 125 of Fig. 15 will be omitted where appropriate.
[0339] "Centroid ID" indicates identification information for identifying the centroid of each blob. "Vector information" indicates vector information of the centroid, which is the representative point of the blob.
[0340] 30, the centroid of blob BL1 identified by blob ID "BL1" is the centroid (centroid C1) identified by centroid ID "C1". The centroid of blob BL9 identified by blob ID "BL9" is the centroid (centroid C9) identified by centroid ID "C9".
[0341] The blob information storage unit 125A is not limited to the above, and may store various types of information depending on the purpose.
[0342] Next, a description will be given of a blob connection graph information storage unit 126A according to a modified example shown in Fig. 31. The blob connection graph information storage unit 126A according to the modified example shown in Fig. 30 has items such as "centroid ID" and "connection centroid information." Note that a description of the same points as those in the blob connection graph information storage unit 126 of Fig. 25 will be omitted as appropriate.
[0343] "Centroid ID" indicates identification information for identifying a centroid, which is the representative point of each blob in a graph. "Connected centroid information" indicates information about centroids (reference centroids) that can be traced from the corresponding centroid. For example, "connected centroid information" includes information such as "reference destination." "Reference destination" indicates information for identifying a reference destination (centroid) that is connected by an edge and can be traced from that centroid. That is, in the example of Figure 31, a centroid ID that identifies a centroid is registered in association with a reference destination (centroid) that can be traced from that centroid by an edge. Note that "connected centroid information" may also include information (edge ID) for identifying an edge connected to a reference destination.
[0344] 31 shows that edges are connected from the centroid (centroid C1) identified by the centroid ID "C1" to five centroids identified by the centroid IDs "C2," "C3," "C4," "C7," and "C8." In other words, it shows that it is possible to trace from the centroid C1 to each of the five centroids, C2, C3, C4, C7, and C8.
[0345] The blob connection graph information storage unit 126A may store various types of information depending on the purpose, without being limited to the above. The blob connection graph may include a program module that receives a query as input, searches for objects (object nodes) by tracing edges in the blob connection graph, and extracts and outputs objects similar to the query. That is, the blob connection graph may be intended for use as a program module that performs search processing using the blob connection graph. For example, the graph GR42, which is a blob connection graph, may be a program that, when vector data is input as a query, extracts and outputs objects corresponding to vector data similar to the vector data using the blob connection graph. For example, the graph GR42 may be data used as a program module that searches for similar images corresponding to a query image. For example, the graph GR42 causes a computer to function to extract and output objects similar to the query in the blob connection graph based on an input query.
[0346] As described above, in each embodiment and modification, the information processing system 1 executes the following process. For example, the information processing system 1 generates blobs by clustering vectors. For example, the information processing system 1 generates a graph index using centroids of blobs as nodes.
[0347] For example, the information processing system 1 acquires a certain number of arbitrary vectors and performs clustering by direct product quantization. That is, the information processing system 1 divides the vectors into partial vectors, performs clustering for each partial vector, and generates a codebook.
[0348] For example, the information processing system 1 performs Cartesian product quantization on objects belonging to blob units and generates an inverted index. For example, the information processing system 1 generates a graph (blob graph) including edges from objects to blobs. For example, the information processing system 1 searches for nearby objects using a general Cartesian product quantization method. The general Cartesian product quantization method here may involve, for example, calculating the distances between objects and all centroids of blobs to obtain a certain number of blobs, and then using Cartesian product quantization to obtain the nearby objects. Note that, if the information processing system 1 has generated a centroid graph as described above, it can use the graph to search for and obtain a certain number of blobs. Alternatively, the information processing system 1 may generate a graph (blob connection graph) in which blobs are connected by edges, rather than the graph (blob graph) including edges from objects to blobs described above.
[0349] 4. Fourth Embodiment In the above example, a case where blobs are used as an example of groups for classifying objects is shown, but the groups for classifying objects are not limited to blobs, and various graphs for classifying objects may be used. For example, in addition to blobs (first group), clusters may be used as second groups that are units for quantizing objects.
[0350] This point will be described below as a fourth embodiment. In the fourth embodiment, the information processing system 1 has an information processing device 100C instead of the information processing device 100, the information processing device 100A, or the information processing device 100B. Note that the description of the same points as those described above in the first, second, and third embodiments will be omitted as appropriate.
[0351] The information processing device 100C generates blob information (first group information) indicating a plurality of blobs that classify a plurality of objects, and cluster information (second group information) indicating clusters that are groups classified differently from the blobs.
[0352] [4-1. Information Processing] First, an overview of information processing according to the fourth embodiment will be described with reference to Fig. 32. Fig. 32 is a diagram showing an example of information processing according to the fourth embodiment. Note that Fig. 32 shows a case in which the information processing device 100C classifies objects into a plurality of blobs and then classifies them into clusters based on the blobs, but the objects may be classified into clusters and then into blobs, which will be described later.
[0353] First, the information processing device 100C generates blob information that classifies multiple objects into multiple blobs (step S71). In the example of FIG. 32, the information processing device 100C generates blob information that classifies multiple objects into blobs BL1 to BL10, etc., using an arbitrary method. The information processing device 100C may classify each node into one of the blobs BL1 to BL10 using an arbitrary clustering method such as k-means. Note that the information processing device 100C may generate blobs using any method as long as it is capable of generating blobs. For example, the information processing device 100C may generate blobs using an index of a graph, etc. For example, the information processing device 100C may generate information indicating blobs BL1 to BL10, etc., using any of the first to fourth methods described above.
[0354] 32, the information processing device 100C generates a graph GR41 that is a blob-connected graph as shown in the spatial information SP41, but the graph GR41 does not have to be generated. For example, when generating the graph GR41, the information processing device 100C generates the graph GR41 by processing using the same information as in FIG. 23, but detailed description thereof will be omitted because it is the same as the description in FIG. 23.
[0355] Then, the information processing device 100C generates cluster information that classifies the multiple blobs into multiple clusters (step S72). For example, the information processing device 100C generates cluster information indicating clusters into which multiple blobs, such as blobs BL1 to BL10, are classified using an arbitrary method. The information processing device 100C may classify the blobs into one of the multiple clusters, such as blobs BL1 to BL10, using an arbitrary clustering method such as k-means. Note that the information processing device 100C may generate blobs using any method as long as it is capable of generating blobs. For example, the information processing device 100C may generate blobs using an index such as a graph.
[0356] In the example of Fig. 32, the information processing device 100C generates cluster information for classifying blobs BL1 to BL10 into clusters CL1 to CL4. For example, the information processing device 100C classifies blobs BL1 and BL7 into cluster CL1. That is, the information processing device 100C classifies objects belonging to blob BL1 and objects belonging to blob BL7 into cluster CL1. Note that although Figs. 32 and 34 show that the boundaries of blobs are on the boundaries of clusters, they may not coincide depending on the division method, and the space may not necessarily be divided by the boundaries depending on the method.
[0357] Furthermore, the information processing device 100C classifies blob BL2 and blob BL4 into cluster CL2. That is, the information processing device 100C classifies objects belonging to blob BL2 and objects belonging to blob BL4 into cluster CL2.
[0358] Furthermore, the information processing device 100C classifies blobs BL3, BL8, and BL10 into cluster CL3. That is, the information processing device 100C classifies objects belonging to blob BL3, objects belonging to blob BL8, and objects belonging to blob BL10 into cluster CL3.
[0359] Furthermore, the information processing device 100C classifies blobs BL5, BL6, and BL9 into cluster CL4. That is, the information processing device 100C classifies objects belonging to blob BL5, objects belonging to blob BL6, and objects belonging to blob BL9 into cluster CL4.
[0360] For example, the information processing device 100C executes a process related to quantization using cluster information. The information processing device 100C executes a process related to quantization for a plurality of objects using the cluster information. For example, the information processing device 100C executes quantization for each object belonging to each of clusters CL1 to CL4, which will be described later.
[0361] As described above, the information processing device 100C generates blob information indicating a plurality of blobs for classifying a plurality of objects to be searched for data, and cluster information indicating a plurality of clusters for classification different from the plurality of blobs. This allows the information processing device 100C to generate two types of groups for classifying a plurality of objects, and to generate groups for classifying objects.
[0362] [4-1-1. Other processing examples] The process shown in FIG. 32 is merely an example, and the information processing device 100C may perform generation by any mode of process as long as it is capable of generating blob information and cluster information.
[0363] For example, the information processing device 100C may generate cluster information for classifying objects into clusters, and then generate blob information for classifying objects into blobs. In this case, the information processing device 100C may generate cluster information for classifying each object into one of multiple clusters by clustering that classifies multiple objects into clusters. For example, the information processing device 100C may classify each object into one of clusters CL1 to CL4 by any clustering such as k-means.
[0364] The information processing device 100C may then divide at least one of the clusters CL1 to CL4 and generate two or more blobs, thereby generating blob information indicating the blobs. For example, the information processing device 100C may cluster objects belonging to each cluster using any clustering method such as k-means, thereby dividing each cluster and generating blobs. For example, the information processing device 100C may divide cluster CL1 into two blobs, thereby generating blobs BL1 and BL7. For example, the information processing device 100C may cluster objects belonging to cluster CL1 using k-means with the number of clusters set to 2, divide cluster CL1, and generate blobs BL1 and BL7.
[0365] For example, the information processing device 100C may generate blobs BL2 and BL4 by dividing cluster CL2 into two blobs. For example, the information processing device 100C may generate blobs BL3, BL8, and BL10 by dividing cluster CL3 into three blobs. For example, the information processing device 100C may generate blobs BL5, BL6, and BL9 by dividing cluster CL4 into three blobs.
[0366] 4-2. Configuration of information processing device Next, the configuration of an information processing device 100C according to the fourth embodiment will be described with reference to Fig. 33. Fig. 33 is a diagram showing an example of the configuration of an information processing device according to the fourth embodiment. As shown in Fig. 33, the information processing device 100C has a communication unit 110, a storage unit 120C, and a control unit 130B. Note that, in the information processing device 100C, descriptions of the same aspects as those of the information processing device 100, the information processing device 100A, or the information processing device 100B will be omitted as appropriate.
[0367] (Storage unit 120C) The storage unit 120C is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in Fig. 33 , the storage unit 120C according to the fourth embodiment includes an object information storage unit 121, a graph information storage unit 122, a quantization information storage unit 123, a codebook information storage unit 124, a blob information storage unit 125, a blob connection graph information storage unit 126, and a cluster information storage unit 127.
[0368] The cluster information storage unit 127 according to the fourth embodiment stores information about clusters that are the second group. The cluster information storage unit 127 stores information about clusters that are units of quantization. The cluster information storage unit 127 may store identification information (such as a cluster ID) that identifies each cluster in association with information (such as a blob ID) that indicates blobs that are the first group and belong to that cluster. The cluster information storage unit 127 may also store identification information (such as a cluster ID) that identifies each cluster in association with information (such as an object ID) that indicates objects that belong to that cluster. The cluster information storage unit 127 may also store identification information (such as a cluster ID) that identifies each cluster in association with information (such as a base point ID) for identifying a base point such as a centroid of the cluster and its vector.
[0369] (control unit 130C) The control unit 130C is a controller, and is realized by, for example, a CPU, an MPU, a GPU, etc., executing various programs (corresponding to examples of information processing programs) stored in a storage device inside the information processing device 100C using RAM as a work area. The control unit 130C is also a controller, and is realized by, for example, an integrated circuit such as an ASIC or FPGA.
[0370] 33, control unit 130C has an acquisition unit 131, a generation unit 132C, a search processing unit 133C, and a provision unit 134, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of control unit 130C is not limited to the configuration shown in Fig. 33, and may be any other configuration as long as it performs the information processing described below.
[0371] The acquiring unit 131 according to the fourth embodiment acquires object information indicating a plurality of objects to be searched for data, and acquires index information indicating an index for searching a plurality of objects.
[0372] The acquisition unit 131 acquires processing information indicating a plurality of blobs into which a plurality of objects to be subjected to data search are classified, and a plurality of clusters, which are groups classified differently from the plurality of blobs and into which objects belonging to the same blob among the plurality of objects are classified into the same group. The acquisition unit 131 acquires processing information indicating a first number of blobs into which the plurality of objects are classified, and a cluster into which the plurality of objects are classified into a second number smaller than the first number. The acquisition unit 131 acquires processing information indicating a plurality of clusters into which object groups belonging to two or more blobs among the plurality of blobs are classified into one cluster.
[0373] The acquisition unit 131 acquires processing information indicating a plurality of blobs to be used for performing a first process, which is a batch process in a search process targeting a plurality of objects, and acquires processing information indicating a plurality of clusters to be used for performing a second process, which is a process related to quantization targeting a plurality of objects.
[0374] The acquisition unit 131 acquires processing information in which a plurality of pieces of first identification information that identify a plurality of blobs are associated with a plurality of pieces of second identification information that identify a plurality of clusters. The acquisition unit 131 acquires processing information in which one piece of second identification information is associated with two or more pieces of first identification information.
[0375] The acquisition unit 131 acquires processing information associated with first identification information of a blob to which each of the plurality of objects belongs. The acquisition unit 131 acquires processing information in which quantized object information obtained by quantizing each of the plurality of objects is associated with the first identification information. The acquisition unit 131 acquires processing information in which quantized object information generated by vector quantization, which quantizes each of partial vectors obtained by dividing a vector corresponding to each of the plurality of objects by product quantization, is associated with the first identification information. The acquisition unit 131 acquires processing information in which quantized object information generated based on partial centroids corresponding to partial regions of each centroid corresponding to the quantization is associated with the first identification information. The acquisition unit 131 acquires processing information in which quantized object information generated based on residual vectors generated by each centroid corresponding to the quantization and each partial centroid corresponding to each partial region of each centroid is associated with the first identification information.
[0376] (Generation unit 132C) The generating unit 132C generates various types of information in the same manner as the generating unit 132, the generating unit 132A, or the generating unit 132B.
[0377] The generation unit 132C generates blob information indicating a plurality of blobs into which a plurality of objects are classified, based on the object information acquired by the acquisition unit 131. The generation unit 132C generates cluster information indicating a plurality of clusters, which are groups classified differently from the plurality of blobs and which classify, among the plurality of objects, objects that belong to the same blob among the plurality of blobs into the same group.
[0378] Generator 132C generates blob information that classifies a plurality of objects into a first number of blobs, and cluster information that classifies a plurality of objects into a second number of clusters that is less than the first number. Generator 132C generates cluster information that indicates a plurality of clusters in which object groups belonging to two or more blobs among the plurality of blobs are classified into one cluster.
[0379] The generation unit 132C generates blob information indicating a plurality of blobs used for performing batch processing in search processing targeting a plurality of objects. The generation unit 132C generates blob information indicating a plurality of blobs used for performing batch calculation of distances related to a plurality of objects. The generation unit 132C generates blob information indicating a plurality of blobs used for performing parallel processing of calculation of distances between a search query used in search processing and an object.
[0380] The generation unit 132C generates cluster information indicating a plurality of clusters used for performing processing related to quantization on a plurality of objects. The generation unit 132C generates cluster information indicating a plurality of clusters that serve as units of quantization. The generation unit 132C generates cluster information indicating a plurality of clusters used for quantizing vectors corresponding to each of a plurality of objects.
[0381] After classifying the multiple blobs, the generation unit 132C classifies the multiple clusters. The generation unit 132C generates blob information indicating the multiple blobs by clustering the multiple objects. The generation unit 132C generates blob information indicating the multiple blobs using index information. The generation unit 132C generates cluster information indicating the multiple clusters using the blob information indicating the blobs. The generation unit 132C generates cluster information indicating the multiple clusters by clustering the multiple blobs.
[0382] After classifying the objects into the multiple clusters, the generation unit 132C classifies the multiple blobs. The generation unit 132C generates cluster information indicating the multiple clusters by clustering the multiple objects. The generation unit 132C generates blob information indicating the multiple blobs using the cluster information indicating the clusters. The generation unit 132C divides at least one of the multiple clusters to generate two or more blobs, thereby generating blob information indicating the multiple blobs. The generation unit 132C generates first group information indicating the multiple first groups by clustering the objects belonging to each second group and dividing each second group into two or more first groups.
[0383] (Search processing unit 133C) The search processing unit 133C functions as a processing unit that executes a first process for objects belonging to a blob and a second process for objects belonging to a cluster. The search processing unit 133C performs various processes related to the search process in the same way as the search processing unit 133, the search processing unit 133A, or the search processing unit 133B. The search processing unit 133C performs processing using the information generated by the generation unit 132C.
[0384] The search processing unit 133C executes a first process targeting objects belonging to a blob and a second process targeting objects belonging to a cluster, using the processing information acquired by the acquisition unit 131. The search processing unit 133C executes the second process targeting a group of objects belonging to one cluster. The search processing unit 133C executes the first process targeting a group of objects belonging to one blob.
[0385] The search processing unit 133C executes a first process that collectively calculates the distance for a group of objects belonging to one blob. The search processing unit 133C executes the first process, which is a parallel process of calculating the distance between each of the group of objects belonging to one blob and a search query used in the search process.
[0386] The search processing unit 133C executes a second process on the group of objects belonging to one cluster. The search processing unit 133C executes a second process, which is quantization, on the group of objects belonging to one cluster. The search processing unit 133C executes a second process that quantizes vectors corresponding to each of the group of objects belonging to one cluster.
[0387] When two or more blobs belong to one cluster, the search processing unit 133C executes the second process on a group of objects belonging to the two or more blobs. When two or more blobs belong to one cluster, the search processing unit 133C executes quantization on all objects belonging to the two or more blobs.
[0388] The search processing unit 133C executes a first process and a second process by referring to the processing information. The search processing unit 133C identifies an object to be processed by referring to the processing information. In a search process targeting a plurality of objects, the search processing unit 133C performs the first process on a blob of one of the first identification information among two or more pieces of first identification information, and generates reference information to be referenced in the search process for a cluster of one of the second identification information. After that, when performing the first process on other blobs of other first identification information other than the one first identification information among the two or more pieces of first identification information, no reference information is generated.
[0389] The search processing unit 133C performs search processing using the blob connection graph. The search processing unit 133C performs search processing to search for nearby objects of a search query using the blob connection graph. The search processing unit 133C performs search processing to search for nearby objects of a search query by tracing edges that connect blobs in the blob connection graph.
[0390] In the search process, the search processing unit 133C uses blob information indicating multiple blobs to calculate the distance between a target object, which is an object belonging to one blob, and the search query using vector information of the vector-quantized target object. The search processing unit 133C calculates the distance between the target object and the search query in parallel and collectively. The search processing unit 133C calculates the distance between the search query and the target object for a batch processing number determined based on the specifications of the information processing device in parallel.
[0391] The search processing unit 133C performs search processing using the blob connection graph as shown in Figures 27 and 28. For example, the search processing unit 133C performs search processing using graph GR41 that connects blobs with edges, as shown in spatial information SP51. For example, the search processing unit 133C performs search processing as shown in Figures 27 and 28 using graph GR41, which is a blob connection graph, for search query QE3, to obtain search results for search query QE3. For example, the search processing unit 133C performs a first process on the blob that is the processing target by tracing the blob connection graph.
[0392] The search processing unit 133C generates reference information when reference information has not been generated for the cluster to which the processing target blob belongs. For example, the search processing unit 133C generates a distance lookup table (LUT) as reference information. For example, the search processing unit 133C generates a distance lookup table for each cluster in the search process. That is, the search processing unit 133C generates a distance lookup table for each cluster in the search process, rather than for each blob. This allows the information processing device 100C to suppress an increase in search time.
[0393] The search processing unit 133C uses the generated reference information to perform a distance calculation for all objects belonging to the blob being processed. For example, if reference information has already been generated for the cluster to which the blob being processed belongs, the search processing unit 133C does not generate reference information. Then, the search processing unit 133C uses the already generated reference information to perform a distance calculation for all objects belonging to the blob being processed.
[0394] [4-3. Examples of processing and information] Next, examples of processing and information used in processing in the fourth embodiment will be described.
[0395] First, the concept of quantization will be described with reference to Fig. 34. Fig. 34 is a conceptual diagram of quantization according to the fourth embodiment.
[0396] For example, base point CL1 indicates the base point used for quantizing cluster CL1. In this case, base point CL1 indicates the centroid of cluster CL1 that includes blobs BL1 and BL7. Also, for example, base point CL2 indicates the base point used for quantizing cluster CL2. In this case, base point CL2 indicates the centroid of cluster CL2 that includes blobs BL2 and BL4.
[0397] For example, base point CL3 indicates the base point used for quantizing cluster CL3. In this case, base point CL3 indicates the centroid of cluster CL3, which includes blobs BL3, BL8, and BL10. Also, for example, base point CL4 indicates the base point used for quantizing cluster CL4. In this case, base point CL4 indicates the centroid of cluster CL4, which includes blobs BL5, BL6, and BL9.
[0398] In this case, the information processing device 100C finds a residual vector in each of the four clusters CL1 to CL4. Then, the information processing device 100C performs quantization using the residual vector found in each of the four clusters CL1 to CL4. For example, the information processing device 100C performs quantization using a residual vector from the base point CL1 for an object belonging to cluster CL1. Although a detailed description of quantization using a residual vector will be omitted, the information processing device 100C performs quantization using a residual vector by appropriately using, for example, a technique disclosed in Patent Document 3.
[0399] 34 shows a state in which the vector is not divided into partial vectors for the purpose of illustration, but the information processing device 100C may derive a base point for each partial space divided by the Cartesian product quantization, and perform quantization using a residual vector based on the derived base point. For example, when the information processing device 100C divides a vector into four partial vectors and performs Cartesian product quantization, it has a base point for each partial space corresponding to each partial vector, and performs quantization using a residual vector from the base point.
[0400] An example of object quantization and quantization processing by the information processing device 100C will be described below with reference to Fig. 34. In Fig. 34, a case where the query object (search query) is search query QE3 will be described as an example. Note that, in the quantization and direct product quantization of objects (vectors) described below, detailed descriptions of points similar to those disclosed in Patent Document 3 will be omitted as appropriate.
[0401] As shown in the spatial information SP52 of FIG. 34, the distance between the search query QE3 and each vector corresponding to each object belonging to the cluster CL1, such as the node N7, is different, but the information processing device 100C does not use the vector corresponding to each object belonging to the cluster CL1, such as the node N7. In this case, the distance between the search query QE3 and each vector is regarded as, for example, the distance between the search query QE3 and the vector corresponding to the base point CL1 of the cluster CL1. That is, the vector corresponding to each object belonging to the cluster CL1 is quantized into a vector corresponding to the base point CL1. Specifically, the information processing device 100C obtains a residual vector, which is the difference between each vector and the base point CL1, and performs direct product quantization on the residual vector (further dividing into subspaces and quantizing).
[0402] For example, the information processing device 100C may further divide each cluster into a plurality of partial regions and generate partial centroids corresponding to the plurality of partial regions as a codebook. For example, the information processing device 100C may generate partial centroids corresponding to the plurality of partial regions into which each cluster is divided, based on information on the residual vectors related to each cluster.
[0403] For example, the information processing device 100C may determine the number of partial regions for each cluster based on the number of objects included in each cluster, or may determine the number of partial regions based on a predetermined setting value. For example, the information processing device 100C may determine the partial regions by appropriately using various conventional clustering techniques. For example, the information processing device 100C calculates a residual vector between the objects included in each partial region of cluster CL1 and the base point CL1 of cluster CL1, and generates a partial centroid for the partial region from the residual vector. For example, the information processing device 100C uses information on the residual vector between the base point of the cluster and the partial centroid for each partial region. For example, the information processing device 100C uses information on the residual vector between the base point CL1 of cluster CL1 and the partial centroid for each partial region. For example, the information processing device 100C stores information on the partial centroid of the partial region to which each object belongs, in association with each object.
[0404] For example, by using the above-described information, the information processing device 100C can quantize the position of each object to the position of the centroid of the partial region to which the vector belongs. This allows the information processing device 100C to refine the distance between the search query QE3 and each object to the distance between the search query QE3 and the vector corresponding to the partial centroid of the partial region to which each object belongs.
[0405] For example, the information processing device 100C uses information indicating the vector of the base point of each cluster and the residual vector related to the partial centroid. This allows the information processing device 100C to calculate the distance to the partial centroid for, for example, the search query QE3. Therefore, the information processing device 100C can determine, for example, the distance between the search query QE3 and each object as the distance to the partial centroid that is closer than the base point of the cluster.
[0406] The information processing device 100C may process the vector of each object by dividing it into a plurality of partial vectors. That is, the information processing device 100C may perform processing using a technique related to so-called direct product quantization. This is the same as in the first embodiment, and therefore a duplicated description will be omitted as appropriate.
[0407] For example, in the case of FIG. 34, the information processing device 100C divides the vector of the search query QE3 into a plurality of partial vectors (also referred to as "partial vectors of the search query QE3"). In this case, the spatial information SP52 is also divided into a plurality of subspaces (also referred to as "subspaces of the spatial information SP52"). For example, the information processing device 100C may divide the vector of the search query QE3 into four, as in FIG. 1. In this case, the spatial information SP52 is also divided into four subspaces, and each cluster is divided into a plurality of subspaces.
[0408] For example, the information processing device 100C stores information indicating a partial centroid (codebook) corresponding to the region to which each object belongs in the four subspaces, in association with each object. For example, the information processing device 100C calculates the distance between each object and the search query QE3 by summing the distances between the four codebooks corresponding to each object and the search query QE3.
[0409] Next, an overview of an inverted index will be explained with reference to Fig. 35 and Fig. 36. Fig. 35 is a diagram showing an example of an inverted index according to the fourth embodiment. Fig. 36 is a diagram showing a specific example of an inverted index according to the fourth embodiment.
[0410] First, the transposed index IND1 shown in Fig. 35 will be described. The transposed index IND1 shown in Fig. 35 includes items such as "blob ID," "cluster ID," "number of quantized objects," "quantized object #1," "quantized object #2," "quantized object #3," and "quantized object #4." Note that Fig. 35 illustrates four objects, "quantized object #1" to "quantized object #4," but the transposed index IND1 includes items such as "quantized object #5," "quantized object #6," and so on, the number of which corresponds to the number of objects belonging to each blob.
[0411] "Blob ID" indicates identification information for identifying a blob, which is the first group. "Cluster ID" indicates identification information for identifying a cluster, which is the second group. "Number of quantized objects" indicates the number of objects belonging to a blob. "Quantized object #1" to "Quantized object #4" indicate quantized objects. For example, "Quantized object #1" to "Quantized object #4" store a list of information indicating the codebook corresponding to each partial vector after the objects are product quantized.
[0412] In the example of FIG. 35, a blob with blob ID "1" belongs to a cluster with cluster ID "2" and the number of objects belonging to the cluster is four. Furthermore, for a blob with blob ID "1," "4, 2, 8, 10" are stored in "Quantization Object #1," indicating that the object is composed of four partial vectors, which are quantized with the codebooks "4," "2," "8," and "10," respectively. For example, for an object corresponding to "Quantization Object #1" with blob ID "1," the first partial vector of the four divided partial vectors is quantized with the codebook "4," the second partial vector is quantized with the codebook "2," the third partial vector is quantized with the codebook "8," and the last partial vector is quantized with the codebook "10."
[0413] The transposed index IND1 may store various information depending on the purpose, not limited to the above. For example, the transposed index IND1 may store information indicating each quantized object (such as an object ID) in association with each quantized object.
[0414] Next, the transposed index IND2 shown in FIG. 36 will be described. FIG. 36 shows a specific example of a transposed index corresponding to the example shown in FIG. 32, etc. Note that the description of the same points as in the transposed index IND1 will be omitted as appropriate. The transposed index IND2 shown in FIG. 36 includes items such as "blob ID," "cluster ID," "number of quantized objects," and "quantized object #1." Note that although only "quantized object #1" is shown in FIG. 36, like the transposed index IND1, it includes items in the number corresponding to the number of objects belonging to each blob.
[0415] In the transposed index IND2, similarly to the spatial information SP51 shown in FIG. 32, it is indicated that the blob BL1 identified by the blob ID "BL1" belongs to the cluster CL1 with the cluster ID "CL1". Also, it is indicated that the number of objects belonging to the blob BL1 is "NM1". Note that although an abstract symbol such as "NM1" is used in FIG. 36, a value indicating a concrete number (for example, 10 or 100) is stored in the "quantized object number" as in the transposed index IND1.
[0416] Furthermore, blob BL1 indicates that "Quantized Object #1" stores "CD51, CD64, CD77, CD82", and that the object is composed of four partial vectors, which are quantized using the codebooks "CD51", "CD64", "CD77", and "CD82", respectively.
[0417] 32, the inverted index IND2 indicates that the blob BL2 identified by the blob ID "BL2" belongs to the cluster CL2 with the cluster ID "CL2". The inverted index IND2 indicates that the blob BL7 identified by the blob ID "BL7" belongs to the cluster CL1 with the cluster ID "CL1".
[0418] The transposed index IND2 is not limited to the above, and may store various information depending on the purpose. For example, the transposed index IND2 may store information indicating each quantized object (such as an object ID) in association with each quantized object. In this way, the transposed index is assigned information identifying a cluster (a cluster ID), and the information processing device 100C calculates a residual vector using the transposed index. For example, the residual vector is used when generating and searching an index, but since this is similar to the processing in conventional techniques such as Patent Document 3, detailed description thereof will be omitted as appropriate.
[0419] [4-4. Information processing flow] Next, the procedure of information processing according to the fourth embodiment will be described with reference to Fig. 37 and Fig. 38. Fig. 37 and Fig. 38 are flowcharts showing an example of information processing according to the fourth embodiment.
[0420] First, Fig. 37 will be described. For example, Fig. 37 is a flowchart showing an example of group generation processing. As shown in Fig. 37, information processing device 100C acquires object information indicating multiple objects to be subjected to data search (step S601). For example, information processing device 100C acquires object information indicating multiple objects from object information storage unit 121.
[0421] The information processing device 100C generates, based on the object information, first group information indicating a plurality of first groups into which the plurality of objects are classified (step S602). For example, the information processing device 100C generates, based on the object information, blob information indicating a plurality of blobs into which the plurality of objects are classified.
[0422] The information processing device 100C generates second group information indicating a plurality of second groups, which are groups classified differently from the plurality of first groups and which classify, among the plurality of objects, objects belonging to the same first group in the plurality of first groups into the same group (step S603). For example, the information processing device 100C generates cluster information indicating a plurality of clusters, at least one of which includes two or more blobs. Note that the information processing device 100C may also generate cluster information indicating clusters in one-to-one correspondence with blobs.
[0423] Next, FIG. 38 will be described. For example, FIG. 38 is a flowchart showing an example of processing using groups. As shown in FIG. 38, the information processing device 100C acquires processing information indicating a plurality of first groups into which a plurality of objects to be subjected to data search are classified, and a plurality of second groups which are groups classified differently from the plurality of first groups and which classify, among the plurality of objects, objects belonging to the same first group in the plurality of first groups into the same group (step S701). For example, the information processing device 100C acquires, as processing information, blob information stored in the blob information storage unit 125 and cluster information stored in the cluster information storage unit 127. For example, the information processing device 100C acquires an inverted index IND1 or an inverted index IND2.
[0424] The information processing device 100C uses the processing information to execute a first process for objects belonging to the first group and a second process for objects belonging to the second group (step S702). For example, the information processing device 100C executes batch processing for objects belonging to blobs and executes quantization for objects belonging to clusters.
[0425] 5. Fifth Embodiment Note that when performing search processing using a blob connection graph (also referred to as a "blob graph") that connects blobs as described above, search ranges from multiple perspectives may be used. This will be described below as a fifth embodiment. In the fifth embodiment, the information processing system 1 has an information processing device 100D instead of the information processing device 100, the information processing device 100A, the information processing device 100B, or the information processing device 100C. Note that descriptions of the same points as those described in the first, second, third, and fourth embodiments will be omitted as appropriate. For example, the blob BL* (* is an arbitrary number) in the following description is a group of objects, similar to the above-described blob BL* (* is an arbitrary number), and therefore detailed descriptions thereof will be omitted. Furthermore, the centroid C* (* is an arbitrary number) in the following description is similar to the above-described centroid C* (* is an arbitrary number), and therefore detailed descriptions thereof will be omitted. In the following description, the cluster CL* (* is an arbitrary number) is a cluster for classifying blobs in the same way as the cluster CL* (* is an arbitrary number) described above, and therefore a detailed description thereof will be omitted.
[0426] 39 and 40, the representation is the same as in FIG. 32 except that the distinction between clusters is indicated by thick solid lines and the connections between blobs are indicated by dotted lines connecting centroids C of blobs BL. For example, among the solid lines, thick lines indicate the division correspondence between clusters and blobs, and thin lines indicate the division correspondence between blobs. For example, among the solid lines, thin lines indicate the division correspondence between blobs within a cluster. Also, in FIGS. 39 and 40, centroids CN11 to CN14 indicate the centroids of each of clusters CL11 to CL14. Cluster CL11 includes three blobs BL: blobs BL11, BL12, and BL13. Blob BL11 includes nodes N31, N205, etc. Cluster CL12 includes two blobs BL: blobs BL14 and BL15. Cluster CL13 includes three blobs BL: blobs BL16, BL17, and BL18. Cluster CL14 includes two blobs BL: blobs BL19 and BL20. Note that the clusters CL, blobs BL, and nodes N shown in Figures 39 and 40 are only a partial list, and the clusters CL, blobs BL, and nodes N to be processed include many more clusters CL, blobs BL, and nodes N in addition to those shown. Also, the symbols for the blobs BL and clusters CL are shown only in Figure 39 and are omitted in Figure 40, but Figures 39 and 40 show the same spatial information (spatial information SP61).
[0427] The information processing device 100D executes a search process using a blob graph, which is a group connection graph in which multiple groups into which multiple objects to be searched for data are classified are connected by edges. The information processing device 100D also executes the search process using a range corresponding to a search for blobs, which are groups (hereinafter also referred to as a "first search range"), and a range corresponding to a search for multiple objects (hereinafter also referred to as a "second search range"). The information processing device 100D thereby extracts nearby objects corresponding to the search query from the multiple objects.
[0428] [5-1. Information Processing] First, an overview of information processing according to the fifth embodiment will be described using Figs. 39 and 40. Figs. 39 and 40 are diagrams showing an example of information processing according to the fifth embodiment. Specifically, Fig. 39 shows an overall overview of search processing executed by the information processing device 100D. Furthermore, Fig. 40 shows changes in the first search range and the second search range in the search processing executed by the information processing device 100D. Thus, Figs. 39 and 40 explain the changes in the first search range and the second search range and an overview of the search processing, and the specific flow of the search processing is shown in Fig. 43.
[0429] The graph GR61 shown in the spatial information SP61 in FIG. 39 is a blob graph connecting blobs, similar to the graph GR41 shown in the spatial information SP51 in FIG. 32. For example, the dotted lines (edges) connecting the centroids C of the blobs BL in the graph GR61 are synonymous with the arrowed lines (edges) connecting the blobs BL to each other in the graph GR41. The graph GR61 in FIG. 39 illustrates a portion of blobs BL11 to BL20. For example, the information processing device 100D may generate information indicating blobs BL11 to BL10, etc., using the method described above, or may acquire information indicating blobs BL11 to BL10, etc., from another device. Note that the spatial information SP61-1 to SP61-4 in FIG. 40 are assigned different reference numerals for the convenience of explaining changes in the first search range and the second search range, but are similar to the spatial information SP61 in FIG. 39, and may be referred to as spatial information SP61 when describing without distinction.
[0430] Note that the graph GR61 may be in a state in which the graph GR41 illustrates portions of the blobs BL11 to BL20. In this case, the spatial information SP51 and the spatial information SP61 target the same multiple objects, and the graph GR61 and the graph GR41 illustrate different portions of the same graph. For example, the graph GR41 in FIG. 32 may illustrate portions of the blobs BL1 to BL10, and the graph GR61 may illustrate portions of the blobs BL11 to BL20. The graph GR61 may also be a graph in which the centroids of the blobs are connected, as in the graph GR42 in FIG. 29. In other words, the graph GR61 may be a graph in which the blobs are connected in any manner, as long as the search process described below can be executed.
[0431] FIG. 39 shows an example of processing when the query is search query QE11 and the number of objects to be extracted as nearby objects (also referred to as the "number of nearby objects to be extracted") is "5." Of the five circles centered on search query QE11, the dashed-dotted circle indicates blob search range SB, which is an example of the first search range. Note that in FIG. 39, of the three dashed-dotted circles indicating blob search range SB, only the circle with the largest radius is labeled "SB," but the three dashed-dotted circles indicate blob search range SB that changes during the search processing.
[0432] For example, blob search range SB is a first search range defined by a radius (blob search radius) centered on search query QE11. As indicated by the dashed-dotted arrow in FIG. 39, the range (radius) of blob search range SB widens as the search process progresses; this point will be described later with reference to FIG. 40. Note that blob search ranges SB1, SB2, and SB3 shown in FIG. 40 correspond to the three blob search ranges SB in FIG. 39. When the blob search ranges SB1, SB2, and SB3 are not to be distinguished from one another, they may be referred to as blob search range SB.
[0433] 39, of the five circles centered on the search query QE11, the two-dot chain circle indicates the object search range SO, which is an example of the first search range. Note that in Fig. 39, of the two two-dot chain circles indicating the object search range SO, only the outer circle is labeled "SO," but the two two-dot chain circles indicate the object search range SO that changes during the search process.
[0434] For example, object search range SO is a second search range defined by a radius (object search radius) centered on search query QE11. As indicated by the two-dot chain arrow in FIG. 39, the range (radius) of object search range SO becomes smaller as the search process progresses, but this point will be described later with reference to FIG. 40. Note that object search ranges SO1 and SO2 shown in FIG. 40 correspond to the two object search ranges SO in FIG. 39. When the object search ranges SO1 and SO2 are not to be distinguished from each other, they may be referred to as object search range SO.
[0435] First, the information processing device 100D starts search processing from the state shown in space information SP61-1. Then, as shown in space information SP61-2, the information processing device 100D determines the blob search range SB to be the blob search range SB1, and determines the object search range SO to be infinity "∞". That is, in FIG. 40, the information processing device 100D determines the blob search range SB to be the blob search range SB1, which is defined by a radius based on the distance between the nearest blob BL20 and the search query QE11. Also in FIG. 40, the information processing device 100D extracts three of the nodes N2, N6, and N8 to which the blob BL20 belongs (objects corresponding to the nodes) as candidates for nearby objects corresponding to the search query QE11, and adds them to the nearby object candidates. Note that there are three nearby object candidates, which is less than or equal to "5" in the nearby object extraction count, so the object search range SO is maintained at infinity "∞".
[0436] Then, as shown in spatial information SP61-3, information processing device 100D determines blob search range SB to be blob search range SB2, and determines object search range SO to be object search range SO1. That is, in FIG. 40, information processing device 100D determines blob search range SB to be blob search range SB2, which is defined by a radius based on the distance between blob BL11, the next nearest blob after blob BL20, and search query QE11. Also in FIG. 40, information processing device 100D extracts three nodes (objects corresponding to) nodes N31, N205, and N312 to which blob BL11 belongs as candidates for nearby objects corresponding to search query QE11, and adds them to the nearby object candidates. As a result, the nearby object candidates include six nodes N2, N6, N8, N31, N205, and N312, exceeding the number of nearby object extractions "5." Therefore, the information processing device 100D keeps five of the nearby object candidates closest to the search query QE11 and excludes the rest. In this case, the information processing device 100D excludes node N312, which is the farthest from the search query QE11, out of the six nodes, and sets the number of nearby object candidates to five. That is, the information processing device 100D sets the number of nearby object candidates to five: nodes N2, N6, N8, N31, and N205. Because the number of nearby object candidates has reached the nearby object extraction number "5," the information processing device 100D determines the object search range SO to be object search range SO1, which is defined by a radius based on the distance between the search query QE11 and node N205, which is the nearby object candidate farthest from the search query QE11.
[0437] Then, as shown in spatial information SP61-4, information processing device 100D determines blob search range SB to be blob search range SB3, and determines object search range SO to be object search range SO2. That is, in FIG. 40, information processing device 100D determines blob search range SB to be blob search range SB3, which is defined by a radius based on the distance between blob BL19, the next nearest blob after blob BL11, and search query QE11. Also in FIG. 40, information processing device 100D extracts three nodes (objects corresponding to) nodes N3, N11, and N114 to which blob BL19 belongs as candidates for nearby objects corresponding to search query QE11, and adds them to the nearby object candidates. As a result, the nearby object candidates include eight nodes N2, N6, N8, N31, N205, N3, N11, and N114, exceeding the number of nearby object extractions "5." Therefore, the information processing device 100D keeps five of the nearby object candidates closest to the search query QE11 and excludes the rest. In this case, the information processing device 100D excludes three nodes N205, N3, and N114 from the farthest distance from the search query QE11, leaving the number of nearby object candidates at five. That is, the information processing device 100D leaves five nearby object candidates: nodes N2, N6, N8, N11, and N31. Then, the information processing device 100D determines the object search range SO to be an object search range SO2, which is defined by a radius based on the distance between the search query QE11 and node N11, which is the nearby object candidate farthest from the search query QE11.
[0438] The information processing device 100D repeats the above-described processing until termination criteria #1 to #3, which will be described later, are met. For example, when termination criteria #1 to #3 are met, the information processing device 100D determines that the objects included in the nearby object candidates at that time are nearby objects of the search query QE11. Then, the information processing device 100D provides information indicating the objects determined as nearby objects of the search query QE11 to the source of the search that specified the search query QE11, etc.
[0439] As described above, the information processing device 100D performs search processing using search ranges from multiple perspectives, such as a first search range and a second search range, and by changing each search range depending on the status of the search processing, it is possible to perform flexible search processing using graphs.
[0440] For example, in the case of a blob graph, a large blob graph may have a slower search speed due to unnecessary blob references. Also, in the past, product quantization (PQ) was used to search a graph to identify nearby blobs, and then calculate approximate distances for the identified blobs, but this method also slows down the search speed due to unnecessary blob references.
[0441] On the other hand, the information processing device 100D searches the blob graph and calculates the approximate distance of objects within each blob for which the ranking in order of proximity has been determined, and terminates the search when a specified termination criterion is met, thereby reducing unnecessary object searches.
[0442] 5-2. Configuration of information processing device Next, the configuration of an information processing device 100D according to the fifth embodiment will be described with reference to Fig. 41. Fig. 41 is a diagram showing an example of the configuration of an information processing device according to the fifth embodiment. As shown in Fig. 41, the information processing device 100D has a communication unit 110, a storage unit 120C, and a control unit 130B. Note that in the information processing device 100D, descriptions of the same aspects as those of the information processing device 100, the information processing device 100A, the information processing device 100B, or the information processing device 100C will be omitted as appropriate.
[0443] (Storage unit 120C) The storage unit 120C is realized by, for example, a semiconductor memory element such as a RAM or a flash memory, or a storage device such as a hard disk or an optical disk. As shown in Fig. 41 , the storage unit 120C according to the fifth embodiment includes an object information storage unit 121, a graph information storage unit 122, a quantization information storage unit 123, a codebook information storage unit 124, a blob information storage unit 125, a blob connection graph information storage unit 126, and a cluster information storage unit 127.
[0444] (control unit 130D) The control unit 130D is a controller, and is realized by, for example, a CPU, an MPU, a GPU, etc., executing various programs (corresponding to examples of information processing programs) stored in a storage device inside the information processing device 100D using RAM as a work area. The control unit 130D is also a controller, and is realized by, for example, an integrated circuit such as an ASIC or FPGA.
[0445] 41, control unit 130D has an acquisition unit 131, a generation unit 132C, a search processing unit 133D, and a provision unit 134, and realizes or executes the functions and actions of information processing described below. Note that the internal configuration of control unit 130D is not limited to the configuration shown in FIG. 41, and may be any other configuration as long as it performs the information processing described below.
[0446] The acquisition unit 131 according to the fifth embodiment acquires various pieces of information from the storage unit 120C. For example, like the acquisition unit 131 according to the fourth embodiment, the acquisition unit 131 according to the fifth embodiment acquires various pieces of information used by the generation unit 132C.
[0447] The acquisition unit 131 acquires a group connection graph in which a plurality of groups into which a plurality of objects to be searched for data are classified are connected by edges. The acquisition unit 131 acquires a search query for a plurality of objects.
[0448] (Search processing unit 133D) The search processing unit 133D executes the search processing shown in Fig. 43. For example, the search processing unit 133D performs various processes related to the search processing in the same way as the search processing unit 133, the search processing unit 133A, the search processing unit 133B, or the search processing unit 133BC. The search processing unit 133D performs processing using information stored in the storage unit 120C. The search processing unit 133D performs processing using information generated by the generation unit 132C.
[0449] The search processing unit 133D executes a search process to extract nearby objects corresponding to a search query from the plurality of objects by searching the group connection graph using a first search range indicating a range corresponding to a search of multiple groups and a second search range indicating a range corresponding to a search of multiple objects. The search processing unit 133D executes the search process using the first search range and the second search range, which is wider than the first search range.
[0450] The search processing unit 133D executes the search process using a first search range that expands as the search process progresses. The search processing unit 133D executes the search process using a first search range determined based on one group that is the target of the search process out of multiple groups. The search processing unit 133D executes the search process using a first search range determined based on one group currently being searched. The search processing unit 133D executes the search process using a first search range that is based on the distance between the search query and the object closest to the search query among objects included in the one group, or the centroid of the one group.
[0451] The search processing unit 133D executes the search process using a second search range that narrows as the search process progresses. The search processing unit 133D executes the search process using the second search range determined based on one object that is the target of the search process among multiple objects. The search processing unit 133D executes the search process using the second search range determined based on one object that is the farthest from the search query among objects extracted as candidates for nearby objects corresponding to the search query during the search process. The search processing unit 133D executes the search process using the second search range based on the distance between the search query and the one object.
[0452] The search processing unit 133D ends the search process when a predetermined termination condition is met. The search processing unit 133D ends the search process when the number of groups that have been the target of the search process meets the condition. The search processing unit 133D ends the search process when the number of groups that have been the target of the search process exceeds a predetermined threshold.
[0453] The search processing unit 133D ends the search process when the distance related to the predetermined group exceeds the second search range. The search processing unit 133D ends the search process when the distance related to the predetermined group during the search process exceeds the second search range. The search processing unit 133D ends the search process when the distance between the search query and an object closest to the search query among the objects included in the predetermined group exceeds the second search range.
[0454] The search processing unit 133D ends the search process when the magnitude relationship between the first search range and the second search range satisfies the condition. The search processing unit 133D ends the search process when the first search range exceeds the second search range.
[0455] [5-3. Information processing flow] Next, the procedure of information processing according to the fifth embodiment will be described with reference to Fig. 42. Fig. 42 is a flowchart showing an example of information processing according to the fifth embodiment.
[0456] First, Fig. 42 will be described. For example, Fig. 42 is a flowchart showing an example of group generation processing. As shown in Fig. 42, the information processing device 100D acquires a group connection graph in which a plurality of groups into which a plurality of objects to be the target of data search are classified are connected by edges (step S801). For example, the information processing device 100D acquires a blob connection graph from the blob connection graph information storage unit 126 as the group connection graph.
[0457] The information processing device 100D acquires a search query for a plurality of objects (step S802). For example, the information processing device 100D acquires the search query from the terminal device 10 (see FIG. 4) used by the user.
[0458] The information processing device 100D executes a search process to extract nearby objects corresponding to the search query from the plurality of objects by searching the group connection graph using a first search range indicating a range corresponding to a search of a plurality of groups and a second search range indicating a range corresponding to a search of a plurality of objects (step S803). For example, the information processing device 100D executes the search process shown in Fig. 43 using a blob search range that is the first search range and an object search range that is the second search range.
[0459] [5-4. Search processing example] Here, an example of the search processing according to the fifth embodiment will be described with reference to FIG. 43. FIG. 43 is a flowchart showing an example of the search processing according to the fifth embodiment. The search processing described below is performed by the search processing unit 133D of the information processing device 100D. Note that the description of the same points as those in the search processing described in the first, second, third, and fourth embodiments will be omitted as appropriate.
[0460] The information processing device 100D acquires a search blob starting set and calculates the distance between each object and the query (step S901). For example, the information processing device 100D acquires a search blob starting set indicating a blob that is the starting point of the search, and calculates the distance between each object included in the blob of the search blob starting set and the query. For example, the information processing device 100D may randomly select the search blob starting set, or may select the search blob starting set using a predetermined index, such as an index of a tree structure with blobs as leaves. For example, the information processing device 100D calculates the approximate distance between each object included in the blob of the search blob starting set and the query by calculating the approximate distance described above.
[0461] The information processing device 100D determines the nearest blob to the query in the search blob starting set as the current nearest blob B (hereinafter simply referred to as "blob B") (step S902). For example, the information processing device 100D determines the blob that is closest to the query in the search blob starting set as the current nearest blob B.
[0462] The information processing device 100D sets various parameters (step S903). In FIG. 43, the information processing device 100D sets the number of blob searches to "1". Furthermore, the information processing device 100D sets the object search radius R indicating the second search range (object search range) to "∞ (infinity)". For example, the object search radius R is a radius centered on the query q that defines the second search range. Furthermore, the information processing device 100D sets the object search radius E, which is an example of the second search range, to "∞ (infinity)".
[0463] Furthermore, the information processing device 100D sets the blob search radius Rb, which is an example of a first search range (blob search range), to "d(q,B)". For example, the blob search radius Rb is a radius centered on the query q that defines the first search range. For example, the information processing device 100D sets the blob search radius Rb to the distance d(q,B) between the query q and the current nearest blob B. For example, the information processing device 100D sets the blob search radius Rb to the distance d(q,B) between the query q and the object closest to the query q among the objects included in the current nearest blob B. Note that the blob search radius Rb may be the distance between the query q and the centroid of the current nearest blob B. Furthermore, the information processing device 100D sets the blob search radius Eb, which indicates the first search range (blob search range), to "Rb×(1+eb)". Here, the parameter eb is a blob search radius expansion rate (also referred to as a "first search range coefficient"), and is set to an arbitrary value such as 0.1. For example, the information processing device 100D sets the blob search radius Eb to a value obtained by multiplying the blob search radius Rb by 1+eb.
[0464] The information processing device 100D adds the search blob starting set other than blob B to the excluded blob set (step S904). For example, the information processing device 100D adds blobs other than the current nearest blob B from the search blob starting set to the excluded blob set.
[0465] The information processing device 100D adds the search blob starting set to the unsearched blob set and the distance-calculated set (step S905). For example, the information processing device 100D adds all blobs in the search blob starting set, including the current nearest blob B, to the unsearched blob set, which indicates that they have not yet been subjected to search processing. For example, the information processing device 100D adds all blobs in the search blob starting set, including the current nearest blob B, to the distance-calculated set, which indicates that distance calculation processing has been performed for objects within the blobs.
[0466] The information processing device 100D determines whether the unsearched blob set is empty (step S906). If the unsearched blob set is not empty (step S906: No), the information processing device 100D acquires a blob T that is closest to the query from the unsearched blob set (step S907). As a result, the unsearched blob set is excluded from the unsearched blob set. For example, the information processing device 100D acquires, as blob T, a blob that has the shortest distance from the query q in the unsearched blob set. For example, the distance from the query q is the distance between the query q and the object that is closest to the query q among the objects included in the blob. Note that the distance from the query q may be the distance between the query q and the centroid of the blob.
[0467] The information processing device 100D determines whether the blob T is within the blob search range (step S908). For example, the information processing device 100D determines whether the distance between the blob T and the query q is equal to or less than the blob search radius Eb.
[0468] If blob T is within the blob search range (step S908: Yes), the information processing device 100D determines whether all blobs connected to blob T by edges have been processed (step S909). For example, the information processing device 100D determines whether all connected blobs of blob T have been processed using information ("processing presence / absence information") indicating whether blobs connected to blob T by edges (also referred to as "connected blobs") have been processed. For example, the information processing device 100D may manage whether connected blobs of blob T have been processed using information such as a flag. For example, the information processing device 100D may update the processing presence / absence information stored in the storage unit 120C according to the processing status, and may determine whether all connected blobs of blob T have been processed by referring to the processing presence / absence information in the storage unit 120C.
[0469] If the information processing device 100D has not processed all of the blobs connected to blob T by edges (step S909: Yes), it acquires blob J connected to blob T by an edge (step S910). For example, the information processing device 100D acquires, among the blobs connected to blob T by an edge, a blob that has not been processed as a blob connected to blob T by an edge (also referred to as an "unprocessed connected blob") as blob J. For example, the information processing device 100D refers to the processing presence / absence information, acquires, among the blobs connected to blob T by an edge, an unprocessed connected blob as blob J, and updates the processing presence / absence information for the blob acquired as blob J to "processed."
[0470] The information processing device 100D determines whether or not the blob J exists in the distance-calculated set (step S911). If the blob J exists in the distance-calculated set (step S911: Yes), the information processing device 100D returns to step S909 and performs the process.
[0471] If blob J does not exist in the distance-calculated set (step S911: No), the information processing device 100D adds blob J to the distance-calculated set (step S912) and adds blob J to the unsearched blob set (step S913). Then, the information processing device 100D calculates the distance De between the centroid of blob J and the query (step S914).
[0472] The information processing device 100D determines whether the distance De is smaller than the blob search radius Rb (step S915). If the distance De is not smaller than the blob search radius Rb (step S915: No), the information processing device 100D adds the blob J to the excluded blob set (step S916) and returns to step S909 to perform the process.
[0473] Furthermore, if the distance De is smaller than the blob search radius Rb (step S915: Yes), the information processing device 100D adds the current nearest neighbor blob B to the set of excluded blobs and sets blob J as the current nearest neighbor blob B (step S917). In this way, if the distance De is smaller than the blob search radius Rb, the information processing device 100D updates the current nearest neighbor blob B to blob J.
[0474] Then, the information processing device 100D updates the parameters related to the blob based on the current nearest neighbor blob B that has been updated to blob J (step S918). In FIG. 43, the information processing device 100D updates the blob search radius Rb to "d(q, B)" based on the current nearest neighbor blob B that has been updated to blob J. Furthermore, the information processing device 100D updates the blob search radius Eb to "Rb×(1+eb)" based on the updated blob search radius Rb. Then, the information processing device 100D returns to step S909 and performs the processing.
[0475] When the information processing device 100D has processed all blobs connected by edges of blob T (step S909: Yes), it returns to step S906 and performs the processing. When the unsearched blob set is empty (step S906: Yes), the information processing device 100D performs the processing of step S920. Furthermore, when blob T is not within the blob search range (step S908: No), the information processing device 100D returns blob T to an unsearched blob (step S919) and performs the processing of step S920.
[0476] In step S920, the information processing device 100D adds 1 to the blob search count (step S920). Then, the information processing device 100D calculates the approximate distances of all objects in blob B, and sets the shortest distance as the blob distance of the blob B (step S921). For example, the information processing device 100D calculates the approximate distance between query q and all objects in blob B, and sets the distance between query q and the object closest to query q among the objects included in the current nearest blob B as the blob distance of blob B. Note that the blob distance of blob B is used in the termination determination process in step S924.
[0477] The information processing device 100D adds and sorts all objects in blob B to a search result set, and sets only those with the highest specified search counts as the search result set (step S922). For example, the information processing device 100D adds all objects in blob B to a search result set that indicates candidate objects near query q. The information processing device 100D then updates the search result set by sorting the objects in the search result set after adding the objects in blob B in order of shortest distance from query q, and excludes from the updated search result set any objects that are lower than the specified search count (e.g., 5, 10, etc.). For example, the information processing device 100D keeps, from the objects in the search result set after adding the objects in blob B, the objects with the shortest distance from query q up to the search count (e.g., 5, 10, etc.) (i.e., the number of objects equal to the search count) and excludes the rest.
[0478] Then, the information processing device 100D updates the parameters related to the objects based on the updated search result set (step S923). In FIG. 43, the information processing device 100D updates the object search radius R to the maximum distance of the objects in the search result set based on the updated search result set. For example, the information processing device 100D updates the object search radius R to the distance of the object that is farthest from the query q in the search result set. Furthermore, the information processing device 100D sets the object search radius E to "R×(1+eo)" based on the updated search result set. Here, the parameter eo is an object search radius expansion rate (also referred to as a "second search range coefficient"), and is set to an arbitrary value such as 0.05. For example, the parameter eo corresponds to the above-mentioned search range coefficient "ε". For example, the information processing device 100D sets the object search radius E to a value obtained by multiplying the object search radius R by 1+eo. The parameters eo and eb may have the same value.
[0479] Then, the information processing device 100D determines whether termination criteria #1 to #3 are met (step S924). The information processing device 100D determines whether the processing status meets any of termination criteria #1 to #3 based on each termination condition. Termination criterion #1 terminates the process when the upper limit number of blob searches is exceeded. For example, termination criterion #1 is a criterion that sets a condition that the number of blob searches, which is a parameter, exceeds the upper limit number of searches (for example, any value such as 5, 10, etc.).
[0480] Termination criterion #2 terminates the process when the blob distance exceeds the object search radius. For example, termination criterion #2 is a criterion that requires the blob distance to exceed the object search radius E. Termination criterion #3 terminates the process when the blob search radius exceeds the object search radius. For example, termination criterion #3 is a criterion that requires the blob search radius Eb to exceed the object search radius E. Note that the above termination criteria #1 to #3 are merely examples, and various termination criteria may be used. Furthermore, multiple termination criteria may be used in combination.
[0481] If the termination criteria #1 to #3 are not met (step S924: No), the information processing device 100D acquires the blob with the shortest distance from the set of excluded blobs and sets it as the current nearest neighbor blob B (step S925). For example, the information processing device 100D acquires the blob from the set of excluded blobs that has the shortest distance to the query q as the current nearest neighbor blob B. As a result, the blob with the shortest distance at that time is removed from the set of excluded blobs. For example, the information processing device 100D updates the current nearest neighbor blob B to the blob from the set of excluded blobs that has the shortest distance to the query q.
[0482] Then, the information processing device 100D updates the parameters related to the blob based on the current nearest neighbor blob B that has been updated to the blob with the shortest distance (step S926). In FIG. 43, the information processing device 100D updates the blob search radius Rb to "d(q, B)" based on the current nearest neighbor blob B that has been updated to the blob with the shortest distance. Furthermore, the information processing device 100D updates the blob search radius Eb to "Rb×(1+eb)" based on the updated blob search radius Rb. Then, the information processing device 100D returns to step S906 and performs processing.
[0483] Furthermore, if the termination criteria #1 to #3 are met (step S924: Yes), the information processing device 100D terminates the process. For example, if the termination criteria #1 to #3 are met, the information processing device 100D determines that the objects included in the search result set at that time are neighboring objects of query q. Then, the information processing device 100D provides information indicating the objects determined to be neighboring objects of query q to the source of the search that specified query q, etc.
[0484] [6. Effects] As described above, the information processing device 100D according to the fifth embodiment includes an acquisition unit 131 and a search processing unit 133D. The acquisition unit 131 acquires a group connection graph in which a plurality of groups into which a plurality of objects to be subjected to data search are classified are connected by edges, and a search query for the plurality of objects. The search processing unit 133D executes a search process to extract nearby objects corresponding to the search query from the plurality of objects by searching the group connection graph using a first search range indicating a range corresponding to a search for a plurality of groups and a second search range indicating a range corresponding to a search for a plurality of objects.
[0485] In this way, the information processing device 100D according to the fifth embodiment performs search processing using two search ranges: a first search range indicating a range corresponding to a search of multiple groups, and a second search range indicating a range corresponding to a search of multiple objects, thereby enabling flexible search processing using graphs.
[0486] The search processing unit 133D executes the search process using a first search range and a second search range that is wider than the first search range.
[0487] In this way, the information processing device 100D can perform flexible search processing using graphs by performing search processing using a first search range that indicates a range corresponding to a search of multiple groups and a second search range that indicates a range corresponding to a search of multiple objects.
[0488] The search processing unit 133D executes the search process using a first search range that widens as the search process progresses.
[0489] In this way, the information processing device 100D executes search processing using the first search range, which expands as the search processing progresses, and can therefore perform flexible search processing using graphs.
[0490] The search processing unit 133D executes the search process using a first search range that is determined based on one group that is the target of the search process among the multiple groups.
[0491] In this way, the information processing device 100D can perform flexible search processing using graphs by performing search processing using a first search range determined based on one group among multiple groups that is the target of the search processing.
[0492] The search processing unit 133D executes the search process using a first search range determined based on one group during the search process.
[0493] In this way, the information processing device 100D executes search processing using the first search range determined based on one group during search processing, thereby enabling flexible search processing using graphs.
[0494] The search processing unit 133D executes the search process using a first search range based on the distance between the search query and the object closest to the search query among the objects included in one group, or the centroid of one group.
[0495] In this way, the information processing device 100D executes search processing using the first search range based on the distance to the search query, thereby enabling flexible search processing using a graph.
[0496] The search processing unit 133D executes the search process using a second search range that narrows as the search process progresses.
[0497] In this way, the information processing device 100D executes search processing using the second search range, which narrows as the search processing progresses, and can therefore perform flexible search processing using a graph.
[0498] Search processing unit 133D executes the search process using a second search range that is determined based on one object that is the target of the search process among the plurality of objects.
[0499] In this way, the information processing device 100D can perform flexible search processing using graphs by executing search processing using a second search range determined based on one object that is the target of the search processing among multiple objects.
[0500] During the search process, the search processing unit 133D performs the search process using a second search range determined based on one object that is the farthest from the search query among the objects extracted as candidates for nearby objects corresponding to the search query.
[0501] In this way, during the search process, the information processing device 100D performs the search process using a second search range that is determined based on the object farthest from the search query among the candidate nearby objects that correspond to the search query, thereby enabling flexible search processing using graphs.
[0502] The search processing unit 133D executes the search process using a second search range based on the distance between the search query and one object.
[0503] In this way, the information processing device 100D executes search processing using the second search range based on the distance to the search query, thereby enabling flexible search processing using a graph.
[0504] When a predetermined termination condition is met, the search processing unit 133D terminates the search process.
[0505] In this way, the information processing device 100D can terminate the search process at an appropriate timing by terminating the search process when a predetermined termination condition is satisfied, and can perform flexible search processing using a graph.
[0506] When the number of groups that have been the subject of the search process satisfies the condition, the search processing unit 133D ends the search process.
[0507] In this way, the information processing device 100D can terminate the search process at an appropriate time by terminating the search process when the number of groups targeted for the search process meets the conditions, thereby enabling flexible search processing using graphs.
[0508] When the number of groups that have been the subject of the search process exceeds a predetermined threshold, the search processing unit 133D ends the search process.
[0509] In this way, the information processing device 100D can terminate the search process at an appropriate time by terminating the search process when the number of groups targeted for the search process exceeds a predetermined threshold, thereby enabling flexible search processing using graphs.
[0510] If the distance for the predetermined group exceeds the second search range, the search processing unit 133D ends the search process.
[0511] In this way, the information processing device 100D can terminate the search process at an appropriate time by terminating the search process when the distance for a specified group exceeds the second search range, thereby enabling flexible search processing using graphs.
[0512] When the distance for a predetermined group during the search process exceeds the second search range, the search processing unit 133D ends the search process.
[0513] In this way, the information processing device 100D can terminate the search process at an appropriate time by terminating the search process when the distance for a specified group during the search process exceeds the second search range, thereby enabling flexible search processing using graphs.
[0514] The search processing unit 133D ends the search process when the distance between the search query and the object that is closest to the search query among the objects included in the predetermined group exceeds the second search range.
[0515] In this way, the information processing device 100D can terminate the search process at an appropriate time by terminating the search process when the distance between the search query and the object closest to the search query among the objects included in a specified group exceeds the second search range, thereby enabling flexible search processing using graphs.
[0516] If the magnitude relationship between the first search range and the second search range satisfies the condition, the search processing unit 133D ends the search process.
[0517] In this way, the information processing device 100D can terminate the search process at an appropriate time by terminating the search process when the magnitude relationship between the first search range and the second search range satisfies the condition, thereby enabling flexible search processing using graphs.
[0518] If the first search range exceeds the second search range, the search processing unit 133D ends the search process.
[0519] In this way, the information processing device 100D can terminate the search process at an appropriate timing by terminating the search process when the first search range exceeds the second search range, and can perform flexible search processing using graphs.
[0520] [7. Hardware Configuration] The information processing devices 100, 100A, 100B, 100C, and 100D according to the above-described embodiments are realized by a computer 1000 having a configuration as shown in Fig. 44, for example. Fig. 44 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing device. The computer 1000 has a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.
[0521] The CPU 1100 operates and controls each unit based on programs stored in the ROM 1300 or the HDD 1400. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs that depend on the hardware of the computer 1000, and the like.
[0522] The HDD 1400 stores programs executed by the CPU 1100, data used by such programs, etc. The communication interface 1500 receives data from other devices via the network N and sends it to the CPU 1100, and transmits data generated by the CPU 1100 to other devices via the network N.
[0523] The CPU 1100 controls output devices such as a display and a printer, and input devices such as a keyboard and a mouse, via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. The CPU 1100 also outputs generated data to the output devices via the input / output interface 1600.
[0524] Media interface 1700 reads a program or data stored in recording medium 1800 and provides it to CPU 1100 via RAM 1200. CPU 1100 loads the program or data from recording medium 1800 onto RAM 1200 via media interface 1700 and executes the loaded program. Recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc) or a PD (Phase Change Rewritable Disc), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0525] For example, when the computer 1000 functions as the information processing device 100 according to the first embodiment, the CPU 1100 of the computer 1000 executes programs loaded onto the RAM 1200 to implement the functions of the control unit 130. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800, but as another example, the CPU 1100 may obtain these programs from another device via the network N.
[0526] Although some of the embodiments of the present application have been described in detail above with reference to the drawings, these are merely examples, and the present invention can be implemented in other forms that incorporate various modifications and improvements based on the knowledge of those skilled in the art, including the aspects described in the Disclosure of the Invention.
[0527] [8. Other] Furthermore, among the processes described in the above embodiments, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically using known methods. In addition, the information including the processing procedures, specific names, various data, and parameters shown in the above documents and drawings can be changed as desired unless otherwise specified. For example, the various information shown in each drawing is not limited to the information shown in the drawings.
[0528] Furthermore, the components of each device shown in the figure are conceptual functional components and do not necessarily have to be physically configured as shown in the figure. In other words, the specific form of distribution and integration of each device is not limited to that shown in the figure, and all or part of them can be functionally or physically distributed and integrated in any unit depending on various loads, usage conditions, etc.
[0529] Furthermore, the processes described in the above-described embodiments can be combined as appropriate within the scope of not causing any contradiction in the process contents.
[0530] Furthermore, the above-mentioned "section, module, unit" can be read as "means" or "circuit," etc. For example, an acquisition unit can be read as an acquisition means or an acquisition circuit. [Explanation of symbols]
[0531] 1. Information Processing Systems 10 Terminal Equipment 50 Information provision device 100, 100A, 100B, 100C, 100D Information processing equipment 120, 120A, 120B, 120C storage section 121 Object information storage unit 122 Graph information storage unit 123, 123A Quantization information storage unit 124 Codebook information storage unit 125, 125A Blob information storage unit 126, 126A Blob connection graph information storage unit 130, 130A, 130B, 130C, 130D control section 131 Acquisition Department 132, 132A, 132B, 132C generation section 133, 133A, 133B, 133C, 133D Search processing unit (processing unit) 134 Provision Department
Claims
1. an acquisition unit that acquires a group connection graph in which a plurality of groups into which a plurality of objects to be searched for data are classified are connected by edges, a search query for the plurality of objects, and a search group starting set including a group that is a starting point for the search; calculate the distance between each object included in the group of the search group starting set and the search query, determine the group of the search group starting set that is closest to the search query as the nearest group, add the search group starting set other than the nearest group to an excluded group set, add the search group starting set to an unsearched group set and a distance-calculated set, and then start subsequent processing; In the post-processing, for a second group connected by the edge to a first group that is closest to the search query and is within a group search radius of the unsearched group set, if the second group does not exist in the distance-calculated set, add the second group to the distance-calculated set and the unsearched group set, and if the second group is within the group search radius from the search query, update the nearest group to the second group, and perform a search process of a group connection graph that repeats the process of adding the nearest group or the second group to the excluded group set based on the distance between the search query and the nearest group, and the process of updating the group search radius and the group search radius, and add 1 to the group search count in accordance with the execution of the search process of the group connection graph; adding all objects in the nearest group to a search result set, and excluding objects other than a predetermined number of objects closest to the search query from the search result set; updating an object search radius and an object search radius based on a maximum distance between each object in the search result set and the search query; determining whether or not any of termination criteria including at least one of the following is satisfied: when the number of group searches exceeds a predetermined upper limit; when a group distance, which is the distance between the search query and an object included in the nearest group that is closest to the search query, exceeds the object search radius; and when the group search radius exceeds the object search radius; If the termination criterion is not met, update the nearest group to the group in the set of excluded groups that has the shortest distance from the search query, update the group search radius and the group search radius based on the group distance, and return to the search process of the group connection graph to continue the process. On the other hand, if the termination criteria are met, by determining the objects included in the search result set as objects in the vicinity of the search query, a search processing unit that repeatedly executes a search process of the group connection graph and an update process of the search result set based on the searched groups until the termination criterion is satisfied; An information processing device comprising:
2. 1. A computer-implemented information processing method, comprising: an acquisition step of acquiring a group connection graph in which a plurality of groups into which a plurality of objects to be searched for data are classified are connected by edges, a search query for the plurality of objects, and a search group starting set including a group that is a starting point for the search; calculate the distance between each object included in the group of the search group starting set and the search query, determine the group of the search group starting set that is closest to the search query as the nearest group, add the search group starting set other than the nearest group to an excluded group set, add the search group starting set to an unsearched group set and a distance-calculated set, and then start subsequent processing; In the post-processing, for a second group connected by the edge to a first group that is closest to the search query and is within a group search radius of the unsearched group set, if the second group does not exist in the distance-calculated set, add the second group to the distance-calculated set and the unsearched group set, and if the second group is within the group search radius from the search query, update the nearest group to the second group, and perform a search process of a group connection graph that repeats the process of adding the nearest group or the second group to the excluded group set based on the distance between the search query and the nearest group, and the process of updating the group search radius and the group search radius, and add 1 to the group search count in accordance with the execution of the search process of the group connection graph; adding all objects in the nearest group to a search result set, and excluding objects other than a predetermined number of objects closest to the search query from the search result set; updating an object search radius and an object search radius based on a maximum distance between each object in the search result set and the search query; determining whether or not any of termination criteria including at least one of the following is satisfied: when the number of group searches exceeds a predetermined upper limit; when a group distance, which is the distance between the search query and an object included in the nearest group that is closest to the search query, exceeds the object search radius; and when the group search radius exceeds the object search radius; If the termination criterion is not met, update the nearest group to the group in the set of excluded groups that has the shortest distance from the search query, update the group search radius and the group search radius based on the group distance, and return to the search process of the group connection graph to continue the process. On the other hand, if the termination criteria are met, by determining the objects included in the search result set as objects in the vicinity of the search query, a search processing step of repeatedly searching the group connection graph and updating the search result set based on the searched groups until the termination criterion is met; An information processing method comprising:
3. an acquisition step of acquiring a group connection graph in which a plurality of groups into which a plurality of objects to be searched for data are classified are connected by edges, a search query for the plurality of objects, and a search group starting set including a group that is a starting point for the search; calculate the distance between each object included in the group of the search group starting set and the search query, determine the group of the search group starting set that is closest to the search query as the nearest group, add the search group starting set other than the nearest group to an excluded group set, add the search group starting set to an unsearched group set and a distance-calculated set, and then start subsequent processing; In the post-processing, for a second group connected by the edge to a first group that is closest to the search query and is within a group search radius of the unsearched group set, if the second group does not exist in the distance-calculated set, add the second group to the distance-calculated set and the unsearched group set, and if the second group is within the group search radius from the search query, update the nearest group to the second group, and perform a search process of a group connection graph that repeats the process of adding the nearest group or the second group to the excluded group set based on the distance between the search query and the nearest group, and the process of updating the group search radius and the group search radius, and add 1 to the group search count in accordance with the execution of the search process of the group connection graph; adding all objects in the nearest group to a search result set, and excluding objects other than a predetermined number of objects closest to the search query from the search result set; updating an object search radius and an object search radius based on a maximum distance between each object in the search result set and the search query; determining whether or not any of termination criteria including at least one of the following is satisfied: when the number of group searches exceeds a predetermined upper limit; when a group distance, which is the distance between the search query and an object included in the nearest group that is closest to the search query, exceeds the object search radius; and when the group search radius exceeds the object search radius; If the termination criterion is not met, update the nearest group to the group in the set of excluded groups that has the shortest distance from the search query, update the group search radius and the group search radius based on the group distance, and return to the search process of the group connection graph to continue the process. On the other hand, if the termination criteria are met, by determining the objects included in the search result set as objects in the vicinity of the search query, a search processing procedure for repeatedly executing a search process of the group connection graph and an update process of the search result set based on the searched groups until the termination criterion is satisfied; An information processing program characterized by causing a computer to execute the above.
Citation Information
Patent Citations
Conductor for connection of semiconductor device
JP1987093335A
Power source plug slip alarm
JP1988000982A
Ultrasonic transformer
JP1988011000A
Image management device, image management method, program and integrated circuit
JP2014093058A
Information processing device, information processing method, and information processing program
JP2020027590A