Information processing apparatus, information processing method, and information processing program
The information processing apparatus addresses the issue of inappropriate edge deletion in existing graph data generation technologies by executing a controlled shortcut edge deletion process, ensuring accurate and uniform edge removal across nodes.
Patent Information
- Application Number
- JP2023213359
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-30
AI Technical Summary
Existing technologies for generating graph data to support information searches, such as deleting shortcut edges, often result in inappropriate edge deletion, lacking sufficient consideration of processing procedures and determination conditions.
An information processing apparatus that acquires first graph information, selects target edges, and executes a shortcut edge deletion process to generate second information indicating a graph with deleted edges, using a specific determination condition to ensure appropriate edge deletion.
The apparatus effectively generates information indicating graphs with appropriately deleted edges, improving the accuracy and reliability of information processing by uniformly deleting edges across nodes.
Smart Images

Figure 2025097203000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus, an information processing method, and an information processing program.
Background Art
[0002] Conventionally, technologies for searching for various information have been provided. For example, a technology for generating graph data in which nodes corresponding to search targets are connected by edges is provided in order to perform a search for a predetermined target. For example, in Patent Document 1, it is determined whether a directed edge connected from one node to another node corresponds to a shortcut edge, and the directed edge determined to be a shortcut edge is deleted, thereby generating graph data in which an increase in the number of edges is suppressed. Further, such a technology is used, for example, in image search and the like.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Non-Patent Documents
[0004]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, there is room for improvement in the above prior art. For example, in the above prior art, although it is possible to suppress an increase in the number of edges by deleting shortcut edges, there may be cases where too many edges are deleted. For example, in the above prior art, when performing the process of deleting shortcut edges, it is hard to say that the processing procedures and determination conditions are sufficiently considered. Therefore, there is room for improvement in appropriately generating information indicating other graphs in which edges in the graph have been deleted.
[0006] The present application has been made in view of the above, and an object thereof is to provide an information processing apparatus, an information processing method, and an information processing program that appropriately generate information indicating other graphs in which edges in a graph have been deleted.
Means for Solving the Problems
[0007] The information processing apparatus according to the present application includes an acquisition unit that acquires first graph information indicating a first graph in which a plurality of nodes corresponding to each of a plurality of objects to be searched are connected by edges, and among the directed edges starting from each of the plurality of nodes included in the first graph, one directed edge that has not been selected as a processing target is selected as a target edge, and a shortcut edge deletion process is executed in which the selected target edge for each of the plurality of nodes is set as a deletion determination target for a shortcut edge, thereby generating second information indicating a second graph in which directed edges have been deleted from the first graph, and is characterized by including the above.
Effects of the Invention
[0008] According to one aspect of the embodiment, there is an effect that information indicating other graphs in which edges in a graph have been deleted can be appropriately generated.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, embodiments for implementing the information processing apparatus, information processing method, and information processing program according to the present application (hereinafter referred to as "embodiments") will be described in detail with reference to the drawings. Note that the information processing apparatus, information processing method, and information processing program according to the present application are not limited by these embodiments. Also, in the following embodiments, the same parts are denoted by the same reference numerals, and redundant descriptions are omitted.
[0011] (Embodiment) [1. Information Processing] An example of information processing according to the embodiment will be described with reference to FIG. 1. FIG. 1 is a diagram showing an example of information processing according to the embodiment. In FIG. 1, for the edges included in a first graph, which is a certain graph data (also simply referred to as "graph") of the information processing apparatus 100 (see FIG. 5), it is determined whether the determination condition of the shortcut edge is satisfied, and second information is generated which shows a graph (also referred to as "second graph") obtained by deleting the edges determined to correspond to the shortcut edges.
[0012] Note that although details of the second information will be described later, any information can be adopted as long as the second information is information showing the second graph. For example, the second information may be information for identifying the edges corresponding to the shortcut edges, such as a flag associated with each of the edges included in the first graph and indicating whether it is a shortcut edge. Also, the second information may be graph data (graph information) generated separately from the first graph and showing the graph (second graph) after the shortcut edges included in the first graph are deleted. Thus, the second information may be any information as long as it can specify the structure of the graph (second graph) after the shortcut edges included in the first graph are deleted. Note that in the following description, generating the second information showing the second graph may sometimes be described as generating the second graph.
[0013] In addition, FIG. 1 shows a case where target information (object) is vectorized and a graph (graph index) is generated for the vectorized object. That is, FIG. 1 shows a case where the information processing apparatus 100 performs processing using a vector as an object value corresponding to the object.
[0014] Note that the information used by the information processing apparatus 100 is not limited to vectors, and may be information in any form as long as it can represent the similarity of each target. For example, the information processing apparatus 100 may use predetermined data or values corresponding to each target. For example, the information processing apparatus 100 may use predetermined numerical values (e.g., binary values or hexadecimal values) generated from each target. For example, the information processing apparatus 100 may use data in any form as long as the distance (similarity) between the data is defined, not limited to vectors. In addition, hereinafter, a case where image information is used as an object will be described as an example, but the object may be various targets such as video information and audio information.
[0015] In addition, in FIG. 1, the information processing apparatus 100 performs information processing on a graph including nodes and directed edges. Here, the directed edge means an edge that can only be traversed in one direction. Hereinafter, the node from which the edge traverses, that is, the starting node, is referred to as the reference source, and the node to which the edge traverses, that is, the ending node, is referred to as the reference destination. For example, a directed edge connected from a predetermined node "A" to a predetermined node "B" indicates that the reference source is the node "A" and the reference destination is the node "B".
[0016] Hereinafter, an edge that refers to the node "A" as the reference source as described above is called an output edge of the node "A". Also, hereinafter, an edge that refers to the node "B" as the reference destination as described above is called an input edge of the node "B". That is, the output edge and the input edge referred to here are differences in how to capture a single directed edge with respect to the two nodes connected by the directed edge. A single directed edge becomes an output edge and an input edge. That is, the output edge and the input edge are relative concepts. For a single directed edge, it becomes an output edge when captured with the reference source node as the center, and becomes an input edge when captured with the reference destination node as the center. Note that in the present embodiment, since the edges are directed edges such as output edges and input edges, hereinafter, a directed edge may be simply described as an "edge".
[0017] Also, each node referred to here corresponds to each object. For example, each of a plurality of local feature amounts extracted from an image may be an object. Also, for example, various data in which the distance between objects is defined may be an object.
[0018] The information processing apparatus 100 performs graph generation processing on nodes corresponding to huge amounts of image information (for example, several million to several hundred million, etc.) within the range that the information processing apparatus 100 can process. However, only a part of it is illustrated in the drawings. In FIG. 1, for the sake of simplicity of explanation, five nodes are illustrated to explain the outline of the processing. Specifically, in FIG. 1, only the nodes N1 to N5 are illustrated, and only some edges such as the edge E1 are illustrated.
[0019] When described as "node N* (* is an arbitrary numerical value)" in this way, it indicates that the node is a node identified by the node ID "N*". For example, when described as "node N1", the node is a node identified by the node ID "N1".
[0020] Also, when described as "edge E* (* is any numerical value)" in this way, it indicates that the edge is the edge identified by the edge ID "E*". For example, when described as "edge E1", the edge is the edge identified by the edge ID "E1". For example, by the edge E1 that connects node N1 as the reference source and node N2 as the reference destination, it is possible to reach from node N1 to node N2. In this case, when the directed edge E1 is identified with node N1 as the center, it becomes an output edge, and when identified with node N2 as the center, it becomes an input edge. In other words, when the directed edge E1 is viewed from the perspective of node N1, it is an edge with an arrow pointing from itself to another edge, that is, an outward edge, and when viewed from the perspective of node N2, it is an edge with an arrow pointing towards itself, that is, an inward edge. That is, the output edge referred to here can be read as an outward edge, and the input edge can be read as an inward edge.
[0021] Also, the spatial information VS1-0 to VS1-6 shown in FIG. 1 is a diagram schematically showing each graph, and the spaces shown in the spatial information VS1-0 to VS1-6 may be the same space. Also, hereinafter, when explaining the spatial information VS1-0 to VS1-6 without particular distinction, it is described as spatial information VS1.
[0022] For example, each of the circles (〇) in the spatial information VS1 of FIG. 1 represents each node. Each node corresponds to each object. Also, the arrow lines connecting the circles (〇) in the spatial information VS1 represent each directed edge.
[0023] Also, the spatial information VS1 in FIG. 1 may be a Euclidean space. Also, the spatial information VS1 shown in FIG. 1 is a conceptual diagram for explaining the distance between each vector, etc., and the spatial information VS1 is a multi-dimensional space. For example, although the spatial information VS1 shown in FIG. 1 is illustrated in a two-dimensional form for illustration on a plane, it is assumed to be a multi-dimensional space such as 100 dimensions or 1000 dimensions.
[0024] Also, the graphs GR11-1 to GR11-6 shown in FIG. 1 are diagrams schematically showing the generation process of the second graph, and the graphs GR11-1 to GR11-6 are the same second graph generated by information processing. Also, hereinafter, when explaining the graphs GR11-1 to GR11-6 without particular distinction, they will be described as graph GR11. For example, in FIG. 1, when the second information is a flag or the like, the information processing apparatus 100 generates information (such as flag information) for specifying an edge included in the second graph (graph GR11). Also, when the information processing apparatus 100 generates the graph data itself of the second graph, it generates the graph data of the second graph (graph GR11).
[0025] In the present embodiment, the distance between each node in the spatial information VS1 is set as the similarity between the corresponding objects. For example, it is assumed that the similarity of the objects (image information) corresponding to each node is mapped as the distance between the nodes in the spatial information VS1. For example, it is assumed that the similarity between the concepts corresponding to each node is mapped as the distance between each node. Here, in the example shown in FIG. 1, the similarity between objects with a short distance between each node in the spatial information VS1 is high, and the similarity between objects with a long distance between each node in the spatial information VS1 is low. For example, in the spatial information VS1 in FIG. 1, the node (node N2) identified by the node ID "N2" and the node (node N3) identified by the node ID "N3" are close to each other, that is, the distance is short. Therefore, it indicates that the similarity between the object corresponding to the node identified by the node ID "N2" and the object corresponding to the node identified by the node ID "N3" is high.
[0026] Also, for example, in the spatial information VS1 in FIG. 1, the node identified by the node ID "N1" and the node identified by the node ID "N5" are remote from each other, that is, the distance is long. Therefore, it indicates that the similarity between the object corresponding to the node identified by the node ID "N1" and the object corresponding to the node identified by the node ID "N5" is low. Note that the distance as an index indicating similarity may be any distance as long as it can be applied as the distance between vectors (N-dimensional vectors). For example, various distances such as Euclidean distance, Mahalanobis distance, and cosine distance may be used.
[0027] [1-1. Information Processing Example] Hereinafter, an example of the information processing executed by the information processing apparatus 100 will be described with reference to FIG. 1. Specifically, FIG. 1 shows an example of the generation of the second information indicating the second graph by the shortcut edge deletion process of the information processing apparatus 100. Note that each step shown in FIG. 1 is a convenient step for explaining the graph generation process, and the actual process may be a process before that or a more detailed process. Note that the information processing performed by the information processing apparatus 100 may be any processing flow as long as the graph GR11 (second graph) as shown in the graph GR11-6 in FIG. 1 is generated.
[0028] In FIG. 1, the information processing apparatus 100 acquires the first graph GR1 (step S1). Note that the first graph GR1 shown in FIG. 1 is merely an example, and any graph can be adopted as the first graph. Further, hereinafter, the case where the information processing apparatus 100 acquires the first graph GR1 from the storage unit 120 (see FIG. 5) will be described as an example. However, the information processing apparatus 100 may generate the first graph GR1 or may acquire the first graph GR1 from an external apparatus such as the information providing apparatus 50.
[0029] Also, in FIG. 1, as an example, a case where a reference (reference CR11) identified by a reference ID "CR11" among the references stored in the reference information storage unit 123 (see FIG. 8) is used will be described. For example, the reference CR11 is a reference of a selection order (processing order) indicating that edges to be processed (also referred to as "target edges") are selected in order from the shortest edge length.
[0030] First, the information processing apparatus 100 selects, as a target edge, one directed edge that is not selected as a processing target (target edge) among the directed edges starting from each of the plurality of nodes included in the first graph GR1 (step S2). In FIG. 1, in the first graph GR1 in the spatial information VS1-0, since there is no edge that has been processed as a target edge, the edge with the shortest length for each of the nodes N1 to N5 is selected as the target edge.
[0031] For example, the information processing apparatus 100 selects the edge E7 with the shortest length as the target edge for the node N1 among the directed edges starting from the node N1 in the first graph GR1 (edges E1, E3, E7 in FIG. 1). Also, the information processing apparatus 100 selects the edge E5 with the shortest length as the target edge for the node N2 among the directed edges starting from the node N2 in the first graph GR1 (edges E2, E5, E8 in FIG. 1).
[0032] For example, the information processing apparatus 100 selects the edge E4 with the shortest length as the target edge for the node N3 among the directed edges starting from the node N3 in the first graph GR1 (edges E4, E10, E11 in FIG. 1). Similarly, the information processing apparatus 100 selects the edges E6 and E14 with the shortest length as the target edges for the nodes N4 and N5 respectively among the directed edges starting from each of the nodes N4 and N5 in the first graph GR1.
[0033] In the graph GR11-1 in the spatial information VS1-2 in FIG. 1, the information processing apparatus 100 indicates the edges E4, E5, E6, E7, and E14 by dotted lines, indicating a state in which the edges E4, E5, E6, E7, and E14 are selected as target edges. That is, the graph GR11-1 shows a state in which no determination has been made as to whether any of the edges in the first graph are to be deleted as shortcut edges, and the second graph does not include any of the edges.
[0034] The information processing apparatus 100 generates second information based on processing for each of the selected target edges (step S3). In FIG. 1, the information processing apparatus 100 generates second information indicating the graph GR11-2 in the spatial information VS1-2 based on processing for each of the selected edges E4, E5, E6, E7, and E14.
[0035] Here, since the shortest edge cannot be a shortcut, the information processing apparatus 100 determines that the edges E4, E5, E6, E7, and E14 do not correspond to shortcut edges. Then, the information processing apparatus 100 generates a graph GR11-2, which is a second graph in which the edges E4, E5, E6, E7, and E14 are added (left) without deleting the edges E4, E5, E6, E7, and E14 as shortcut edges. Details of the determination conditions for shortcut edges will be described later.
[0036] Next, the information processing apparatus 100 selects, as a target edge, one directed edge that has not been selected as a processing target (target edge) among the directed edges starting from each of the plurality of nodes included in the first graph GR1 (step S4). In FIG. 1, in the first graph GR1 in the spatial information VS1-0, since the shortest edge for each node has already been processed as a target edge, the information processing apparatus 100 selects, as the target edge, the second shortest edge for each of the nodes N1 to N5.
[0037] For example, among the directed edges starting from node N1 in the first graph GR1 (edges E1, E3, E7 in FIG. 1), the information processing apparatus 100 selects the edge E1 with the second shortest length as the target edge for node N1. Also, among the directed edges starting from node N2 in the first graph GR1 (edges E2, E5, E8 in FIG. 1), the information processing apparatus 100 selects the edge E8 with the second shortest length as the target edge for node N2.
[0038] For example, among the directed edges starting from node N3 in the first graph GR1 (edges E4, E10, E11 in FIG. 1), the information processing apparatus 100 selects the edge E10 with the second shortest length as the target edge for node N3. Similarly, among the directed edges starting from each of nodes N4 and N5 in the first graph GR1, the information processing apparatus 100 selects the edges E13 and E12 with the second shortest length as the target edges for nodes N4 and N5, respectively.
[0039] In the graph GR11-3 in the spatial information VS1-3 in FIG. 1, the information processing apparatus 100 indicates the edges E1, E8, E10, E13, and E12 by dotted lines, showing the state in which the edges E1, E8, E10, E13, and E12 are selected as the target edges. That is, the graph GR11-3 shows the state in which the edges E4, E5, E6, E7, and E14 in the first graph are added to the second graph by the process of step S3, and the second graph includes the edges E4, E5, E6, E7, and E14.
[0040] The information processing apparatus 100 generates second information based on the process for each of the selected target edges (step S5). In FIG. 1, the information processing apparatus 100 generates second information indicating the graph GR11-4 in the spatial information VS1-4 based on the process for each of the selected edges E1, E8, E10, E13, and E12.
[0041] Here, in graph GR11-2 which is the second graph before the execution of step S5, since none of the edges E1, E8, E10, E13, E12 have a detour path yet (each target edge does not become a shortcut edge), the information processing apparatus 100 determines that the edges E1, E8, E10, E13, E12 do not correspond to shortcut edges. Then, the information processing apparatus 100 generates graph GR11-4 which is the second graph with edges E1, E8, E10, E13, E12 added (left) without deleting the edges E1, E8, E10, E13, E12 as shortcut edges.
[0042] Next, the information processing apparatus 100 selects, as a target edge, one unselected directed edge among the directed edges starting from each of the plurality of nodes included in the first graph GR1 (step S6). In FIG. 1, in the first graph GR1 in the spatial information VS1-0, since the shortest edge and the second shortest edge for each node have already been processed as target edges, the information processing apparatus 100 selects, as the target edge for each of nodes N1 to N5, the edge with the third shortest length.
[0043] For example, the information processing apparatus 100 selects edge E3, which has the third shortest length, as the target edge for node N1 among the directed edges starting from node N1 in the first graph GR1 (edges E1, E3, E7 in FIG. 1). Also, the information processing apparatus 100 selects edge E2, which has the third shortest length, as the target edge for node N2 among the directed edges starting from node N2 in the first graph GR1 (edges E2, E5, E8 in FIG. 1).
[0044] For example, the information processing apparatus 100 selects, as the target edge for node N3, the edge E11 which is the third shortest in length among the directed edges starting from node N3 in the first graph GR1 (edges E4, E10, E11 in FIG. 1). Similarly, the information processing apparatus 100 selects, as the target edges for nodes N4 and N5 respectively, the edges E9 and E15 which are the third shortest in length among the directed edges starting from nodes N4 and N5 in the first graph GR1.
[0045] In the graph GR11-5 in the spatial information VS1-5 in FIG. 1, the information processing apparatus 100 indicates the edges E2, E3, E9, E11, E15 by dotted lines, showing the state in which the edges E2, E3, E9, E11, E15 are selected as the target edges. That is, the graph GR11-5 shows the state in which the edges E1, E8, E10, E13, E12 in the first graph are added to the second graph by the process of step S5, and the second graph includes the edges E1, E4, E5, E6, E7, E8, E10, E13, E14, E12.
[0046] The information processing apparatus 100 generates second information based on the process for each of the selected target edges (step S7). In FIG. 1, the information processing apparatus 100 generates second information indicating the graph GR11-6 in the spatial information VS1-6 based on the process for each of the selected edges E2, E3, E9, E11, E15.
[0047] Here, in graph GR11-4, which is the second graph before the execution of step S7, there is a path that bypasses edge E3. In FIG. 1, from node N1, which is the starting point of edge E3, to node N3, one can reach node N2 via edge E1 from node N1 and then reach node N3 via edge E5 from node N2. Thus, there is a path (route) other than edge E3 from node N1 to node N3. In this way, it is possible to bypass edge E3 by passing through one node N2 from node N1 to node N3. Also, in graph GR11-4, which is the second graph before the execution of step S7, from node N5 to node N2, it is possible to bypass edge E15 by passing through edges E12 and E4 in this order, that is, by passing through one node N3, and there is a path that bypasses edge E15.
[0048] When there is a path that bypasses a target edge, such as the above-mentioned edges E3 and E15, the information processing apparatus 100 determines whether to delete the target edge as a shortcut edge. The information processing apparatus 100 may determine whether there is a path that bypasses the target edge by any process.
[0049] In FIG. 1, for example, the information processing apparatus 100 extracts a node (also referred to as a "bypass node") that has an edge with node N3 as the output destination (end point) among the nodes to which other edges starting from node N1, which is the starting point of edge E3, are input. Since the information processing apparatus 100 extracts node N2 as the bypass node of edge E3, it determines that there is a path that bypasses edge E3 and makes a determination as to whether to delete edge E3 as a shortcut edge.
[0050] Also, in FIG. 1, as an example, the case where the condition (condition CD1) identified by the condition ID "CD1" among the criteria stored in the condition information storage unit 122 (see FIG. 7) is used will be described. For example, the condition CD1 is a condition as shown in the condition information CND1 in FIG. 2. FIG. 2 is a conceptual diagram showing an example of the determination of a shortcut according to the embodiment. Note that the condition is not limited to the condition shown in FIG. 2, and any condition may be used, but this will be described later.
[0051] First, the outline of the condition CD1 shown in the condition information CND1 in FIG. 2 will be described using the schematic diagram shown on the left side of FIG. 2. In FIG. 2, the node that is the start point of the target edge (also referred to as the "first node") is denoted as node Nsrc, the node that is the end point of the target edge (also referred to as the "second node") is denoted as node Ndst, and the detour node (also referred to as the "third node") on the path that detours the target edge is denoted as node Npass for explanation. Also, in FIG. 2, the target edge (also referred to as the "first directed edge") is denoted as edge E1st, the edge whose start point is the first node and whose end point is the third node (also referred to as the "second directed edge") is denoted as edge E2nd, and the edge whose start point is the third node and whose end point is the second node (also referred to as the "third directed edge") is denoted as edge E3rd for explanation.
[0052] The range AR11 shown in FIG. 2 indicates a range centered on the node Nsrc which is the first node and having a radius of the distance to the node Ndst which is the second node (also referred to as the "first distance"). The distance (first distance) between the node Nsrc and the node Ndst corresponds to the length of the edge E1st which is the first directed edge.
[0053] The condition (also referred to as the "first condition") indicated by "distance(Nsrc, Npass) < distance(Nsrc, Ndst)" among the condition information CND1 in FIG. 2 corresponds to the range AR11. Among the first conditions in the condition information CND1 in FIG. 2, "distance(Nsrc, Npass)" corresponds to the distance (also referred to as the "second distance") between the node Nsrc (first node) and the node Npass which is the third node. The distance (second distance) between the node Nsrc and the node Npass corresponds to the length of the edge E2nd which is the second directed edge.
[0054] Also, among the first conditions in the condition information CND1 in FIG. 2, "distance (Nsrc, Ndst)" corresponds to the distance (first distance) between the node Nsrc (first node) and the node Ndst (second node). That is, when the node Npass (third node) is within the range AR11, it indicates that the target edge (edge E1st in FIG. 2) satisfies the first condition among the condition information CND1.
[0055] Also, the range AR21 shown in FIG. 2 indicates a range centered on the node Ndst (second node) with the distance (first distance) to the node Nsrc (first node) as the radius.
[0056] The condition indicated by "distance (Npass, Ndst) < distance (Nsrc, Ndst)" among the condition information CND1 in FIG. 2 (also referred to as the "second condition") corresponds to the range AR21. Among the second conditions in the condition information CND1 in FIG. 2, "distance (Npass, Ndst)" corresponds to the distance (also referred to as the "third distance") between the node Npass (third node) and the node Ndst (second node). The distance (third distance) between the node Npass and the node Ndst corresponds to the length of the edge E3rd which is the third directed edge.
[0057] Also, among the second conditions in the condition information CND1 in FIG. 2, "distance (Nsrc, Ndst)" corresponds to the distance (first distance) between the node Nsrc (first node) and the node Ndst (second node). That is, when the node Npass (third node) is within the range AR21, it indicates that the target edge (edge E1st in FIG. 2) satisfies the second condition among the condition information CND1.
[0058] In addition, the range AR31 shown in FIG. 2 indicates a range obtained by multiplying, by a predetermined coefficient α, the range with the distance between the node Nsrc (first node) and the node Ndst (second node) as the diameter. For example, the range AR31 shown in FIG. 2 indicates a range obtained by expanding, by a factor of α, the range passing through the nodes Nsrc and Ndst. Note that any value can be set for the predetermined coefficient α. For example, α may be 1, or a value greater than 1 such as 1.1 may be set, or a value less than 1 such as 0.9 may be set.
[0059] Among the condition information CND1 in FIG. 2, the condition indicated by "distance(Nsrc, Npass)+distance(Npass, Ndst)<α*distance(Nsrc, Ndst)" (also referred to as the "third condition") corresponds to the range AR21. Among the third condition in the condition information CND1 in FIG. 2, "distance(Nsrc, Npass)" corresponds to the distance (second distance) between the node Nsrc (first node) and the node Npass which is the third node. Among the third condition in the condition information CND1 in FIG. 2, "distance(Npass, Ndst)" corresponds to the distance (third distance) between the node Npass (third node) and the node Ndst (second node).
[0060] In addition, among the third condition in the condition information CND1 in FIG. 2, "α*distance(Nsrc, Ndst)" is obtained by multiplying the distance (first distance) between the node Nsrc (first node) and the node Ndst (second node) by a predetermined coefficient α. That is, when the node Npass (third node) is within the range AR21, it indicates that the target edge (edge E1st in FIG. 2) satisfies the third condition among the condition information CND1.
[0061] The corresponding range CA1 indicated by the hatching in FIG. 2 is the range where range AR11, range AR21, and range AR31 overlap, and corresponds to the range that satisfies the condition CD1 indicated by the condition information CND1. That is, when the node Npass, which is the third node, is within the corresponding range CA1, it indicates that the target edge (edge E1st in FIG. 2) satisfies the condition CD1 indicated by the condition information CND1. Thus, the condition CD1 indicated by the condition information CND1 in FIG. 2 is the condition for determining that the target edge corresponds to a shortcut edge when it satisfies all of the above-described first condition, second condition, and third condition.
[0062] In FIG. 1, the information processing apparatus 100 determines whether or not a target edge for which it has determined that there is a detour path satisfies the above-described condition CD1. For example, when the information processing apparatus 100 determines that a certain target edge satisfies the above-described condition CD1, it determines that the target edge corresponds to a shortcut edge. For example, when the information processing apparatus 100 determines that a certain target edge does not satisfy the above-described condition CD1, it determines that the target edge does not correspond to a shortcut edge. The information processing apparatus 100 deletes a target edge determined to correspond to a shortcut edge as a shortcut edge to generate a second graph.
[0063] Here, in FIG. 1, when node N1 is the first node, node N3 is the second node, and node N2 is the third node, edge E3 becomes the first directed edge, edge E1 becomes the second directed edge, and edge E5 becomes the third directed edge. The information processing apparatus 100 uses the length of edge E3, which is the first distance, the length of edge E1, which is the second distance, the length of edge E5, which is the third distance, and the determination formula indicated by the condition information CND1 to determine whether or not to delete edge E3 as a shortcut edge. In FIG. 1, since edge E3 satisfies the determination formula indicated by the condition information CND1, the information processing apparatus 100 determines to delete edge E3 as a shortcut edge. Similarly, the information processing apparatus 100 determines to delete edge E15 as a shortcut edge since edge E15 satisfies the determination formula indicated by the condition information CND1.
[0064] In addition, the information processing apparatus 100 determines that the other target edges, namely edges E2, E9, and E11, do not correspond to shortcut edges. For example, edge E11 that leads from node N3 to node N5 is an edge (shortcut) for which there exists a path (detour path) that reaches node N5 by traversing edges E10 and E13 from node N3, but the information processing apparatus 100 determines not to delete it as a shortcut edge. For example, when the information processing apparatus 100 processes edge E11 as a target edge, since edge E11 does not satisfy the determination formula shown in the condition information CND1, the information processing apparatus 100 determines not to delete edge E11 as a shortcut edge.
[0065] Then, the information processing apparatus 100 deletes edges E3 and E15 as shortcut edges, and generates a graph GR11-6, which is a second graph with edges E2, E9, and E11 added (left intact) without deleting edges E2, E9, and E11 as shortcut edges. As a result, in FIG. 1, the information processing apparatus 100 generates a graph GR11, which is a second graph with edge E1 deleted as a shortcut edge from graph GR1, which is the first graph.
[0066] 〔1-2. Effects, etc.〕 In this way, the information processing apparatus 100 can appropriately generate information indicating another graph in which an edge in the graph has been deleted by generating second information indicating a second graph in which the target edge determined to be a shortcut edge has been deleted from the first graph.
[0067] For example, regarding the graph GR1 shown in FIG. 1, when processing is performed in node order, there may be a situation where the edge E1 is not deleted as a shortcut edge. For example, when selecting one node in the order of nodes N1 to N5, processing all the edges starting from the selected node as target edges, and then selecting the next node, there may be a situation where the edge E1 is not deleted as a shortcut edge. Specifically, when first selecting node N1 and processing all the edges E1, E3, and E7 starting from the selected node N1 as target edges, at the stage where edge E3 is the target edge, since there is no detour path for the target edge E3, the edge E3 will not be deleted as a shortcut edge. Then, the edges of the nodes selected later are more likely to have detour paths, and the edges of the nodes selected later may be deleted as shortcut edges, resulting in a bias in the graph structure.
[0068] On the other hand, as described above, the information processing apparatus 100 can reduce the possibility that the edges of a specific node are deleted as shortcut edges by selecting one edge from each node and performing processing on that edge as a target edge, and can uniformly delete edges for each node. Thereby, the information processing apparatus 100 can generate an appropriate graph.
[0069] As described above, the reduction of shortcut edges results in different generated graphs depending on the order. When performing reduction processing starting from the longest edges at each node, longer edges will be deleted compared to the case of performing reduction processing starting from the shortest edges. Also, when performing reduction processing on all the edges of each node unit, the edges of the nodes at the beginning of the reduction processing will be reduced more. Therefore, in the example described above, the information processing apparatus 100 performs reduction processing on the shortest edges of all the nodes first, and then performs reduction processing on the second shortest edges of all the nodes, and repeats this sequentially. Thereby, the information processing apparatus 100 can make the number of reduced edges per node equal.
[0070] In addition, in vector approximation neighborhood search using a graph, edges can be reduced by deleting shortcut edges, and performance improvement can be achieved by reducing the number of edges to be referred to during search. Based on the conditions as described above, the information processing apparatus 100 can improve the search performance using the generated graph by optimizing the reduction range of shortcut edges. Further, the information processing apparatus 100 can generate a flag as the second information, have a flag indicating the presence or absence of each edge for all edges, and perform deletion processing in one graph by discriminating (controlling) the presence or absence of the edge with that flag.
[0071] [1-3. Conditions, criteria, etc.] In the above-described example, the case of using the conditions shown in FIG. 2 was shown, but the conditions shown in FIG. 2 are merely an example of the determination conditions for shortcut edges, and various conditions may be used. Examples in this regard will be described below. Note that descriptions of the same points as those described above will be omitted as appropriate.
[0072] First, the conditions shown in FIG. 3 will be described. FIG. 3 is a conceptual diagram showing an example of the determination of a shortcut according to the embodiment. Note that descriptions of the same points as those in FIG. 2 will be omitted as appropriate. For example, the node Nsrc, node Ndst, node Npass, edge E1st, edge E2nd, edge E3rd, range AR11, and range AR21 in FIG. 3 are the same as those in FIG. 2, and thus the description thereof will be omitted.
[0073] FIG. 3 shows a case where, among the criteria stored in the condition information storage unit 122 (see FIG. 7), the condition (condition CD2) identified by the condition ID "CD2" is used. The outline of the condition CD2 shown in the condition information CND2 in FIG. 3 will be described using the schematic diagram shown on the left side of FIG. 3.
[0074] The condition CD2 (also referred to as the "fourth condition") shown in the condition information CND2 of FIG. 3 is a condition using a determination formula based on the comparison between the value calculated using the cosine theorem TR1 (also referred to as the "calculated value") and a threshold value. In FIG. 3, the case where the condition is satisfied when the calculated value calculated using the cosine theorem TR1 is greater than the threshold value is shown. Note that any value can be set as the threshold value. For example, any value within the range of values that the calculated value calculated using the cosine theorem TR1 can take can be set as the threshold value.
[0075] Among the condition information CND2 of FIG. 3, "distance (Nsrc, Ndst)" corresponds to the distance (first distance) between the node Nsrc (first node) and the node Ndst (second node). The distance (first distance) between the node Nsrc and the node Ndst corresponds to the length of the edge E1st which is the first directed edge.
[0076] Among the condition information CND2 of FIG. 3, "distance (Nsrc, Npass)" corresponds to the distance (second distance) between the node Nsrc (first node) and the node Npass (third node). The distance (second distance) between the node Nsrc and the node Npass corresponds to the length of the edge E2nd which is the second directed edge.
[0077] Among the condition information CND2 of FIG. 3, "distance (Npass, Ndst)" corresponds to the distance (third distance) between the node Npass (third node) and the node Ndst (second node). The distance (third distance) between the node Npass and the node Ndst corresponds to the length of the edge E3rd which is the third directed edge.
[0078] For example, the information processing apparatus 100 calculates, as a calculated value, the value of the cosine of the angle formed by the second directed edge (edge E2nd) corresponding to the second distance and the third directed edge (edge E3rd) corresponding to the third distance using the cosine theorem TR1. In FIG. 3, the information processing apparatus 100 calculates the cosine value (cosine value) of the target angle TC in FIG. 3 using the first distance between the node Nsrc and the node Ndst, the second distance between the node Nsrc and the node Npass, and the third distance between the node Npass and the node Ndst.
[0079] The range AR41 shown in FIG. 3 corresponds to the range that satisfies the condition CD2 shown in the condition information CND2. That is, when the node Npass, which is the third node, is within the range AR41, it indicates that the target edge (edge E1st in FIG. 3) satisfies the condition CD2 shown in the condition information CND2.
[0080] In FIG. 3, the case where the range AR41 is included in the range AR11 is illustrated. However, the range AR41 may partially overlap with the range AR11, and the range other than that part does not have to overlap with the range AR11. Also, in FIG. 3, the case where the range AR41 is included in the range AR21 is illustrated. However, the range AR41 may partially overlap with the range AR21, and the range other than that part does not have to overlap with the range AR21.
[0081] In FIG. 1, the information processing apparatus 100 determines whether or not a target edge for which it has determined that there is a detour path satisfies the above-described condition CD2. For example, when the information processing apparatus 100 determines that a certain target edge satisfies the above-described condition CD2, it determines that the target edge corresponds to a shortcut edge. For example, when the information processing apparatus 100 determines that a certain target edge does not satisfy the above-described condition CD2, it determines that the target edge does not correspond to a shortcut edge. The information processing apparatus 100 deletes the target edge determined to correspond to the shortcut edge as the shortcut edge to generate a second graph.
[0082] Note that the conditions shown in FIGS. 2 and 3 are merely examples, and the information processing apparatus 100 may use any conditions. For example, the information processing apparatus 100 may use as a condition that the ranges AR11, AR21, and AR41 (the fourth condition) overlap. That is, in the example of FIG. 1, the information processing apparatus 100 may use the third condition (range AR31) among the conditions CD1 shown in the condition information CND1 in FIG. 2 as the fourth condition (range AR41 shown in FIG. 3).
[0083] In this way, the information processing apparatus 100 may determine that the target edge corresponds to a shortcut edge when all three of the first condition, the second condition, and the fourth condition are satisfied. Since the processing is the same as that described in FIG. 1 except for the different conditions, detailed description thereof is omitted.
[0084] As described above, the information processing apparatus 100 can appropriately determine a shortcut edge by determining whether the target edge corresponds to a shortcut edge based on the relationships among the three edges, i.e., the first directed edge, the second directed edge, and the third directed edge, which are target edges. Also, the information processing apparatus 100 can perform a flexible determination and appropriately determine a shortcut edge by combining a plurality of conditions to determine whether the target edge corresponds to a shortcut edge. Then, the information processing apparatus 100 generates a graph with the shortcut edge deleted based on the determination of the shortcut edge. Thereby, the information processing apparatus 100 can appropriately generate information indicating another graph in which the edge in the graph has been deleted.
[0085] Note that the above-described example shows the case where the target edges are selected in order from the shortest distance, but this is only an example of the above criterion, and various criteria may be used. Examples in this regard are described below.
[0086] For example, the information processing apparatus 100 may select the target edges in order from the longest distance. In this case, the information processing apparatus 100 sets each of the nodes included in the first graph as a starting point, and selects the longest directed edge among the unselected directed edges to be processed and executes the shortcut edge deletion process.
[0087] Also, for example, the information processing apparatus 100 may select the target edges based on a criterion other than the distance. In this case, the information processing apparatus 100 sets each of the nodes included in the first graph as a starting point, and may randomly select a directed edge from the unselected directed edges to be processed and execute the shortcut edge deletion process.
[0088] Note that when generating the first graph, the information processing apparatus 100 may generate a k-nearest neighbor graph as the first graph, or may generate an approximate k-nearest neighbor graph as the first graph. For example, a k-nearest neighbor graph is a graph in which edges to k nodes are connected in order from the nodes closer to each node. For example, an approximate k-nearest neighbor graph is a concept including a graph (also referred to as "ANNG") that approximates a k-nearest neighbor graph generated by performing a k-nearest neighbor search using the graph being generated during graph index (graph) generation, etc. Note that the approximate k-nearest neighbor graph generated by the above-described processing may be a k-nearest neighbor graph. That is, the approximate k-nearest neighbor graph is a concept including the k-nearest neighbor graph. Also, for the generation of the approximate k-nearest neighbor graph, any processing as disclosed in Patent Document 1, Non-Patent Document 1, etc. can be adopted, and detailed description thereof is omitted.
[0089] Also, in the search process, the information processing apparatus 100 may use the starting point information GINF11 regarding the tree structure (tree structure) as shown in FIG. 12 as the starting point information (starting point index). FIG. 12 is a diagram showing an example of the starting point information used in the information processing according to the embodiment. For example, the starting point information GINF11 is an index having a tree structure reachable to the nodes in the first graph GR1. Note that the starting point information such as the starting point information GINF11 may be generated by the information processing apparatus 100, or the information processing apparatus 100 may acquire the starting point information from another external device such as the information providing apparatus 50.
[0090] When the information processing apparatus 100 acquires the starting point information from another external device, the information processing apparatus 100 provides a graph to the other external device. Then, the information processing apparatus 100 acquires the starting point information generated by the other external device that has received the graph from the other external device. For example, when the information processing apparatus 100 acquires the starting point information GINF11 from the information providing apparatus 50, the information processing apparatus 100 transmits the first graph GR1 to the information providing apparatus 50. Then, the information processing apparatus 100 acquires the starting point information GINF11 generated by the information providing apparatus 50 that has received the first graph GR1 from the information providing apparatus 50.
[0091] Further, the information processing apparatus 100 may determine a start node using start information GINF11 as shown in the start information GINF11 in FIG. 12. In the example of FIG. 12, the information processing apparatus 100 determines a start node corresponding to the query QE1 based on the start information GINF11. The query QE1 may be, for example, a node corresponding to an object to be newly added or a target for performing a search using the first graph GR1. That is, the information processing apparatus 100 determines a start node using the start information GINF11 during graph generation or search.
[0092] Specifically, the information processing apparatus 100 determines a start node using the start information GINF11 stored in the storage unit 120 (see FIG. 5). For example, the information processing apparatus 100 determines (identifies) a start node that is a candidate in the vicinity of the start information GINF11 by tracing the start information GINF11 from top (root RT) to bottom based on the query QE1. Thereby, the information processing apparatus 100 can efficiently determine a start node corresponding to the search query (query QE1). For example, the information processing apparatus 100 can quickly determine an appropriate start node corresponding to the query QE1 which is the target node.
[0093] Note that the information processing apparatus 100 is not limited to the above, and may use various start indexes. That is, the start information (start index) shown in the example of FIG. 12 is an example, and the information processing apparatus 100 may search graph information using various start information. The information processing apparatus 100 may generate a start index used for determining a start node at the time of search. For example, the information processing apparatus 100 generates a search index (start information) for quickly searching for a high-dimensional vector. The high-dimensional vector mentioned here may be, for example, a vector of several hundred dimensions to several thousand dimensions, or a vector of more dimensions. Note that the start index as described above is an example, and the information processing apparatus 100 may generate a start index having any data structure as long as it can quickly identify a query in the graph.
[0094] 〔2. Configuration of the Information Processing System〕 As shown in FIG. 4, the information processing system 1 includes a terminal device 10, an information providing device 50, and an information processing device 100. The terminal device 10, the information providing device 50, and the information processing device 100 are communicably connected by wire or wirelessly via a predetermined network N. FIG. 4 is a diagram showing a configuration example of the information processing system according to the embodiment. Note that the information processing system 1 shown in FIG. 4 may include a plurality of terminal devices 10, a plurality of information providing devices 50, and a plurality of information processing devices 100.
[0095] The terminal device 10 is an information processing device used by a user. The terminal device 10 receives various operations by the user. Note that hereinafter, the terminal device 10 may be referred to as the user. That is, hereinafter, the user can also be read as the terminal device 10. Note that the above-described terminal device 10 is realized by, for example, a smartphone, a tablet terminal, a notebook PC (Personal Computer), a desktop PC, a mobile phone, a PDA (Personal Digital Assistant), or the like.
[0096] The information providing device 50 is an information processing device in which information for providing various information to a user or the like is stored. For example, the information providing device 50 stores an object ID based on character information or the like collected from various external devices such as a web server. For example, the information providing device 50 is an information processing device that provides an image search service to a user or the like. For example, the information providing device 50 stores each information for providing an image search service. For example, the information providing device 50 provides vector information corresponding to an image that is the target of the image search service to the information processing device 100. Further, the information providing device 50 transmits a query to the information processing device 100, and thereby receives an object ID or the like indicating an image corresponding to the query from the information processing device 100.
[0097] The information processing apparatus 100 is a computer that executes a generation process for generating information related to a graph. The information processing apparatus 100 is a generation apparatus that generates a second graph using the first graph. The information processing apparatus 100 selects, as a target edge, one unselected directed edge to be processed among the directed edges starting from each of the plurality of nodes included in the first graph, and executes a shortcut edge deletion process in which the selected target edge for each of the plurality of nodes is a target for deletion determination as a shortcut edge, thereby generating second information indicating a second graph from which directed edges have been deleted from the first graph.
[0098] The information processing apparatus 100 executes a shortcut edge deletion process using the determination condition of the shortcut edge. When the relationship among a first directed edge starting from a first node and ending at a second node, a second directed edge starting from the first node and ending at a third node different from the second node, and a third directed edge starting from the third node and ending at the second node among the edges included in the first graph satisfies the determination condition of the shortcut edge, the information processing apparatus 100 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge.
[0099] For example, when the information processing apparatus 100 receives query information (hereinafter, also simply referred to as "query") from a terminal device, it searches for an object (such as vector information) similar to the query and provides the search result to the terminal device. Also, for example, the data provided by the information processing apparatus 100 to the terminal device may be the data itself such as image information, or may be information for referring to corresponding data such as a URL (Uniform Resource Locator). Also, the query and the data to be searched may be any type of data such as images, audio, and text data. In the present embodiment, a case where the information processing apparatus 100 searches for an image will be described as an example.
[0100] [3. Configuration of Information Processing Apparatus] Next, with reference to FIG. 5, the configuration of the information processing apparatus 100 according to the embodiment will be described. FIG. 5 is a diagram showing a configuration example of the information processing apparatus 100 according to the embodiment. As shown in FIG. 5, the information processing apparatus 100 includes a communication unit 110, a storage unit 120, and a control unit 130. Note that the information processing apparatus 100 may include an input unit (for example, a keyboard, a mouse, etc.) that receives various operations from an administrator or the like of the information processing apparatus 100, and a display unit (for example, a liquid crystal display, etc.) that displays various information.
[0101] (Communication Unit 110) The communication unit 110 is realized by, for example, a NIC (Network Interface Card) or the like. Then, the communication unit 110 is connected to a network (for example, the network N in FIG. 4) by wire or wirelessly, and transmits and receives information to and from the terminal device 10 and the information providing device 50.
[0102] (Storage Unit 120) The storage unit 120 is realized by, for example, a semiconductor memory element such as a RAM (Random Access Memory) or a flash memory, or a storage device such as a hard disk or an optical disk. The storage unit 120 according to the embodiment includes, as shown in FIG. 5, an object information storage unit 121, a condition information storage unit 122, a reference information storage unit 123, and a graph information storage unit 124.
[0103] (Object Information Storage Unit 121) The object information storage unit 121 according to the embodiment stores various information related to the object. For example, the object information storage unit 121 stores an object ID and vector data. FIG. 6 is a diagram showing an example of the object information storage unit according to the embodiment. The object information storage unit 121 shown in FIG. 6 includes items such as "object ID" and "vector information".
[0104] "Object ID" indicates identification information for identifying an object. Also, "vector information" indicates vector information corresponding to the object identified by the Object ID. That is, in the example of FIG. 6, vector data (vector information) corresponding to the object is associated and registered with the Object ID for identifying the object.
[0105] For example, in the example of FIG. 6, it shows that the object (target) identified by the ID "OB1" is associated with multi-dimensional vector information of "10, 24, 51, 2...".
[0106] Note that the object information storage unit 121 is not limited to the above, and may store various information according to the purpose.
[0107] (Condition information storage unit 122) The condition information storage unit 122 according to the embodiment stores various information related to conditions regarding the difficulty of retrieval. FIG. 7 is a diagram showing an example of the condition information storage unit according to the embodiment. The condition information storage unit 122 shown in FIG. 7 has items such as "Condition ID" and "Condition information".
[0108] "Condition ID" indicates information for identifying a condition. "Condition information" stores information used for determining (deciding) whether to delete an edge as a shortcut edge. FIG. 7 shows an example where conceptual information such as "CND1" is stored in "Condition information", but actually, information indicating specific conditions such as a judgment formula as shown in parentheses, or a file path name indicating its storage location, etc. is stored.
[0109] In the example of FIG. 7, the condition (condition CD1) identified by the condition ID "CD1" indicates that it is a condition as shown in the condition information CND1. For example, condition CD1 is that the second distance is smaller than the first distance, the third distance is smaller than the first distance, and the sum of the second distance and the third distance is smaller than the value obtained by multiplying the first distance by a predetermined coefficient α. That is, condition CD1 is a condition for determining that the target edge is a shortcut edge when all of the conditions that the second distance is smaller than the first distance, the third distance is smaller than the first distance, and the sum of the second distance and the third distance is smaller than the value obtained by multiplying the first distance by a predetermined coefficient α are satisfied.
[0110] Also, the condition (condition CD2) identified by the condition ID "CD2" indicates that it is a condition as shown in the condition information CND2. For example, condition CD2 is that the calculated value calculated by the cosine theorem is larger than the threshold value. That is, condition CD2 is a condition for determining that the target edge is a shortcut edge when the calculated value calculated based on the first distance, the second distance, and the third distance and the cosine theorem is larger than the threshold value. For example, condition CD2 is a condition for determining that the target edge is a shortcut edge when the calculated value, which is the value of the cosine of the angle formed by the second directed edge corresponding to the second distance and the third directed edge corresponding to the third distance, is larger than the threshold value. For example, condition CD2 is a condition using the determination formula shown in the condition information CND2 in FIG. 1.
[0111] Note that the condition information storage unit 122 is not limited to the above, and may store various information according to the purpose. For example, the condition information storage unit 122 may store information indicating which condition to use. For example, the condition information storage unit 122 may store information indicating the designation of the condition to be used set by the administrator or the like of the information processing apparatus 100.
[0112] (Reference information storage unit 123) The reference information storage unit 123 according to the embodiment stores various information related to the reference of various processes. For example, the reference information storage unit 123 stores various information related to the reference of the selection order of edges. FIG. 8 is a diagram showing an example of the reference information storage unit according to the embodiment. The reference information storage unit 123 shown in FIG. 8 includes items such as "reference ID", "target", and "reference content". For example, the reference information storage unit 123 stores various information related to the reference and conditions for executing various processes.
[0113] "Reference ID" indicates information for identifying a reference. "Target" indicates the target of the reference for the selection order of edges. Also, "reference content" indicates the specific content used as the corresponding reference. In FIG. 8, although the "reference content" is illustrated with abstract symbols such as "CINF11", "CINF12", and "CINF13", it shall be information, conditional expressions, etc. that are specific references.
[0114] In FIG. 8, the reference (reference CR11) identified by the reference ID "CR11" indicates that it is a reference for the selection order based on distance. The target of reference CR11 is an edge, and its reference content is shown as "CINF11". The reference content CINF11 in FIG. 8 indicates that the selection order of edges is in ascending order, that is, it is determined in order from the shorter edge length.
[0115] In FIG. 8, the reference (reference CR12) identified by the reference ID "CR12" indicates that it is a reference for the selection order based on distance. The target of reference CR12 is an edge, and its reference content is shown as "CINF12". The reference content CINF12 in FIG. 8 indicates that the selection order of edges is in descending order, that is, it is determined in order from the longer edge length.
[0116] In FIG. 8, the reference (reference CR13) identified by the reference ID "CR13" indicates that there is no specific target. The reference content of reference CR13 is shown as "CINF13". The reference content CINF13 in FIG. 8 indicates that the selection order of edges is determined randomly.
[0117] Note that the reference information storage unit 123 is not limited to the above, and may store various information according to the purpose. For example, the reference information storage unit 123 may store information indicating which reference is used. For example, the reference information storage unit 123 may store information indicating the designation of the reference set by the administrator or the like of the information processing apparatus 100.
[0118] (Graph information storage unit 124) The graph information storage unit 124 according to the embodiment stores various information related to a graph (graph data). For example, the graph information storage unit 124 stores graph information. FIG. 9 shows a case where the graph information storage unit 124 stores a first graph and second information. For example, in FIG. 9, a case where the graph information storage unit 124 stores a graph (first graph) that is a target of the shortcut edge deletion process and a "flag" that is second information corresponding to the first graph is shown as an example. For example, the graph information storage unit 124 stores graph data of a k-nearest neighbor graph as the first graph. For example, the graph information storage unit 124 stores graph data of an approximate k-nearest neighbor graph as the first graph. For example, the graph information storage unit 124 stores the data of the first graph GR1 in FIG. 1.
[0119] Note that the information processing apparatus 100 may store a second graph separately from the first graph. In this case, the storage unit 120 may have a second graph information storage unit that stores the second graph, and the graph information storage unit 124 may not store the second information (information corresponding to the item "flag" in FIG. 9). The second graph information storage unit stores a graph (second graph) in which an edge (edge whose flag is "0" in FIG. 9) determined to correspond to a shortcut edge has been deleted from the first graph described in the graph information storage unit 124.
[0120] FIG. 9 is a diagram showing an example of the graph information storage unit according to the embodiment. The graph information storage unit 124 shown in FIG. 9 has items such as "node ID", "object ID", and "directed edge information". Further, the "directed edge information" includes information such as "edge ID", "reference destination", and "flag (second information)".
[0121] "Node ID" indicates identification information for each node (target) in the graph data. Also, "Object ID" indicates identification information for identifying an object.
[0122] Also, "Directed edge information" indicates information regarding the edges connected to the corresponding node. In the example of FIG. 9, "Directed edge information" indicates information regarding the output edges output from the corresponding node. Also, "Edge ID" indicates identification information for identifying the edges connecting nodes. Also, "Reference destination" indicates information indicating the reference destination (node) connected by the edge. That is, in the example of FIG. 9, for the node ID that identifies a node, information for identifying the object (target) corresponding to that node and the reference destination (node) to which the directed edge (output edge) from that node is connected are registered in association with each other.
[0123] Also, "Flag (second information)" indicates the second information generated by the shortcut edge deletion process. For example, "Flag (second information)" indicates whether the edge is valid or invalid in the second graph. For example, "Flag (second information)" indicates the case where "1" is assigned when the edge is valid in the second graph and "0" is assigned when it is invalid.
[0124] In FIG. 9, "Flag (second information)" indicates whether the edge is determined to correspond to a shortcut edge by the shortcut edge deletion process. For example, "Flag (second information)" stores "1" when it is determined by the shortcut edge deletion process that the edge does not correspond to a shortcut edge, that is, when the edge is not deleted in the second graph and exists in the second graph. Also, "Flag (second information)" stores "0" when it is determined by the shortcut edge deletion process that the edge corresponds to a shortcut edge, that is, when the edge is deleted in the second graph and does not exist in the second graph.
[0125] Also, in the "Flag (Second Information)" of FIG. 9, in order to indicate the state after the generation of the second information, it shows a state in which either "1" indicating valid or "0" indicating invalid is assigned to all edges. However, information indicating unprocessed may be associated with the edges before they become the target edges of the shortcut edge deletion process. For example, in the "Flag (Second Information)", for the edges before they become the target edges of the shortcut edge deletion process, that is, the unprocessed (unselected) edges, a flag indicating before processing (for example, any numerical value other than 0 and 1, Null, etc.) may be assigned. Thereby, the information processing apparatus 100 can generate information for specifying the structure of the second graph based on the second information, and can also specify the processing status of the generation process of the second information such as which edges have been processed.
[0126] In the example of FIG. 9, the node (node N1) identified by the node ID "N1" indicates that it corresponds to the object (target) identified by the object ID "OB1". Also, from node N1, it is shown that the edge (edge E1) identified by the edge ID "E1" is connected to the node (node N2) identified by the node ID "N2". That is, in the example of FIG. 9, it shows that in the first graph, it is possible to reach node N2 from node N1 via edge E1. Also, the flag of edge E1 is "1", indicating that edge E1 is not deleted as a shortcut edge and exists in the second graph. That is, in the example of FIG. 9, it shows that in the second graph, it is also possible to reach node N2 from node N1 via edge E1.
[0127] Also, from node N1, it is shown that an edge (edge E3) identified by edge ID "E3" is connected to a node (node N3) identified by node ID "N3". That is, in the example of FIG. 9, it is shown that from node N1 in the first graph, node N3 can be reached via edge E3. Also, the flag of edge E3 is "0", indicating that edge E3 is deleted as a shortcut edge and does not exist in the second graph. That is, in the example of FIG. 9, it is shown that in the second graph, node N3 cannot be directly reached from node N1 via edge E3.
[0128] Note that the graph information storage unit 124 is not limited to the above, and may store various information according to the purpose. For example, the graph information storage unit 124 may store the length of the edge connecting between each node (vector). That is, the graph information storage unit 124 may store information indicating the distance between each node (vector). Also, for example, the graph information storage unit 124 may store information indicating the number of input edges to each node.
[0129] Also, the graph data may include a program module that takes a query as input, searches for nodes by traversing the edges in the graph data, extracts and outputs nodes similar to the query. That is, the graph data may be assumed to be used as a program module for performing search processing using the graph. For example, the graph data may be a program that extracts and outputs, from the graph, nodes corresponding to vector data similar to the input vector data when vector data is input as a query. For example, the graph data may be data used as a program module for searching for similar images corresponding to a query image. For example, the graph data causes a computer to function so as to extract and output nodes similar to the input query in the graph based on the input query.
[0130] (Control Unit 130) Returning to the description of FIG. 5, the control unit 130 is a controller, which is realized, for example, by a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a GPU (Graphics Processing Unit), etc., when various programs (corresponding to an example of an information processing program) stored in a storage device inside the information processing apparatus 100 are executed with the RAM as a working area. Further, the control unit 130 is a controller, which is realized, for example, by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).
[0131] As shown in FIG. 5, the control unit 130 includes an acquisition unit 131, a search unit 132, a generation unit 133, and a provision unit 134, and realizes or executes the functions and operations of information processing described below. Note that the internal configuration of the control unit 130 is not limited to the configuration shown in FIG. 5, and other configurations may be used as long as they can perform the information processing described later.
[0132] (Acquisition Unit 131) The acquisition unit 131 acquires various information. For example, the acquisition unit 131 acquires various information from the storage unit 120. For example, the acquisition unit 131 acquires various information from the object information storage unit 121, the condition information storage unit 122, the reference information storage unit 123, the graph information storage unit 124, etc. Further, the acquisition unit 131 acquires various information from an external information processing apparatus.
[0133] The acquisition unit 131 acquires first graph information indicating a first graph in which a plurality of nodes corresponding to each of a plurality of objects to be searched are connected by edges. The acquisition unit 131 acquires condition information indicating a determination condition for shortcut edges to be deleted in the graph.
[0134] The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges based on the positional relationships of each of the first node, the second node, and the third node. The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges based on a first distance between the first node and the second node, a second distance between the first node and the third node, and a third distance between the second node and the third node.
[0135] The acquisition unit 131 acquires condition information that is the determination conditions of shortcut edges based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the comparison between the combination of the second distance and the third distance and the first distance. The acquisition unit 131 acquires condition information that sets, as the determination conditions of shortcut edges, that the second distance is smaller than the first distance, the third distance is smaller than the first distance, and the total distance obtained by adding the second distance and the third distance is smaller than the value obtained by multiplying the first distance by a predetermined coefficient.
[0136] The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges using the first distance, the second distance, and the third distance and a theorem related to triangles. The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges using the first distance, the second distance, and the third distance and a function related to triangles. The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges based on the comparison between a calculated value and a threshold value.
[0137] The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges based on the first distance, the second distance, and the third distance and the cosine theorem. The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges based on the first distance, the second distance, and the third distance and the cosine function. The acquisition unit 131 acquires condition information indicating the determination conditions of shortcut edges based on a threshold value corresponding to the value of cosine.
[0138] The acquisition unit 131 acquires condition information that serves as a determination condition for shortcut edges based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the first distance, the second distance, and the third distance and a function related to a triangle. The acquisition unit 131 acquires condition information that sets, as the determination condition for shortcut edges, that the second distance is smaller than the first distance, the third distance is smaller than the first distance, and a calculated value calculated based on the first distance, the second distance, and the third distance and the cosine function is larger than a threshold value.
[0139] The acquisition unit 131 acquires condition information that serves as a determination condition for shortcut edges based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the first distance, the second distance, and the third distance and a theorem related to a triangle. The acquisition unit 131 acquires condition information that sets, as the determination condition for shortcut edges, that the second distance is smaller than the first distance, the third distance is smaller than the first distance, and a calculated value calculated based on the first distance, the second distance, and the third distance and the cosine theorem is larger than a threshold value.
[0140] The acquisition unit 131 acquires the first graph information from the graph information storage unit 124. The acquisition unit 131 acquires the first graph that is a k-nearest neighbor graph. The acquisition unit 131 acquires the first graph that is an approximate k-nearest neighbor graph. For example, the information processing apparatus 100 acquires the first graph GR1 in FIG. 1. For example, the information processing apparatus 100 may acquire the first graph information such as the first graph GR1 from an external device such as the information providing apparatus 50.
[0141] For example, the acquisition unit 131 acquires information related to a search query. For example, the acquisition unit 131 acquires a search query related to image search. For example, the acquisition unit 131 acquires a query from the terminal device 10 to be used. For example, the acquisition unit 131 acquires a query from the information providing apparatus 50 that has received the query from the terminal device 10 to be used.
[0142] (Search unit 132) The search unit 132 searches for various information. The search unit 132 executes a search process using a graph. The search unit 132 functions as an extraction unit that extracts various information. The search unit 132 extracts various kinds of information. For example, the search unit 132 functions as a search unit that provides a search service regarding an object. The search unit 132 explores various information. The search unit 132 searches for various information. For example, the search unit 132 searches for an object by exploring graph data.
[0143] The search unit 132 executes a search process for searching a graph in response to an instruction from the generation unit 133. For example, when information indicating an object (node) to be processed is given, the search unit 132 extracts an object (node) similar to the target object (node) by exploring the graph based on the processing procedure shown in FIG. 10. The search unit 132 extracts neighboring objects (neighboring nodes) of a target object (target node) by performing a search process for searching a graph with one object (node) among a plurality of objects (nodes) as the target object (target node).
[0144] For example, the search unit 132 extracts various information from the object information storage unit 121, the condition information storage unit 122, the reference information storage unit 123, the graph information storage unit 124, and the like. For example, the search unit 132 acquires the starting point information GINF11 from the storage unit 120. For example, the search unit 132 extracts various information based on the information acquired by the acquisition unit 131.
[0145] The search unit 132 extracts a predetermined number (for example, the number of searches, etc.) of nodes from a plurality of nodes as neighboring nodes. The search unit 132 performs a search process for extracting neighboring nodes by exploring the graph. The search unit 132 performs a search process for extracting a predetermined number of nodes as neighboring nodes based on the relationship with an additional node among a plurality of nodes. The search unit 132 performs a search process for extracting a predetermined number of nodes as neighboring nodes based on the distance between each of the plurality of nodes and the additional node.
[0146] For example, when the query acquired by the acquisition unit 131 is acquired, the search unit 132 searches for an object similar to the query by searching the graph data. For example, the search unit 132 extracts an object similar to the query by searching the graph data. For example, the search unit 132 extracts an object similar to the query by searching the graph data based on the processing procedure as shown in FIG. 10.
[0147] The search unit 132 extracts neighboring nodes by searching the graph. The search unit 132 extracts a predetermined number of neighboring nodes by searching the graph with the additional node as a query. The search unit 132 extracts neighboring nodes by searching the graph by the search process as shown in FIG. 10.
[0148] For example, the search unit 132 executes a search process using the generated graph GR11. For example, the search unit 132 executes a search process using the generated graph GR1.
[0149] (Generation unit 133) The generation unit 133 executes various processes related to graph generation. The generation unit 133 executes a generation process for generating a graph. The generation unit 133 executes a process related to deleting shortcuts. The generation unit 133 generates information indicating a graph with shortcuts deleted. The generation unit 133 executes a selection process for selecting a node (object) to be processed. The generation unit 133 causes the search unit 132 to execute a search process by instructing the search unit 132, and acquires a search result from the search unit 132.
[0150] The generation unit 133 generates various information. For example, the generation unit 133 generates various information (data) from the information (data) stored in the storage unit 120. For example, the generation unit 133 generates various information from the object information storage unit 121, the condition information storage unit 122, the reference information storage unit 123, the graph information storage unit 124, and the like.
[0151] For example, the generation unit 133 generates various information based on the information acquired by the acquisition unit 131. The generation unit 133 generates various information using the result of the search process by the search unit 132.
[0152] When the relationships regarding the first directed edge that starts from the first node and ends at the second node, the second directed edge that starts from the first node and ends at a third node different from the second node, and the third directed edge that starts from the third node and ends at the second node among the edges included in the first graph satisfy the determination condition for a shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted assuming it is a shortcut edge. When the positional relationships of each of the first node, the second node, and the third node satisfy the determination condition for a shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted assuming it is a shortcut edge.
[0153] The generation unit 133 executes a shortcut edge deletion process targeting the target edge based on the target edge, which is the first directed edge, the second directed edge that starts from the first node that is the start point of the first directed edge and ends at a third node different from the second node that is the end point of the first directed edge, and the third directed edge that starts from the third node and ends at the second node. When the relationships of the first distance, the second distance, and the third distance satisfy the determination condition for a shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted assuming it is a shortcut edge.
[0154] When the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the comparison between the combination of the second distance and the third distance and the first distance satisfy the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. When the second distance is smaller than the first distance, the third distance is smaller than the first distance, and the total distance obtained by adding the second distance and the third distance is smaller than the value obtained by multiplying the first distance by a predetermined coefficient, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge.
[0155] When the calculated value calculated based on the first distance, the second distance, the third distance, and the theorem satisfies the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. When the calculated value calculated based on the first distance, the second distance, the third distance, and the function satisfies the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. When the calculated value is larger than the threshold value, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge.
[0156] When the calculated value calculated based on the first distance, the second distance, the third distance, and the law of cosines satisfies the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. The generation unit 133 calculates a calculated value that is the value of the cosine of the angle formed by the second directed edge corresponding to the second distance and the third directed edge corresponding to the third distance using the law of cosines, and when the calculated value satisfies the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge.
[0157] When the calculated value calculated based on the first distance, the second distance, the third distance, and the cosine function satisfies the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. The generation unit 133 calculates a calculated value that is the value of the cosine of the angle formed by the second directed edge corresponding to the second distance and the third directed edge corresponding to the third distance using the cosine function. When the calculated value satisfies the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. When the calculated value is greater than the threshold value, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge.
[0158] When the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the calculated value calculated based on the first distance, the second distance, the third distance, and the theorem satisfy the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. When the second distance is smaller than the first distance, the third distance is smaller than the first distance, and the calculated value calculated based on the first distance, the second distance, and the third distance and the cosine theorem is greater than the threshold value, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge.
[0159] When the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the calculated value calculated based on the first distance, the second distance, the third distance, and the function satisfy the determination condition of the shortcut edge, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge. When the second distance is smaller than the first distance, the third distance is smaller than the first distance, and the calculated value calculated based on the first distance, the second distance, and the third distance and the cosine function is greater than the threshold value, the generation unit 133 generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge.
[0160] The generation unit 133 selects, as a target edge, one unselected directed edge to be processed among the directed edges starting from each of the plurality of nodes included in the first graph, and executes a shortcut edge deletion process in which the selected target edges for each of the plurality of nodes are set as targets for deletion determination as shortcut edges, thereby generating second information indicating a second graph from which directed edges have been deleted from the first graph. The generation unit 133 generates the second information by repeatedly executing the shortcut edge deletion process with the directed edge selected as the processing target in the shortcut edge deletion process being set as a selected directed edge and selecting target edges for each of the plurality of nodes.
[0161] The generation unit 133 generates the second information by repeatedly executing the shortcut edge deletion process until there are no unselected directed edges to be processed for each of the plurality of nodes. The generation unit 133 selects target edges for each of the plurality of nodes based on a selection criterion for selecting as the target edge and executes the shortcut edge deletion process.
[0162] The generation unit 133 starts from each of the nodes included in the first graph, selects the shortest directed edge among the unselected directed edges to be processed, and executes the shortcut edge deletion process. The generation unit 133 starts from each of the nodes included in the first graph, selects the longest directed edge among the unselected directed edges to be processed, and executes the shortcut edge deletion process.
[0163] The generation unit 133 starts from each of the nodes included in the first graph, randomly selects a directed edge from the unselected directed edges to be processed, and executes the shortcut edge deletion process. The generation unit 133 executes the shortcut edge deletion process based on the determination condition of the shortcut edge indicated by the condition information.
[0164] The generation unit 133 generates, as second information, information for identifying an edge corresponding to a shortcut edge among the edges included in the first graph. The generation unit 133 generates, as second information, a flag associated with each of the edges included in the first graph and indicating whether or not it is a shortcut edge. The generation unit 133 generates, as second information, graph information in which the shortcut edges included in the first graph are deleted.
[0165] (Provision unit 134) The provision unit 134 provides various information. For example, the provision unit 134 transmits various information to the terminal device 10 and the information provision device 50. For example, the provision unit 134 provides an object ID corresponding to a query as a search result. For example, the provision unit 134 provides the object ID retrieved by the search unit 132 to the information provision device 50. For example, the provision unit 134 provides the object ID extracted by the search unit 132 by search to the information provision device 50. The provision unit 134 provides the object ID extracted by the search unit 132 to the information provision device 50 as information indicating a vector corresponding to the query.
[0166] Further, the provision unit 134 may provide the graph generated by the generation unit 133 to an external information processing device. For example, the provision unit 134 may transmit the graph GR11 generated by the generation unit 133 to the information provision device 50.
[0167] [4. Information processing flow] Next, with reference to FIGS. 10 and 11, the information processing procedure by the information processing system 1 according to the embodiment will be described. FIGS. 10 and 11 are flowcharts showing an example of information processing according to the embodiment.
[0168] First, FIG. 10 will be described. As shown in FIG. 10, the information processing apparatus 100 acquires first graph information indicating a first graph in which a plurality of nodes corresponding to a plurality of objects to be searched are connected by edges (step S101). For example, the information processing apparatus 100 acquires the first graph from the information provision device 50.
[0169] Then, the information processing apparatus 100 acquires condition information indicating the determination condition of the shortcut edge to be deleted in the graph (step S102). For example, the information processing apparatus 100 acquires the condition information from the information providing apparatus 50.
[0170] Then, among the edges included in the first graph, when the relationship regarding the first directed edge starting from the first node and ending at the second node, the second directed edge starting from the first node and ending at a third node different from the second node, and the third directed edge starting from the third node and ending at the second node satisfies the determination condition of the shortcut edge, the information processing apparatus 100 generates second information indicating the second graph in which the first directed edge is deleted as the shortcut edge (step S103). For example, the information processing apparatus 100 generates a flag indicating whether each edge in the first graph is to be deleted as the second information.
[0171] First, FIG. 11 will be described. As shown in FIG. 11, the information processing apparatus 100 acquires first graph information indicating a first graph in which a plurality of nodes corresponding to each of a plurality of objects to be searched are connected by edges (step S201). For example, the information processing apparatus 100 acquires the first graph from the information providing apparatus 50.
[0172] Then, among the directed edges starting from each of the plurality of nodes included in the first graph, the information processing apparatus 100 selects one unselected directed edge to be processed as the target edge, and executes a shortcut edge deletion process in which the selected target edge for each of the plurality of nodes is set as the target for deletion determination as the shortcut edge, thereby generating second information indicating the second graph in which the directed edge is deleted from the first graph (step S102). For example, the information processing apparatus 100 generates a flag indicating whether each edge in the first graph is to be deleted as the second information.
[0173] 〔5. Search example〕 Here, an example of a search using the above-described graph data is shown. Note that the search using the generated graph data is not limited to the following, and may be performed by various procedures. This point will be described using FIG. 13 as an example. FIG. 13 is a flowchart showing an example of a search process using graph data. The search process described below is performed by the search unit 132 of the information processing apparatus 100. Also, the object referred to below may be read as a node. For example, the information processing apparatus 100 (for example, the search unit 132) performs a search process. The search query of the process described below may be a target node, an object specified by the user, or the like.
[0174] Here, the neighborhood object set N(G, y) is a set of neighboring objects associated by the edges assigned to the node y. "G" may be predetermined graph data (for example, the first graph GR1 or the like). For example, the information processing apparatus 100 executes a k-nearest neighbor search process.
[0175] For example, the information processing apparatus 100 sets the radius r of the hypersphere to ∞ (infinity) (step S300), and extracts a subset S from the existing object set (step S301). For example, the information processing apparatus 100 may extract the object (node) selected as the root node as the subset S. Also, for example, a hypersphere is a virtual sphere indicating a search range. Note that the objects included in the object set S extracted in step S301 are also included in the initial set of the object set R of the search results at the same time.
[0176] Next, the information processing apparatus 100 extracts the object with the shortest distance from the object y when the search query object is y among the objects included in the object set S, and sets it as the object s (step S302). For example, when only the object (node) selected as the root node is an element of S, the root node is ultimately extracted as the object s by the information processing apparatus 100. Next, the information processing apparatus 100 excludes the object s from the object set S (step S303).
[0177] Next, the information processing apparatus 100 determines whether the distance d(s, y) between the object s and the object y exceeds r(1 + ε) (step S304). Here, ε is an expansion factor, and r(1 + ε) is a value indicating the radius of the search range (only the nodes within this range are searched. The accuracy can be improved by making it larger than the search range). When the distance d(s, y) between the object s and the object y exceeds r(1 + ε) (step S304: Yes), the information processing apparatus 100 outputs the object set R as the set of neighboring objects of the object y (step S305) and ends the process.
[0178] When the distance d(s, y) between the object s and the search query object y does not exceed r(1 + ε) (step S304: No), the information processing apparatus 100 selects one object that is not included in the object set C from among the objects that are elements of the set N(G, s) of neighboring objects of the object s, and stores the selected object u in the object set C (step S306). The object set C is provided for the sake of convenience to avoid duplicate searches and is set to an empty set at the start of the process.
[0179] Next, the information processing apparatus 100 determines whether the distance d(u, y) between the object u and the object y is less than or equal to r(1 + ε) (step S307). When the distance d(u, y) between the object u and the object y is less than or equal to r(1 + ε) (step S307: Yes), the information processing apparatus 100 adds the object u to the object set S (step S308). Also, when the distance d(u, y) between the object u and the object y is not less than or equal to r(1 + ε) (step S307: No), the information processing apparatus 100 performs the determination (process) of step S309.
[0180] Next, the information processing apparatus 100 determines whether the distance d(u, y) between the object u and the object y is less than or equal to r (step S309). If the distance d(u, y) between the object u and the object y exceeds r, the information processing apparatus 100 performs the determination (process) in step S315. Also, if the distance d(u, y) between the object u and the object y is not less than or equal to r (step S309: No), the information processing apparatus 100 performs the determination (process) in step S315.
[0181] If the distance d(u, y) between the object u and the object y is less than or equal to r (step S309: Yes), the information processing apparatus 100 adds the object u to the object set R (step S310). Then, the information processing apparatus 100 determines whether the number of objects included in the object set R exceeds ks (step S311). The predetermined number ks is a natural number arbitrarily determined. For example, ks may be the number of candidates. For example, ks may be any value such as ks = 2. If the number of objects included in the object set R does not exceed ks (step S311: No), the information processing apparatus 100 performs the determination (process) in step S313.
[0182] If the number of objects included in the object set R exceeds ks (step S311: Yes), the information processing apparatus 100 excludes from the object set R the object having the longest (farthest) distance from the object y among the objects included in the object set R (step S312).
[0183] Next, the information processing apparatus 100 determines whether the number of objects included in the object set R matches ks (step S313). If the number of objects included in the object set R does not match ks (step S313: No), the information processing apparatus 100 performs the determination (process) in step S315. If the number of objects included in the object set R matches ks (step S313: Yes), the information processing apparatus 100 sets the distance between the object having the longest (farthest) distance from object y among the objects included in the object set R and object y as the new r (step S314).
[0184] Then, the information processing apparatus 100 determines whether it has finished selecting all objects from the objects that are elements of the neighborhood object set N(G, s) of object s and storing them in the object set C (step S315). If it has not finished selecting all objects from the objects that are elements of the neighborhood object set N(G, s) of object s and storing them in the object set C (step S315: No), the information processing apparatus 100 returns to step S306 to repeat the process.
[0185] When all the objects are selected from the objects that are elements of the set N(G, s) of neighboring objects of the object s and stored in the object set C (step S315: Yes), the information processing apparatus 100 determines whether the object set S is an empty set (step S316). When the object set S is not an empty set (step S316: No), the information processing apparatus 100 returns to step S302 and repeats the process. When the object set S is an empty set (step S316: Yes), the information processing apparatus 100 outputs the object set R and ends the process (step S317). For example, the information processing apparatus 100 extracts the objects (nodes) of the candidate number included in the object set R as neighboring nodes corresponding to the target node (input object y). For example, the information processing apparatus 100 extracts the objects (nodes) included in the object set R as a node group corresponding to the target node (input object y). Also, for example, the information processing apparatus 100 may provide the objects (nodes) included in the object set R to the terminal device or the like that has performed the search as a search result corresponding to the search query (input object y).
[0186] [6. Effects] As described above, the information processing apparatus according to the embodiment (corresponding to "information processing apparatus 100" in the embodiment) includes an acquisition unit (corresponding to "acquisition unit 131" in the embodiment) and a generation unit (corresponding to "generation unit 133" in the embodiment). The acquisition unit acquires first graph information indicating a first graph in which a plurality of nodes corresponding to each of a plurality of objects to be searched are connected by edges. The generation unit selects, as a target edge, one unselected directed edge to be processed among the directed edges starting from each of the plurality of nodes included in the first graph, and executes a shortcut edge deletion process in which the selected target edge for each of the plurality of nodes is a target for deletion determination as a shortcut edge, thereby generating second information indicating a second graph from which the directed edges have been deleted from the first graph.
[0187] In this way, the information processing apparatus 100 according to the embodiment selects, as a target edge, one unselected directed edge among the directed edges starting from each of the plurality of nodes included in the first graph, executes the shortcut edge deletion process, and generates second information indicating a second graph from which the directed edge has been deleted from the first graph, thereby being able to appropriately generate information indicating another graph in which the edge in the graph has been deleted.
[0188] Also, in the information processing apparatus 100 according to the embodiment, the generation unit generates the second information by repeatedly selecting, for each of the plurality of nodes, the target edge as a selected directed edge and executing the shortcut edge deletion process on the selected directed edge.
[0189] Thereby, the information processing apparatus 100 according to the embodiment generates the second information by repeatedly selecting, for each of the plurality of nodes, the target edge as a selected directed edge and executing the shortcut edge deletion process on the selected directed edge, thereby being able to appropriately generate information indicating another graph in which the edge in the graph has been deleted.
[0190] Also, in the information processing apparatus 100 according to the embodiment, the generation unit generates the second information by repeatedly executing the shortcut edge deletion process until there are no unselected directed edges as processing targets for each of the plurality of nodes.
[0191] Thereby, the information processing apparatus 100 according to the embodiment generates the second information by repeatedly executing the shortcut edge deletion process until there are no unselected directed edges as processing targets for each of the plurality of nodes, thereby being able to appropriately generate information indicating another graph in which the edge in the graph has been deleted.
[0192] Further, in the information processing apparatus 100 according to the embodiment, the generation unit selects a target edge for each of the plurality of nodes based on a selection criterion for selecting the target edge, and executes shortcut edge deletion processing.
[0193] Thereby, the information processing apparatus 100 according to the embodiment can appropriately generate information indicating another graph in which an edge in the graph is deleted by selecting a target edge for each of the plurality of nodes based on a selection criterion for selecting the target edge and executing shortcut edge deletion processing.
[0194] Further, in the information processing apparatus 100 according to the embodiment, the generation unit selects, as a starting point, each of the nodes included in the first graph, and among the unselected directed edges to be processed, selects the shortest directed edge and executes shortcut edge deletion processing.
[0195] Thereby, the information processing apparatus 100 according to the embodiment can appropriately generate information indicating another graph in which an edge in the graph is deleted by selecting, as a starting point, each of the nodes included in the first graph, and among the unselected directed edges to be processed, selecting the shortest directed edge and executing shortcut edge deletion processing.
[0196] Further, in the information processing apparatus 100 according to the embodiment, the generation unit selects, as a starting point, each of the nodes included in the first graph, and among the unselected directed edges to be processed, selects the longest directed edge and executes shortcut edge deletion processing.
[0197] Thereby, the information processing apparatus 100 according to the embodiment can appropriately generate information indicating another graph in which an edge in the graph is deleted by selecting, as a starting point, each of the nodes included in the first graph, and among the unselected directed edges to be processed, selecting the longest directed edge and executing shortcut edge deletion processing.
[0198] Further, in the information processing apparatus 100 according to the embodiment, the generation unit starts from each of the nodes included in the first graph, randomly selects a directed edge from the unselected directed edges to be processed, and executes the shortcut edge deletion process.
[0199] Thereby, the information processing apparatus 100 according to the embodiment starts from each of the nodes included in the first graph, randomly selects a directed edge from the unselected directed edges to be processed, and executes the shortcut edge deletion process, whereby information indicating another graph in which an edge in the graph has been deleted can be appropriately generated.
[0200] Further, in the information processing apparatus 100 according to the embodiment, the acquisition unit acquires condition information indicating a determination condition for a shortcut edge to be deleted in the graph. The generation unit executes the shortcut edge deletion process based on the determination condition for the shortcut edge indicated by the condition information.
[0201] Thereby, the information processing apparatus 100 according to the embodiment executes the shortcut edge deletion process based on the determination condition for the shortcut edge indicated by the condition information, whereby information indicating another graph in which an edge in the graph has been deleted can be appropriately generated.
[0202] Further, in the information processing apparatus 100 according to the embodiment, the generation unit executes the shortcut edge deletion process targeting the target edge based on the first directed edge that is the target edge, the first node that is the start point of the first directed edge as the start point, the second directed edge having a third node different from the second node that is the end point of the first directed edge as the end point, and the third directed edge having the third node as the start point and the second node as the end point.
[0203] As a result, the information processing apparatus 100 according to the embodiment determines a shortcut edge based on the relationships of three edges, i.e., a first directed edge, a second directed edge, and a third directed edge corresponding to three nodes, namely a first node, a second node, and a third node, and generates second information indicating a second graph in which the determined shortcut edge is deleted, thereby being able to appropriately generate information indicating another graph in which an edge in the graph is deleted.
[0204] Also, in the information processing apparatus 100 according to the embodiment, the acquisition unit acquires condition information indicating a determination condition of a shortcut edge based on the positional relationship of each of the first node, the second node, and the third node. When the positional relationship of each of the first node, the second node, and the third node satisfies the determination condition of the shortcut edge, the generation unit generates second information indicating a second graph in which the first directed edge is deleted assuming that it is a shortcut edge.
[0205] As a result, when the positional relationship of each of the first node, the second node, and the third node satisfies the determination condition of the shortcut edge, the information processing apparatus 100 according to the embodiment deletes the first directed edge assuming that it is a shortcut edge, thereby being able to appropriately generate information indicating another graph in which an edge in the graph is deleted.
[0206] Also, in the information processing apparatus 100 according to the embodiment, the acquisition unit acquires condition information indicating a determination condition of a shortcut edge based on a first distance between the first node and the second node, a second distance between the first node and the third node, and a third distance between the second node and the third node. When the relationship of the first distance, the second distance, and the third distance satisfies the determination condition of the shortcut edge, the generation unit generates second information indicating a second graph in which the first directed edge is deleted assuming that it is a shortcut edge.
[0207] As a result, when the relationship among the first distance, the second distance, and the third distance satisfies the determination condition of the shortcut edge, the information processing apparatus 100 according to the embodiment can appropriately generate information indicating another graph in which an edge in the graph is deleted by deleting the first directed edge as the shortcut edge.
[0208] Further, in the information processing apparatus 100 according to the embodiment, the acquisition unit acquires condition information that is a determination condition of the shortcut edge based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the comparison between the combination of the second distance and the third distance and the first distance. The generation unit generates second information indicating a second graph in which the first directed edge is deleted as the shortcut edge when the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the comparison between the combination of the second distance and the third distance and the first distance satisfy the determination condition of the shortcut edge.
[0209] As a result, when the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the comparison between the combination of the second distance and the third distance and the first distance satisfy the determination condition of the shortcut edge, the information processing apparatus 100 according to the embodiment can appropriately generate information indicating another graph in which an edge in the graph is deleted by deleting the first directed edge as the shortcut edge.
[0210] Further, in the information processing apparatus 100 according to the embodiment, the acquisition unit acquires condition information indicating a determination condition of the shortcut edge using the first distance, the second distance, and the third distance and a function related to a triangle. The generation unit generates second information indicating a second graph in which the first directed edge is deleted as the shortcut edge when the calculated value calculated based on the first distance, the second distance, and the third distance and the function satisfies the determination condition of the shortcut edge.
[0211] As a result, when the calculated value calculated based on the first distance, the second distance, the third distance, and the function satisfies the determination condition for the shortcut edge, the information processing apparatus 100 according to the embodiment deletes the first directed edge as the shortcut edge, thereby appropriately generating information indicating another graph in which the edge in the graph has been deleted.
[0212] Also, in the information processing apparatus 100 according to the embodiment, the acquisition unit acquires condition information that is a determination condition for the shortcut edge based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the first distance, the second distance, and the third distance, and a function related to a triangle. The generation unit generates second information indicating a second graph in which the first directed edge is deleted as a shortcut edge when the calculated value calculated based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the first distance, the second distance, and the third distance, and the function satisfies the determination condition for the shortcut edge.
[0213] As a result, when the calculated value calculated based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the first distance, the second distance, and the third distance, and the function satisfies the determination condition for the shortcut edge, the information processing apparatus 100 according to the embodiment deletes the first directed edge as the shortcut edge, thereby appropriately generating information indicating another graph in which the edge in the graph has been deleted.
[0214] Also, in the information processing apparatus 100 according to the embodiment, the generation unit generates, as the second information, information for identifying an edge corresponding to the shortcut edge among the edges included in the first graph.
[0215] As a result, the information processing apparatus 100 according to the embodiment appropriately generates information indicating another graph in which the edge in the graph has been deleted by generating, as the second information, information for identifying an edge corresponding to the shortcut edge among the edges included in the first graph.
[0216] Further, in the information processing apparatus 100 according to the embodiment, the generation unit generates, as second information, a flag associated with each edge included in the first graph and indicating whether the edge is a shortcut edge.
[0217] Thereby, the information processing apparatus 100 according to the embodiment can appropriately generate information indicating another graph in which an edge in the graph has been deleted by generating, as second information, a flag associated with each edge included in the first graph and indicating whether the edge is a shortcut edge.
[0218] Further, in the information processing apparatus 100 according to the embodiment, the generation unit generates, as second information, graph information in which a shortcut edge included in the first graph has been deleted.
[0219] Thereby, the information processing apparatus 100 according to the embodiment can appropriately generate information indicating another graph in which an edge in the graph has been deleted by generating, as second information, graph information in which a shortcut edge included in the first graph has been deleted.
[0220] 〔7. Hardware Configuration〕 The information processing apparatus 100 according to the above-described embodiment is realized by a computer 1000 having a configuration as shown in FIG. 14, for example. FIG. 14 is a hardware configuration diagram showing an example of a computer that realizes the functions of the information processing apparatus. The computer 1000 includes a CPU 1100, a RAM 1200, a ROM (Read Only Memory) 1300, an HDD (Hard Disk Drive) 1400, a communication interface (I / F) 1500, an input / output interface (I / F) 1600, and a media interface (I / F) 1700.
[0221] The CPU 1100 operates based on programs stored in the ROM 1300 or the HDD 1400 and controls each part. The ROM 1300 stores a boot program executed by the CPU 1100 when the computer 1000 starts up, programs dependent on the hardware of the computer 1000, and the like.
[0222] The HDD 1400 stores programs executed by the CPU 1100, data used by such programs, and the like. The communication interface 1500 receives data from other devices via the network N and sends it to the CPU 1100, and sends data generated by the CPU 1100 to other devices via the network N.
[0223] The CPU 1100 controls output devices such as displays and printers, and input devices such as keyboards and mice via the input / output interface 1600. The CPU 1100 acquires data from the input devices via the input / output interface 1600. Also, the CPU 1100 outputs the generated data to the output devices via the input / output interface 1600.
[0224] The media interface 1700 reads a program or data stored in the recording medium 1800 and provides it to the CPU 1100 via the RAM 1200. The CPU 1100 loads such a program from the recording medium 1800 onto the RAM 1200 via the media interface 1700 and executes the loaded program. The recording medium 1800 is, for example, an optical recording medium such as a DVD (Digital Versatile Disc), a PD (Phase change rewritable Disk), a magneto-optical recording medium such as an MO (Magneto-Optical disk), a tape medium, a magnetic recording medium, or a semiconductor memory.
[0225] For example, when the computer 1000 functions as the information processing apparatus 100 according to the embodiment, the CPU 1100 of the computer 1000 realizes the functions of the control unit 130 by executing the program loaded on the RAM 1200. The CPU 1100 of the computer 1000 reads and executes these programs from the recording medium 1800. As another example, these programs may be acquired from other devices via the network N.
[0226] As described above, some of the embodiments of the present application have been described in detail with reference to the drawings. However, these are merely examples, and the present invention can be implemented in other forms with various modifications and improvements based on the knowledge of those skilled in the art, starting from the aspects described in the disclosure of the invention.
[0227] [8. Others] In addition, among the processes described in the above embodiments, all or part of the processes described as being automatically performed can be manually performed, or all or part of the processes described as being manually performed can be automatically performed by a known method. In addition, the processing procedures, specific names, and information including various data and parameters shown in the above documents and drawings can be arbitrarily changed unless otherwise specified. For example, the various information shown in each figure is not limited to the illustrated information.
[0228] In addition, each component of each device shown in the drawings is a functional concept, and it is not necessarily physically configured as shown in the drawings. That is, the specific form of the distribution and integration of each device is not limited to that shown in the drawings, and all or part of it can be functionally or physically distributed and integrated in any unit according to various loads and usage situations.
[0229] In addition, the processes described in each of the above-described embodiments can be appropriately combined as long as the processing contents do not conflict.
[0230] In addition, the "section (section, module, unit)" described above can be read as "means", "circuit", etc. For example, the acquisition section can be read as an acquisition means or an acquisition circuit.
Explanation of symbols
[0231] 1 Information processing system 100 Information processing device 121 Object information storage unit 122 Condition information storage unit 123 Reference information storage unit 124 Graph information storage unit 130 Control unit 131 Acquisition unit 132 Search unit 133 Generation unit 134 Provision unit 10 Terminal device 50 Information providing device N Network
Claims
1. An acquisition unit that acquires first graph information indicating a first graph in which a plurality of nodes corresponding to each of a plurality of objects to be searched are connected by edges; A generation unit that generates second information indicating a second graph from which directed edges have been deleted from the first graph by executing a shortcut edge deletion process of selecting, as a target edge, one directed edge that has not been selected as a processing target among the directed edges starting from each of the plurality of nodes included in the first graph, and setting the selected target edge as a target for deletion determination as a shortcut edge; An information processing apparatus comprising the above.
2. The generation unit: By repeatedly executing the shortcut edge deletion process by selecting the target edge for each of the plurality of nodes with the directed edge selected as the processing target as a selected directed edge, the second information is generated. The information processing apparatus according to claim 1, characterized in that.
3. The generation unit: By repeatedly executing the shortcut edge deletion process until there are no unselected directed edges as the processing target for each of the plurality of nodes, the second information is generated. The information processing apparatus according to claim 2, characterized in that.
4. The generation unit: Based on a selection criterion for selecting as the target edge, the target edge is selected for each of the plurality of nodes and the shortcut edge deletion process is executed. The information processing apparatus according to claim 1, characterized in that.
5. The generation unit: Starting from each of the nodes included in the first graph, among the unselected directed edges as the processing target, the shortest directed edge is selected and the shortcut edge deletion process is executed. The information processing apparatus according to claim 1, characterized in that.
6. The generation unit: Starting from each of the nodes included in the first graph, among the unselected directed edges as the processing target, the longest directed edge is selected and the shortcut edge deletion process is executed. The information processing apparatus according to claim 1, characterized in that.
7. The generation unit: Starting from each of the nodes included in the first graph, a directed edge is randomly selected from the unselected directed edges as the processing target and the shortcut edge deletion process is executed. The information processing apparatus according to claim 1, characterized in that...
8. The acquisition unit acquires condition information indicating a determination condition of the shortcut edge to be deleted in the graph, The generation unit executes the shortcut edge deletion process based on the determination condition of the shortcut edge indicated by the condition information. The information processing apparatus according to claim 1, characterized in that...
9. The generation unit executes the shortcut edge deletion process for the target edge based on the first directed edge that is the target edge, the second directed edge having a third node different from the second node that is the end point of the first directed edge and having the first node that is the start point of the first directed edge as the start point, and the third directed edge having the second node as the end point and the third node as the start point. The information processing apparatus according to claim 8, characterized in that...
10. The acquisition unit acquires the condition information indicating the determination condition of the shortcut edge based on the positional relationship of each of the first node, the second node, and the third node, The generation unit generates the second information indicating the second graph in which the first directed edge is deleted as the shortcut edge when the positional relationship of each of the first node, the second node, and the third node satisfies the determination condition of the shortcut edge. The information processing apparatus according to claim 9, characterized in that...
11. The acquisition unit acquires the condition information indicating the determination condition of the shortcut edge based on the first distance between the first node and the second node, the second distance between the first node and the third node, and the third distance between the second node and the third node, The generation unit generates the second information indicating the second graph in which the first directed edge is deleted as the shortcut edge when the relationship between the first distance, the second distance, and the third distance satisfies the determination condition of the shortcut edge. The information processing apparatus according to claim 9, characterized in that...
12. The acquisition unit acquires the condition information that is the determination condition of the shortcut edge based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the comparison between the combination of the second distance and the third distance and the first distance, The generation unit When the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the comparison between the combination of the second distance and the third distance and the first distance satisfy the determination condition of the shortcut edge, the second information indicating the second graph in which the first directed edge is deleted as the shortcut edge is generated. The information processing apparatus according to claim 11, characterized in that.
13. The acquisition unit acquires the condition information indicating the determination condition of the shortcut edge using the first distance, the second distance, the third distance, and a function related to a triangle, The generation unit When the calculated value calculated based on the first distance, the second distance, the third distance, and the function satisfies the determination condition of the shortcut edge, the second information indicating the second graph in which the first directed edge is deleted as the shortcut edge is generated. The information processing apparatus according to claim 11, characterized in that.
14. The acquisition unit acquires the condition information serving as the determination condition of the shortcut edge based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the first distance, the second distance, the third distance, and a function related to a triangle, The generation unit When the calculated value calculated based on the comparison between the second distance and the first distance, the comparison between the third distance and the first distance, and the first distance, the second distance, the third distance, and the function satisfies the determination condition of the shortcut edge, the second information indicating the second graph in which the first directed edge is deleted as the shortcut edge is generated. The information processing apparatus according to claim 11, characterized in that.
15. The generation unit generates, as the second information, information for identifying an edge corresponding to the shortcut edge among the edges included in the first graph. The information processing apparatus according to claim 1, characterized in that.
16. The generation unit generates, as the second information, a flag associated with each of the edges included in the first graph and indicating whether the edge is a shortcut edge. The information processing apparatus according to claim 1, characterized in that.
17. The generation unit generates, as the second information, graph information in which the shortcut edge included in the first graph is deleted. The information processing apparatus according to claim 1, characterized in that...
18. An information processing method executed by a computer, comprising: an acquisition step of acquiring first graph information indicating a first graph in which a plurality of nodes corresponding to each of a plurality of objects to be searched are connected by edges; a generation step of generating second information indicating a second graph from which directed edges have been deleted from the first graph by executing a shortcut edge deletion process of selecting, as a target edge, one unselected directed edge among the directed edges starting from each of the plurality of nodes included in the first graph, and setting the selected target edge for each of the plurality of nodes as a target for deletion determination as a shortcut edge; An information processing method, characterized by including the above.
19. an acquisition procedure of acquiring first graph information indicating a first graph in which a plurality of nodes corresponding to each of a plurality of objects to be searched are connected by edges; a generation procedure of generating second information indicating a second graph from which directed edges have been deleted from the first graph by executing a shortcut edge deletion process of selecting, as a target edge, one unselected directed edge among the directed edges starting from each of the plurality of nodes included in the first graph, and setting the selected target edge for each of the plurality of nodes as a target for deletion determination as a shortcut edge; An information processing program, characterized by causing a computer to execute the above.
Citation Information
Patent Citations
Conductor for connection of semiconductor device
JP1987093335A
Information processing device, information processing method, and information processing program
JP7080803B2
Information processing device, information processing method, and information processing program
JP7130019B2