Generation method, search method and generation device
By generating and storing the index information of the compressed vector in the directed graph, the problem of excessive memory usage in the prior art is solved, and efficient approximate nearest neighbor search is achieved.
Patent Information
- Application Number
- CN202411180305.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-29
- Filing Date
- 2024-08-27
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art requires a large amount of memory resources when performing approximate nearest neighbor searches, resulting in an increase in memory usage.
By generating the node information in the directed graph is allocated and stored in the index information, only the compression vector required for the jumping action is stored in the memory, reducing the amount of memory usage.
It realizes the reduction of memory usage when performing approximate nearest neighbor search, improves memory access efficiency, and reduces the demand for memory devices.
Smart Images

Figure CN120386782A_ABST
Abstract
Description
Technical Field
[0001] This embodiment relates to a generation method, a search method, and a generation device. Background Art
[0002] As one of the graph-based approximate nearest neighbor search algorithms, an algorithm such as DiskANN (Disk-based Approximate Nearest Neighbor search) is known. According to DiskANN, a directed graph is created by treating each multi-dimensional vector in the search range, that is, the multi-dimensional vector group, as a node, and the index information generated based on the structure of the directed graph is stored in a memory device. Then, based on the index information in the memory device, a search operation according to the directed graph is performed.
[0003] [Background Art Documents]
[0004] [Non-Patent Documents]
[0005] [Non-Patent Document 1] Suhas Jayaram Subramanya, Devvrit, Rohan Kadekodi, Ravishankar Krishaswamy, and Harsha Vardhan Simhadri, "DiskANN: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node", [online], November 2019, NeurIPS, [retrieved on December 11, 2022], retrieved from the following Internet: <URL: https: / / suhasjs.github.io / files / diskann_neurips19.pdf> Summary of the Invention
[0006] [Problems to be Solved by the Invention]
[0007] An object of one embodiment is to provide a method for generating index information that can perform a search operation with less memory usage, a search method using the index information, and a generation device.
[0008] [Technical Means for Solving the Problems]
[0009] According to one embodiment, a generation method is a method for generating a processor configured to process data represented by a directed graph. The generation method includes setting and writing. The setting is to set one of a plurality of first nodes corresponding to a plurality of first vectors included in a search range as a second node. The plurality of first nodes are a plurality of nodes to which each ID included in the directed graph is assigned. The writing is to write an information piece including a second vector, an ID associated with each of one or more third nodes, and a vector corresponding to the third node, that is, a third vector. The second vector is the first vector corresponding to the second node among the plurality of first vectors. The one or more third nodes are outer adjacent nodes of the second node among the plurality of first nodes. The information piece is an element related to the second node in index information corresponding to the directed graph. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 FIG. is a schematic diagram showing an example of the configuration of a search device according to an embodiment.
[0011] Figure 2 FIG. is a diagram for explaining the configuration of a directed graph and index information according to an embodiment.
[0012] Figure 3 FIG. is a schematic diagram showing an example of information stored in an SSD and a DRAM when a search device according to an embodiment performs a search.
[0013] Figure 4 FIG. is a schematic diagram showing an example of the configuration of a generation device according to an embodiment.
[0014] Figure 5 FIG. is a schematic diagram showing an example of functions implemented by a processor included in a generation device according to an embodiment.
[0015] Figure 6 FIG. is a flowchart showing an example of the operation of a generation device according to an embodiment.
[0016] Figure 7 FIG. is a schematic diagram showing an example of functions implemented by a processor included in a search device according to an embodiment.
[0017] Figure 8 FIG. is a flowchart showing an example of the operation of a search device according to an embodiment.
[0018] Figure 9 FIG. is a diagram showing an example of the configuration of node information according to Variation 1.
[0019] Figure 10 FIG. is a diagram showing an example of the configuration of node information according to Variation 2. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Hereinafter, with reference to the accompanying drawings, a method for generating an embodiment, a search method, and a generating apparatus will be described in detail. In addition, the present invention is not limited by the described embodiment.
[0021] (Embodiment)
[0022] First, an example of an apparatus (referred to as a search apparatus) that executes the search method of the embodiment will be described. Figure 1 It is a schematic diagram showing an example of the configuration of a search apparatus for an embodiment.
[0023] Figure 1 In the example shown, the search apparatus 2 includes a processor 21, an interface 22, an SSD (Solid State Drive) 23, a DRAM (Dynamic Random Access Memory) 24, and a bus 25. The processor 21, the interface 22, the SSD 23, and the DRAM 24 are electrically connected to the bus 25.
[0024] The interface 22 is a device for inputting and outputting information to and from the search apparatus 2. The interface 22 includes an interface for communication via a network, an interface to which a storage device can be connected, an interface to which an input device such as a keyboard can be connected, and the like. The search apparatus 2 can receive the input of a query via the interface.
[0025] The SSD 23 is a large-capacity non-volatile memory device that functions as a memory device of the search apparatus 2. The SSD 23 includes a NAND (Not and) type flash memory as a storage device. In addition, the memory device applicable to the search apparatus 2 is not limited to the SSD. The search apparatus 2 may also include a disk device as a memory device. An example of the disk device is an HDD (Hard Disk Drive).
[0026] The DRAM 24 is a memory that operates faster than the memory device. The DRAM 24 functions as a cache area, a buffer area, a working area, or the like. In addition, the memory applicable to the search apparatus 2 and operating faster than the memory device is not limited to the DRAM.
[0027] The processor 21 is an arithmetic unit that can execute a computer program and realizes the functions specified by the computer program. The processor 21 is, for example, a CPU (Central Processing Unit). In addition, according to the functions to be realized, one or more than one processor 21 may be provided. In the search device 2, the processor 21 executes a search operation according to DiskANN based on a search program (search program SPG described later). The search program SPG is stored, for example, in the SSD 23 or a device external to the search device 2. The processor 21 loads the search program SPG from the SSD 23 or a device external to the search device 2 into the DRAM 24 in the environment provided by the operating system. And the processor 21 executes the search program SPG loaded into the DRAM 24.
[0028] The search operation is an operation to find the data closest to the query in a specific data group. Each data has N (where N is an integer of 1 or more) elements. In other words, each data is an N-dimensional vector. Each data is an image, a file, or any other type of data, or data generated from these data. In one example, each data is N feature amounts extracted from an image. In the whole data and the query described later, the number of elements N is common. Hereinafter, the data is denoted as a vector or a full-precision vector. In addition, the group of data (that is, vectors) set as the search range is denoted as a vector set.
[0029] Regarding the vector set, a directed graph is pre-generated with each vector constituting the vector set regarded as a node. In the search operation, the processor 21 searches for the vector closest to the query according to the directed graph.
[0030] The technology compared with the embodiment is described. The technology compared with the embodiment is denoted as a comparative example. According to the comparative example, compressed vectors related to each vector included in the vector set are generated, and the compressed vectors of the whole vector part included in the vector set are stored in the DRAM. That is, the set of compressed vectors corresponding to the vector set is stored in the DRAM. The composition of the directed graph generated from the vector set is recorded in the index information. According to the directed graph specified by the index information, a search path is selected for each node. Each time a search path is selected, the required compressed vectors are selected from the compressed vectors of the whole vector part stored in the DRAM, and according to the result of distance calculation using the selected compressed vectors, the next search path is selected.
[0031] According to the above comparative example, a high-speed search operation can be performed. On the other hand, a DRAM with a very large capacity is required. And the usage amount of the DRAM increases according to the number of vectors in the search range.
[0032] In the embodiment, in order to reduce the usage amount of the DRAM, the compressed vectors required for distance calculation are recorded in the index information.
[0033] Figure 2This is a diagram for explaining the configuration of the directed graph GF and the index information IDX of the embodiments.
[0034] Node IDs (hereinafter abbreviated as NIDs) are assigned to each vector V included in the vector set. The method of assigning NIDs to each vector V is not limited to a specific method. Hereinafter, a vector V (in other words, a node) to which an NID x (where x is numerical information) is assigned will be denoted as node NIDx.
[0035] Figure 2 The directed graph GF generated from the vector set is shown. In the example, at node NID1, edges having node NID20 as the head, edges having node NID7 as the head, edges having node NID13 as the head, and edges having node NID12 as the head are connected. At node NID20, edges having node NID3 as the head, edges having node NID6 as the head, edges having node NID15 as the head, and edges having node NID11 as the head are connected. At node NID7, edges having node NID10 as the head, edges having node NID19 as the head, edges having node NID16 as the head, and edges having node NID5 as the head are connected. At node NID13, edges having node NID8 as the head, edges having node NID21 as the head, edges having node NID4 as the head, and edges having node NID2 as the head are connected. At node NID12, edges having node NID14 as the head, edges having node NID17 as the head, edges having node NID9 as the head, and edges having node NID18 as the head are connected.
[0036] When node NIDA and node NIDB are connected by an edge having node NIDB as the head, node NIDB is called an out-neighbor of node NIDA. The number of adjacent nodes of node NIDA is called the out-degree of node NIDA. In this specification, the out-neighbor is denoted as an adjacent node.
[0037] In addition, Figure 2 In the example shown, the directed graph GF has a tree-like shape. The shape of the directed graph GF is not limited to a tree-like shape. The out-degrees of multiple nodes may not be common. Multiple nodes may also be connected in a ring. Among many nodes, the out-degree is 1 or more, but there may also be a node with an out-degree of 0. In the embodiments, for the sake of simplicity of explanation, the case where the out-degree of all nodes is 1 or more and the out-degrees of all nodes are common will be described.
[0038] In the search operation, the processor 21 performs an operation to search for the node closest to the query based on an arbitrary search algorithm according to the directed graph GF specified by the index information IDX. As the operation algorithm for the search, any algorithm including Greedy search or Beam search can be adopted. If a simple example is described, the processor 21 sequentially switches the candidates for the node closest to the query, that is, the search target nodes, according to the directed graph GF among multiple nodes. Each time the search target node is switched, the processor 21 calculates the distance between each adjacent node of the search target node and the query. And the processor 21 sets the node closest to the query among one or more adjacent nodes of one or more search target nodes close to the query at the current time point as the next new search target node. The processor 21 sequentially performs the switching of the search target nodes according to the directed graph GF until reaching the vector V presumed to be closest to the query. Sometimes the process of switching the search target nodes according to the directed graph GF is denoted as a "jumping operation".
[0039] The distance in this specification is the length representing the similarity between data (including vectors and queries). Mathematically, the distance is, for example, the Euclidean distance. In addition, the mathematical definition of the distance is not limited to the Euclidean distance. Furthermore, the index for evaluating the distance is not limited to the Euclidean distance or the like, and any index can be used as long as it corresponds to the distance.
[0040] The processor 21 refers to the index information IDX corresponding to the directed graph GF in order to specify each adjacent node of the search target node. That is, the index information IDX has a data structure capable of deriving the structure of the directed graph GF.
[0041] The index information IDX is a set of node information. Each node information is an element of the index information IDX, that is, an information piece. Each node information corresponds one-to-one with one node. The index information IDX has a structure in which all node information is arranged in the order of NID.
[0042] All node information has a common composition. As a representative of all node information, the composition of the node information of a certain node NIDi is described.
[0043] The node information of the node NIDi includes the vector of the node NIDi (that is, the full-precision vector). In addition, the node information of the node NIDi includes the NID and the compressed vector related to all adjacent nodes of the node NIDi. The compressed vector is a vector generated from the full-precision vector assigned to the node NIDi. If the out-degree of each node is denoted as R, the node information of the node NIDi includes the NID and the compressed vector related to each of the R adjacent nodes.
[0044] The compressed vector is generated by compressing the full-precision vector. The compression algorithm is not limited to a specific algorithm. In one example, Product Quantization is used as the compression algorithm.
[0045] Figure 3 It is a schematic diagram showing an example of the information stored in the SSD 23 and the DRAM 24 when the search device 2 for explaining the embodiment performs a search.
[0046] The index information IDX shown is stored in the SSD 23. Figure 2 The index information IDX shown.
[0047] The search program SPG is loaded into the DRAM 24. The processor 21 executes a search operation in accordance with the search program SPG loaded into the DRAM 24.
[0048] In addition, in the DRAM 24, a storage area for temporarily storing information required for the search operation, that is, the cache area 241, is allocated.
[0049] The processor 21 executes a search operation under the control of the search program SPG. In the search operation, the processor 21 searches for the vector closest to the query input from the outside among the vector sets.
[0050] As described above, the index information IDX has a configuration in which the compressed vectors of all adjacent nodes are recorded in each node information. The processor 21 reads one node information related to the search target node from the SSD 23 and stores it in the cache area 241 of the DRAM 24. Then, the processor 21 completes one jump operation based on the compressed vectors of all adjacent nodes included in the stored node information. The processor 21 specifies the vector closest to the query by repeating the jump operation.
[0051] In this way, according to the embodiment, the index information IDX has a configuration in which the compressed vectors of all adjacent nodes required for one jump operation are recorded in each node information. Therefore, it is not necessary to previously store the compressed vectors related to all vectors in the search range in the DRAM 24. Thus, a search operation with a reduced memory usage can be performed.
[0052] When the processor 21 reads the index information IDX under the control of the search program SPG, it issues an IO request for reading from the SSD 23. Specifically, the storage area of the SSD 23 observed from the processor 21 under the control of the search program SPG is divided into a plurality of unit storage areas having a common size. The IO request is a request for accessing (reading or writing) one desired unit storage area among the plurality of unit storage areas. That is, the unit storage area can be regarded as the unit for accessing the SSD 23. The unit storage area can be a page, a block, a cluster, a sector, etc. in the NAND flash memory, or can be different from any of these. When the SSD 23 includes one or more flash memory dies, one page or a part of one page included in a single flash memory die can also be set as the unit storage area. When the SSD 23 includes a plurality of flash memory dies, a group of pages included in each of two or more flash memory dies capable of reading data simultaneously or in parallel, or a partial group of each page can also be set as the unit storage area. As the memory device of the search device 2, when the HDD is provided in place of the SSD 23 in the search device 2, a continuous data storage area at a position within a single data track segment can also be set as the unit storage area. One or a plurality of continuously positioned data sectors within a single data track segment can also be set as the unit storage area. Alternatively, a plurality of adjacent data track segments capable of jointly reading data can also be set as the unit storage area. The processor 21 can obtain the information stored in the target unit storage area in the index information IDX according to one IO request.
[0053] As Figure 2 shown, in the embodiment, the index information IDX has a configuration in which only one node information is stored in one unit storage area. According to the configuration, in the cache area 241 of the DRAM 24, only one node information obtained by one IO request is stored. Since the processor 21 can obtain the node information required for the jump operation with one IO request in the cache area 241 and can obtain only one node information in the cache area 241, the access efficiency to the SSD 23 is high, and the usage amount of the DRAM 24 can be suppressed.
[0054] Next, an apparatus (denoted as generation apparatus 1) for implementing a method for generating the index information IDX will be described.
[0055] Figure 4 is a schematic diagram showing an example of the configuration of the generation apparatus 1 of the embodiment.
[0056] The generation apparatus 1 includes a processor 11, a first interface 12, a second interface 13, a DRAM 14, and a bus 15. The first interface 12, the processor 11, the DRAM 14, and the second interface 13 are electrically connected to the bus 15.
[0057] The first interface 12 is a circuit that receives data from an external device of the generation device 1. In the example, the first interface 12 is a device for communicating with an external device via the network 3. For example, the first interface 12 is an Ethernet TM adapter, or a Wi-Fi TM adapter, etc.
[0058] The second interface 13 is a circuit that outputs data to an external device. In the example, the second interface 13 is an adapter for connecting to the storage device 4. The type of the storage device 4 is not limited to a specific type. The storage device 4 can also be, for example, an SSD, an HDD, or a UFS (Universal Flash Storage).
[0059] The processor 11 is an arithmetic device having a function of generating index information IDX. The processor 11 can also be, for example, a CPU. When the processor 11 is a CPU, the processor 11 realizes the function of generating index information IDX by executing a specified program. Specifically, the processor 11 obtains the generation program GPG from a specified position, loads the obtained generation program GPG into the DRAM 14. And, the processor 11 executes the search program SPG loaded into the DRAM 14. The processor 11 generates index information IDX under the control of the search program SPG. In addition, according to the functions to be implemented, one or more processors 11 can be provided.
[0060] Figure 4 In the example shown, the generation device 1 receives data from an external device via the network 3 and the first interface 12, and outputs data to an external device (here, the storage device 4) via the second interface 13. The reception of data and the output of data can also be executed via the same interface. In addition, the generation device 1 has a storage device, and can also obtain data from the storage device, or output data to the storage device.
[0061] In addition, Figure 4 In the example shown, the generation device 1 is set as a device different from the search device 2. The search device 2 can also function as the generation device 1 by executing the generation program GPG.
[0062] Figure 5 is a schematic diagram showing an example of the functions realized by the processor 11 provided in the generation device 1 of the embodiment.
[0063] The processor 11 functions as an object node setting unit 101, an information acquisition unit 102, a compressed vector acquisition unit 103, a node information generation unit 104, and a memory writing unit 105.
[0064] The processor 11 can obtain the directed graph GF and the vector set via the network 3 and the first interface 12.
[0065] The object node setting unit 101 sets one node in the obtained directed graph GF as the object node. The index information IDX is generated by repeatedly executing a loop process for generating one node information. The object node is a node temporarily set as the generation object of the node information in one loop process.
[0066] The object node setting unit 101 sends the NID of the object node to the information acquisition unit 102.
[0067] If the information acquisition unit 102 identifies the object node based on the NID received from the object node setting unit 101, then it acquires the NIDs of all adjacent nodes of the object node from the directed graph GF. And, the information acquisition unit 102 sends the acquired NIDs of all adjacent nodes to the compressed vector acquisition unit 103 and the node information generation unit 104.
[0068] In addition, the information acquisition unit 102 acquires the full-precision vector of the object node from the vector set. And, the information acquisition unit 102 sends the acquired full-precision vector of the object node to the node information generation unit 104.
[0069] The compressed vector acquisition unit 103 acquires the full-precision vectors of the adjacent nodes from the vector set and compresses the acquired full-precision vectors. As described above, the compression algorithm is not limited to a specific algorithm. The compressed vector acquisition unit 103 acquires the compressed vectors related to all adjacent nodes of the object node and sends the acquired compressed vectors of all adjacent nodes to the node information generation unit 104.
[0070] In addition, the generation of the compressed vector may not necessarily be performed by the compressed vector acquisition unit 103. Prepare the compressed vector in advance for each full-precision vector contained in the vector set, and the compressed vector acquisition unit 103 may also acquire the compressed vector via the network 3 and the first interface 12.
[0071] The node information generation unit 104 generates the node information of the object node including the full-precision vector of the object node, the NIDs of all adjacent nodes of the object node, and the compressed vectors of all adjacent nodes of the object node. The node information generation unit 104 sends the generated node information of the object node to the memory writing unit 105.
[0072] The memory writing unit 105 writes the node information of the target node into the storage device 4 via the second interface 13. The memory writing unit 105 writes the node information of the target node into the area of the unit storage area corresponding to the storage destination in the memory device (for example, the SSD 23 of the search device 2). More specifically, in the storage device 4, a plurality of areas corresponding to different unit storage areas are provided. The memory writing unit 105 completes the index information IDX in the storage device 4 by writing the node information of different nodes into each of the plurality of areas. When the index information IDX is transmitted to the SSD 23 according to the index information IDX written in this way, each node information is stored in a different unit storage area.
[0073] In addition, a part or all of the target node setting unit 101, the information acquisition unit 102, the compressed vector acquisition unit 103, the node information generation unit 104, and the memory writing unit 105 may also be implemented by a hardware circuit such as an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit).
[0074] Figure 6 It is a flowchart showing an example of the operation of the generation device 1 of the embodiment.
[0075] First, the target node setting unit 101 sets one node as the target node based on the directed graph GF (S101). The setting method of the target node is not limited to a specific method. For example, the target node setting unit 101 may also set the target node in the order of NID.
[0076] The information acquisition unit 102 acquires the NID of the target node, the NIDs of all adjacent nodes of the target node, and the full-precision vector of the target node (S102). The information acquisition unit 102 identifies the target node by receiving the NID of the target node from the target node setting unit 101. The information acquisition unit 102 acquires the NIDs of all adjacent nodes of the target node from the directed graph GF. The information acquisition unit 102 acquires the full-precision vector of the target node from the vector set.
[0077] The information acquisition unit 102 may also acquire the NIDs of all adjacent nodes of the target node via the network 3 and the first interface 12. Alternatively, the directed graph GF may be transmitted to the DRAM 14 via the network 3 and the first interface 12, and then the information acquisition unit 102 may acquire the NIDs of all adjacent nodes of the target node from the directed graph GF stored in the DRAM 14.
[0078] Similarly, the information acquisition unit 102 can also acquire the full-precision vector of the target node via the network 3 and the first interface 12. Alternatively, the vector set can be transmitted to the DRAM 14 via the network 3 and the first interface 12, and then the information acquisition unit 102 can acquire the full-precision vector of the target node from the vector set stored in the DRAM 14.
[0079] The compressed vector acquisition unit 103 acquires the compressed vectors of all adjacent nodes of the target node based on the NIDs of all adjacent nodes of the target node (S103).
[0080] The node information generation unit 104 generates the node information of the target node (S104). The node information generation unit 104 generates an information piece including the full-precision vector of the target node, the NIDs of all adjacent nodes of the target node, and the compressed vectors of all adjacent nodes of the target node as the node information of the target node.
[0081] The memory writing unit 105 writes the generated node information of the target node into the area corresponding to the unit storage area of the storage destination in the memory device (e.g., the SSD 23 of the search device 2) (S105).
[0082] The target node setting unit 101 determines whether there is a node in the plurality of nodes included in the directed graph GF that has not been set as a target node (S106). In the case where there is a node that has not been set as a target node (S106: Yes), the target node setting unit 101 sets the node that has not been set as a target node as a target node (S107), and the control transfers to S102.
[0083] In the case where there is no node that has not been set as a target node (S106: No), the operation of the generation device 1 ends.
[0084] Figure 7 It is a schematic diagram showing an example of the functions implemented by the processor 21 included in the search device 2 of the embodiment.
[0085] The processor 21 functions as a start node acquisition unit 201, a node information acquisition unit 202, and an arithmetic unit 203.
[0086] The start node acquisition unit 201 acquires the NID of the start node. And the start node acquisition unit 201 sends the acquired start node NID to the node information acquisition unit 202.
[0087] The method for obtaining the NID of the start node is not limited to a specific method. The start node acquisition unit 201 may also regard the node with the smallest NID value as the start node and obtain the node information of the start node from the index information IDX. The NID of the start node may also be given from outside the search device 2, or the NID of the start node may be stored in the SSD 23 in advance, and the start node acquisition unit 201 acquires the NID.
[0088] As described above, in the search operation, a jump operation is performed to sequentially switch the search target node according to the directed graph GF defined by the index information IDX. The start node is the node set as the search target node for the first jump operation. The node information acquisition unit 202 and the arithmetic unit 203 regard the start node as the search target node for the first jump operation and start the first jump operation.
[0089] The node information acquisition unit 202 acquires the node information of the search target node from the index information IDX stored in the SSD 23 in advance. In addition, before performing the first jump operation, the node indicated by the NID received from the start node acquisition unit 201 is regarded as the search target node. After performing the first jump operation, the node indicated by the NID received from the arithmetic unit 203 is regarded as the search target node.
[0090] The node information acquisition unit 202 stores the acquired node information in the cache area 241. In addition, the node information acquisition unit 202 sends the NID of the search target node and the NIDs of all adjacent nodes of the search target node to the arithmetic unit 203.
[0091] The arithmetic unit 203 acquires the compressed vector and full-precision vector of the search target node and the compressed vector of each of all adjacent nodes from the cache area 241. And the arithmetic unit 203 calculates the distance to the query for each of the search target node and each of all adjacent nodes of the search target node based on the compressed vectors of the respective nodes. The distance obtained by this calculation using the compressed vector is denoted as the first distance.
[0092] In addition, the compressed vector of the current search target node is not included in the node information of the current search target node. The compressed vector of the current search target node is included in the node information of one adjacent node of the previous search target node. Thus, in the cache area 241, in each jump operation, among the compressed vectors of all adjacent nodes included in the node information of the search target node, at least the compressed vector of the jump destination node remains in an effective state. That is to say, in one jump operation, in addition to the node information of the current search target node, the compressed vector of the current search target node included in the node information of the previous search target node can also be obtained from the cache area 241.
[0093] Data in the cache area 241 being "valid" means that the data can be read. Data in the cache area 241 being "invalid" means that the data cannot be read from the cache area 241. Specifically, data in the cache area 241 being "invalid" means that the data has been erased from the cache area 241, or the location storing the data can be used to store other data.
[0094] The arithmetic unit 203 determines whether to end the jump operation, for example, based on the first distance calculated for each of the search target node and all adjacent nodes of the search target node. When the first distance from the search target node to the query is shorter than the first distance from any adjacent node of the search target node to the query, the arithmetic unit 203 determines to end the jump operation. When there is an adjacent node whose first distance from any adjacent node of the search target node to the query is shorter than the first distance from the search target node to the query, the arithmetic unit 203 determines not to end the jump operation. In addition, the method for determining whether to end the jump operation is not limited to this.
[0095] When it is determined not to end the jump operation, the arithmetic unit 203 determines the node of the jump destination from all adjacent nodes of the search target node. The arithmetic unit 203 determines the adjacent node with the shortest first distance among all adjacent nodes of the search target node as the node of the jump destination. The arithmetic unit 203 sends the NID of the node of the jump destination to the node information acquisition unit 202.
[0096] In addition, the arithmetic unit 203 calculates the distance from the search target node to the query using the full-precision vector of the search target node. The distance obtained by this calculation using the full-precision vector is denoted as the second distance. The arithmetic unit 203 stores the second distance in the cache area 241.
[0097] Before the jump operation ends, the second distances related to all nodes on the path of the jump operation are stored in the cache area 241 in a valid state. If it is determined to end the jump operation, then the arithmetic unit 203 specifies the node closest to the query among all nodes on the jump path based on the second distances of all nodes on the jump path. And, the arithmetic unit 203 outputs the information indicating the specified node as the search result.
[0098] The search result output by the arithmetic unit 203 is not limited to specific information as long as it corresponds to the node specified as the node closest to the query. For example, the arithmetic unit 203 can output the NID or full-precision vector of the node specified as the node closest to the query as the search result. Here, as an example, the arithmetic unit 203 outputs the full-precision vector of the node specified as the node closest to the query as the search result.
[0099] In addition, part or all of the start node acquisition unit 201, the node information acquisition unit 202, and the arithmetic unit 203 can also be implemented by a hardware circuit such as an FPGA or an ASIC.
[0100] Figure 8 It is a flowchart showing an example of the operation of the search device 2 according to the embodiment.
[0101] If the arithmetic unit 203 acquires a query (S201), then the start node acquisition unit 201 acquires the NID of the start node (S202). If the node information acquisition unit 202 identifies the start node by the NID of the start node, then the start node is set as the search target node (S203).
[0102] The node information acquisition unit 202 reads the node information of the search target node in the index information IDX from the SSD 23, and stores the read node information of the search target node in the cache area 241 (S204).
[0103] The arithmetic unit 203 calculates the distance (first distance) between each of the search target node and all adjacent nodes of the search target node and the query (S205). In S205, the arithmetic unit 203 calculates the first distance related to each node by using the compressed vector of the search target node included in the node information stored in the cache area 241 and the compressed vectors of all adjacent nodes of the search target node.
[0104] In addition, the arithmetic unit 203 calculates the distance (second distance) between the search target node and the query by using the full-precision vector of the search target node included in the node information stored in the cache area 241 (S206).
[0105] The arithmetic unit 203 determines whether to end the jump operation (S207). In the case where it is determined not to end the jump operation (S207: NO), the arithmetic unit 203 determines the node of the jump destination (S208). The arithmetic unit 203 determines the node of the jump destination based on the first distance related to each node calculated by the process of S205.
[0106] The arithmetic unit 203 stores the second distance related to the search target node in the cache area 241 (S209). The process of S209 can also be executed after the process of S206 and before the process of S207. Further, the arithmetic unit 203 invalidates the data in the cache area 241 except for the second distance related to each node on the jump path and the compressed vector of the jump destination (S210).
[0107] The NID of the node at the jump destination is sent from the arithmetic unit 203 to the node information acquisition unit 202, and the node information acquisition unit 202 sets the node at the jump destination as the search target node (S211). Then, the control transfers to S204.
[0108] In addition, when the control transfers to S204 after passing through S211, the first distance between the new search target node and the query has been obtained through the process of S205 executed last time. Thus, in the process of S205, the arithmetic unit 203 can also omit the calculation of the first distance between the new search target node and the query.
[0109] When it is determined that the jump action has ended (S207: Yes), the arithmetic unit 203 specifies the node closest to the query based on the second distances related to the respective nodes on the jump path (S212). Then, the arithmetic unit 203 outputs the full-precision vector of the specified node as the search result (S213), and the operation of the search device 2 ends.
[0110] As described above, according to the embodiment, the generation device 1 sets one of the multiple nodes included in the directed graph GF as the target node (for example, refer to Figure 6 S101 and S107). The generation device 1 writes node information including the full-precision vector of the target node in the vector set, the ID related to each of all adjacent nodes of the target node in the vector set, and the compressed vector generated by compressing the corresponding full-precision vector (for example, refer to Figure 6 S102 to S105).
[0111] The index information IDX generated as described above has a configuration in which the compressed vectors of all adjacent nodes required for one jump action are recorded in each node information. Thus, it is not necessary to pre-store the compressed vectors related to all vectors in the search range in the DRAM. Thus, a search operation with reduced memory usage can be performed.
[0112] In addition, according to the embodiment, the generation device 1 executes a loop process including setting the target node and writing the node information of the target node multiple times (for example Figure 6 S102 to S107). In each of the multiple loop processes, the generation device 1 writes the node information in different regions among the multiple regions of the storage device 4 corresponding to the unit of access to the memory device (that is, the unit storage area).
[0113] Thus, during the search operation, the processor 21 of the search device 2 can obtain only the node information required for the jump action in the index information IDX with one IO request to the memory device, that is, the SSD23. Thus, during the search operation, the access efficiency to the memory device can be improved.
[0114] In addition, according to an embodiment, the search device 2 obtains a query (e.g., S201 in reference Figure 8 ). The search device 2 sets candidates for nodes closest to the query, i.e., search target nodes, according to the directed graph GF specified by the index information IDX (e.g., S203, S208, and S211 in reference Figure 8 ). Each time the search device 2 sets a search target node, it reads the node information of the search target node from the memory device, i.e., SSD23, and stores the read node information of the search target node in a memory, i.e., DRAM24, that is faster than the memory device for access operations (e.g., S204 in reference Figure 8 ). Based on the compressed vectors of all adjacent nodes included in the node information of the search target node stored in DRAM24, the search device 2 determines the node to jump to (e.g., S205 and S208 in reference Figure 8 ), and sets the node to jump to as a new search target node (e.g., S211 in reference Figure 8 ).
[0115] Thereby, each jump operation can be executed without storing the compressed vectors related to all vectors in the search range in DRAM. That is, a search operation with reduced memory usage can be performed.
[0116] In addition, according to the first embodiment, the search device 2 calculates a second distance using the full-precision vectors included in the node information of the search target node stored in DRAM24 (e.g., S206 in reference Figure 8 ). The search device 2 specifies the vector closest to the query based on the second distances of the respective nodes when they are set as search target nodes, i.e., the respective nodes on the jump path (e.g., S212 in reference Figure 8 ).
[0117] Thereby, each jump operation can be executed without storing the compressed vectors related to all vectors in the search range in DRAM. That is, a search operation with reduced memory usage can be performed.
[0118] Hereinafter, several modification examples of the embodiment will be described. In each modification example, matters different from the embodiment will be described. Description of matters the same as those in the embodiment will be omitted or briefly described.
[0119] (Modification Example 1)
[0120] In Variation 1, the node information generation unit 104 generates information pieces as node information. The information pieces include, in addition to the full-precision vector of the target node, the ID and the compressed vector associated with each of all adjacent nodes of the target node, the compressed vector of the target node is additionally appended. The compressed vector of the target node is generated by compressing the full-precision vector of the target node. The compressed vector of the target node is obtained by the compressed vector acquisition unit 103. The compressed vector acquisition unit 103 can obtain the compressed vector of the target node from the outside, or can obtain the compressed vector of the target node by generating it through compression using the full-precision vector of the target node.
[0121] Figure 9 FIG. is an example showing the configuration of the node information in Variation 1. Figure 9 The configuration of the node information of the node NIDi is shown as a representative of all node information. As Figure 9 shown, the node information in Variation 1 has a configuration in which the compressed vector of the node NIDi is appended to the node information in the embodiment shown in Figure 2 FIG.
[0122] Thus, each node information includes, in addition to the full-precision vector of the target node, the IDs and the compressed vectors of all adjacent nodes of the target node, the compressed vector of the target node. Accordingly, when each jump operation is completed, the arithmetic unit 203 of the search device 2 can invalidate all the node information of the search target node in the cache area 241.
[0123] (Variation 2)
[0124] Figure 10 FIG. is an example showing the configuration of the node information in Variation 2. In this figure, the configuration of the node information of the node NIDi is shown as a representative of all node information. As shown in this figure, the node information in Variation 2 has a configuration in which the out-degree R of the node NIDi is appended to the node information in the embodiment shown in Figure 2 FIG.
[0125] As described above, in one unit storage area, only one node information is stored. However, the size of the node information does not necessarily match the capacity of the unit storage area. There may be free areas left in the unit storage area where one node information is stored. Such free areas may also be filled. And the out-degree R may be different for each node.
[0126] When the node information acquisition unit 202 of the search device 2 reads the node information of the search target node from the unit storage area, it is necessary to specify the end of the node information stored in the unit storage area. According to Variation 2, the node information acquisition unit 202 can specify the end of the node information based on the out-degree R included in the node information.
[0127] For example, if the size of a full-precision vector is denoted as bfull bytes, the size of one compressed vector is denoted as bcomp bytes, the size of the numerical information of NID is denoted as b NID bytes, and the size of the numerical information of out-degree R is denoted as b R bytes, then the size b ND of the node information in Variation 2 can be expressed by the following formula (1).
[0128] bND = bfull + R × (bcomp + bNID) + bR (1)
[0129] The node information acquisition unit 202 uses formula (1) to calculate the size b ND of the node information. And the node information acquisition unit 202 acquires the information stored in the range from the starting position of the unit storage area to the position offset by b ND bytes from the starting position as the node information of the search target node.
[0130] In addition, in Variation 2, the out-degree R is recorded in each node information. When the out-degree R of all nodes is common, the out-degree R can be omitted from each node information, and the out-degree R can also be stored as a parameter common to all nodes in the memory device, that is, SSD23.
[0131] (Variation 3)
[0132] When generating the directed graph GF, the designer can make the out-degree of all nodes common or not. The designer can also determine the capacity of the unit storage area after determining the maximum out-degree.
[0133] Taking Variation 2 as an example. If the maximum out-degree is denoted as Rma x , then the maximum size b NDmax of the node information can be expressed by the following formula (2).
[0134] bNDmax = bfull + Rmax(bcomp + bNID) + bR (2)
[0135] The designer uses formula (2) to calculate the maximum size b ND max of the node information. And the designer sets the maximum size b NDmax of the node information as the capacity of the unit storage area. Thus, one node information can be stored in each unit storage area for sure.
[0136] In addition, the designer can also first determine the capacity of the unit storage area and determine the maximum out-degree based on the capacity of the unit storage area.
[0137] In the embodiment, Variation 1, Variation 2, and Variation 3, only one node information is stored in one unit storage area. However, the number of node information stored in one unit storage area is not limited to only one. Multiple node information may be stored in one unit storage area. Also, the size of one node information may be larger than the capacity of the unit storage area, and one node information may be stored within a range spanning two or more unit storage areas. However, one node information is stored within a logically continuous range in the SSD 23.
[0138] The logically continuous range is a range that is continuous in the logical address space provided by the memory device to the processor. Generally, the response speed of a memory device such as an SSD or a disk device when reading data from a logically continuous range is faster than when reading data from two or more logically discontinuous ranges. Thus, by storing each node information within a logically continuous range in the SSD 23, the time required to read each node information from the SSD 23 is suppressed.
[0139] In the description of the embodiment, Variation 1, Variation 2, and Variation 3, the node information includes the compression vectors of each adjacent node. The vectors of each adjacent node included in the node information are not limited to compression vectors. In the node information, instead of or in addition to the compression vectors of each adjacent node, full-precision vectors of each adjacent node may be included.
[0140] According to the above-described embodiment, Variation 1, Variation 2, and Variation 3, for example, the configurations of Supplementary Note 1 to Supplementary Note 3 below can also be obtained.
[0141] (Supplementary Note 1)
[0142] A generation method is a generation method of a processor configured to process data represented by a directed graph, including:
[0143] setting one of the multiple first nodes, each assigned an ID included in the directed graph, that is, one of the multiple first nodes corresponding to the multiple first vectors included in the search range, as a second node; and
[0144] writing an information piece that includes a first vector corresponding to the second node among the multiple first vectors, that is, a second vector, IDs associated with each of all the outer adjacent nodes of the second node among the multiple first nodes, that is, one or more third nodes, and a third vector; and the third vector is a vector corresponding to the third node, and the information piece is an element related to the second node in the index information corresponding to the directed graph.
[0145] (Supplementary Note 2)
[0146] According to the generation method described in Supplementary Note 1, wherein
[0147] The writing includes generating the information piece that arranges the second vector at the beginning.
[0148] (Appendix 3)
[0149] The generation method according to Appendix 1 or 2, wherein
[0150] The information piece does not include the ID of the second vector.
[0151] Several embodiments of the present invention have been described, but these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be implemented in various other ways, and various omissions, substitutions, and changes can be made without departing from the gist of the invention. These embodiments or their variations are included in the scope or gist of the invention and are included in the invention described in the claims and its equivalents.
[0152] [Symbol Explanation]
[0153] 1 Generation device
[0154] 2 Search device
[0155] 3 Network
[0156] 4 Storage device
[0157] 11 Processor
[0158] 12 First interface
[0159] 13 Second interface
[0160] 14 DRAM
[0161] 15 Bus
[0162] 21 Processor
[0163] 22 Interface
[0164] 23 SSD
[0165] 24 DRAM
[0166] 25 Bus
[0167] 101 Object node setting unit
[0168] 102 Information acquisition unit
[0169] 103 Compressed vector acquisition unit
[0170] 104 Node information generation unit
[0171] 105 Memory writing unit
[0172] 201 Start Node Acquisition Unit
[0173] 202 Node Information Acquisition Unit
[0174] 203 Arithmetic Unit
[0175] 241 Cache Area
[0176] GF Directed Graph
[0177] GPG Generation Program
[0178] IDX Index Information
[0179] SPG Search Program
Claims
1. A generation method is a method for generating a processor configured to process data represented by a directed graph, comprising: setting, as a second node, one of a plurality of first nodes each assigned an ID included in the directed graph, that is, one of the plurality of first nodes corresponding to a plurality of first vectors included in a search range; and writing an information piece that includes a second vector, which is the first vector corresponding to the second node among the plurality of first vectors, IDs associated with each of all outer adjacent nodes of the second node among the plurality of first nodes, that is, one or more third nodes, and a third vector, where the third vector is a vector corresponding to the third node, and the information piece is an element related to the second node in index information corresponding to the directed graph.
2. The generation method according to claim 1, wherein the third vector is the first vector corresponding to the third node, or a vector generated by compressing the first vector corresponding to the third node.
3. The generation method according to claim 1, wherein performing a first action multiple times, each of the multiple first actions including the setting and the writing, in each of the multiple first actions, the setting is to set a first node among the plurality of first nodes that has not been set as the second node, performing the multiple first actions until there is no first node that has not been set as the second node.
4. The generation method according to claim 1, wherein the writing includes adding the number of outer adjacent nodes of the second node to the information piece.
5. The generation method according to any one of claims 1 to 4, wherein performing a first action multiple times, each of the multiple first actions including the setting and the writing, in the writing of each of the multiple first actions, writing the information piece in a different first storage area among a plurality of first storage areas, each of the plurality of first storage areas corresponding to a unit for accessing a memory device.
6. The generation method according to claim 1, wherein the writing includes adding a fourth vector generated by compressing the second vector to the information piece.
7. A search method, comprising: obtaining a query; and setting, according to a directed graph specified by index information, a candidate for a first node closest to the query; the directed graph includes a plurality of first nodes corresponding to a search range, that is, a plurality of first vectors, the index information is stored in a memory device, the index information includes a plurality of first information pieces, each of the plurality of first information pieces includes a second vector, which is the first vector corresponding to one first node among the plurality of first vectors, IDs associated with each of all outer adjacent nodes of the one first node among the plurality of first nodes, that is, one or more second nodes, and a third vector, where the third vector is a vector corresponding to the second node; and Setting the candidate includes: reading a first information piece, which is a second information piece, related to a first node, which is a third node, of the candidate from a first storage area corresponding to a unit of access to the memory device, and storing the read second information piece in a memory that is faster than the memory device in access operation; and Based on one or more of the third vectors included in the second information piece stored in the memory, setting a new candidate first node.
8. The search method according to claim 7, wherein The third vector is a first vector corresponding to the second node, or a vector generated by compressing the first vector corresponding to the second node.
9. The search method according to claim 7, further comprising:[[]] After storing the second information piece in the memory, calculating the distance between the candidate first node and the query using the first vector included in the stored second information piece; and Based on the distances between each of the multiple first nodes set as the candidate and the query, specifying the vector closest to the query.
10. A generation device, comprising:[[]] An interface circuit configured to receive a directed graph including multiple first nodes each assigned an ID, and multiple first vectors included in a search range; and the multiple first nodes correspond to the multiple first vectors; and A processor configured to execute:[[]] Setting one of the multiple first nodes as a second node; and Writing an information piece that includes a first vector, which is a second vector, corresponding to the second node among the multiple first vectors, IDs related to all outer adjacent nodes, which are one or more third nodes, of the second node among the multiple first nodes, and a third vector; and the third vector is a vector corresponding to the third node, and the information piece is an element related to the second node in the index information corresponding to the directed graph.