Monotonic path aware vector diagram index disk layout optimization method and monotonic path aware vector diagram index disk layout optimization system

By constructing a target weighted undirected graph and optimizing disk layout, the problem of insufficient data locality in vector graph indexes is solved, thus improving query efficiency.

CN120872241APending Publication Date: 2025-10-31HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510929839.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing vector graph index disk layout methods do not consider the local characteristics of node paths, resulting in insufficient data locality and low query efficiency.

Method used

By constructing a target weighted undirected graph, the disk layout is optimized based on the target weights of nodes and edges. The target disk layout optimization function is used to allocate nodes to disk pages, reducing the number of I/O units and improving query efficiency.

Benefits of technology

It improves the data locality and query relevance of vector graph indexes on disk, reduces the number of I/O units during disk queries, and improves query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872241A_ABST
    Figure CN120872241A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of disk storage, and provides a monotonous path aware vector diagram index disk layout optimization method and system, and the method comprises the steps: preferentially distributing target nodes at the two ends of a front edge to a same target disk page through the sorting of target weights of edges corresponding to all nodes, according to the method, the data locality when each node in the vector diagram index is stored in the disk can be improved, so that the number of I / O units used when specified data in the disk is queried can be effectively reduced, and the query efficiency when the disk is queried is improved; and obtaining a total contribution value of each edge through the target disk layout optimization function and the target weight of the edge of each node, and sequentially distributing each out-of-neighbor node into the target disk page according to the total contribution value of the edge corresponding to the out-of-neighbor node of each target node in the target disk page to which the target node is distributed. And furthermore, the locality of the single target disk page data can be improved, so that the query efficiency when the target disk is queried can be further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of disk storage technology, and more specifically, relates to a method and system for optimizing the layout of a monotonic path-aware vector graph index disk. Background Technology

[0002] With the technological innovation of large-scale models in fields such as natural language processing and computer vision, massive amounts of cross-modal data, such as text, images, and videos, are transformed into high-dimensional vectors through embedding techniques to meet model training requirements. The rapid expansion of these large-scale model data sets has significantly increased the demands on query efficiency when searching data on disk. These models utilize vector graph indexing methods based on Approximate Nearest Neighbor (ANN) strategies for data retrieval, leading to a surge in disk storage technologies and storage layout optimization methods specifically for vector graph indexing.

[0003] Existing disk layout methods for vector graph indexes mostly fail to consider the local characteristics of node data or only consider the local characteristics between adjacent nodes during the layout process, rather than optimizing the local characteristics of each node at the path level. In other words, they do not consider the correlation between edges between nodes in the query path. This results in a severe lack of data locality when vector graph indexes are stored on disk, leading to a large number of redundant I / O units when querying disk data, thus reducing query efficiency. Summary of the Invention

[0004] To address the aforementioned shortcomings of existing technologies, this application provides a method and system for optimizing the disk layout of a monotonic path-aware vector graph index, aiming to solve the problem of reduced query efficiency caused by insufficient data locality in disk-stored vector graph indexes in existing methods.

[0005] In a first aspect, this application provides a method for optimizing disk layout of a monotonic path-aware vector graph index, including: S1. Construct a vector graph index based on the nearest neighbor relationship of the data vector to be indexed, and obtain the target weight of the edge corresponding to each node in the vector graph index. The node stores the number and number of the outgoing neighbor nodes corresponding to the node, and the edge is used to represent the connection relationship between the nodes. S2. Obtain the target weighted undirected graph based on the vector graph index and the target weight of each edge; S3. Establish the target disk layout optimization function, and based on the target disk layout optimization function and the target weighted undirected graph, allocate each node in the target weighted undirected graph to each disk page of the target disk.

[0006] Furthermore, a vector graph index is constructed based on the nearest neighbor relationships of the data vectors to be indexed, including: An initial set of nodes is constructed based on the nearest neighbor relationships of the data vector to be indexed, and an initial empty graph is constructed that does not contain any nodes or edges. Select nodes sequentially from the initial node set, obtain the edges corresponding to the nodes, and insert the nodes and their corresponding edges into the initial empty graph to obtain the vector graph index.

[0007] Furthermore, obtain the edges corresponding to the nodes, and insert the nodes and their corresponding edges into the initial empty graph, including: S11. Obtain the neighborhood of a node based on the target search algorithm, and select the nearest outgoing neighbor node from the neighborhood as the first node. The neighborhood includes all outgoing neighbor nodes corresponding to the node. S12. Construct the first positive edge from node to the first node, and select nodes from the neighborhood whose distance to the node is greater than the distance to the first node as the second node; S13. Construct the first reverse edge from the first node to the node. If the number of first reverse edges constructed at the first node does not reach the out-degree limit of the first node, insert the node, the first reverse edge, and the first forward edge into the initial empty graph. S14. Remove the first node and the second node from the neighborhood of the node, and repeat step S11 until the number of out-neighbor nodes in the neighborhood is empty or the number of the first positive edge corresponding to the node reaches the upper limit of the out-degree of the node.

[0008] Furthermore, the target weights of the edges corresponding to each node in the vector graph index are obtained using the following formula: ; in, This indicates that the node can be reached monotonically. The number of edges, Represents a node To the node The corresponding edge, Representing an edge The number of nodes that can be monotonically reached. Representing an edge The target weight.

[0009] The process of obtaining the target weights of the edges corresponding to each node is carried out synchronously with the process of constructing the vector graph index. This allows for the effective use of intermediate data in the vector graph index construction process, reducing the computational cost of additional target weight calculations and improving the execution efficiency of this method.

[0010] Furthermore, the target disk layout optimization function is shown in the following formula:

[0011] in, Represents a node To the node The corresponding edge, Representing an edge The target weight, Represents all nodes within the vector graph index. To the node The corresponding set of edges, Represents a node Indicates the indicator function, in and When they exist in the same disk page, the value is represented by the default indicator value. This represents the total contribution value of all edges in E.

[0012] The target disk layout optimization function is designed based on the target weight of each edge. By using the target disk layout optimization function, the total contribution of the edges corresponding to each node to the query speed can be taken into account during the disk layout process, thereby improving the data locality when storing each node in the vector graph index to the disk.

[0013] Furthermore, the nodes in the target weighted undirected graph are allocated to the disk pages of the target disk, including: S31. Sort each edge according to its target weight, and add the edges whose nodes at both ends have not been assigned to the set to be assigned. S32. Select any empty disk page of the target disk as the target disk page, and select the nodes at both ends of the edge with the largest target weight from the set to be allocated as the target nodes and store them in the target disk page. S33. Based on the target disk layout optimization function, calculate the total contribution value of all edges between the target node and the corresponding out-neighbor nodes of the target node, and store the corresponding out-neighbor nodes of the target node into the target disk page in order of total contribution value from high to low, until the target disk page is full or all the corresponding out-neighbor nodes of the target node have been stored into the target disk page. S34. For the edges in the set to be allocated that have not been selected, execute step S32 until all target disk pages or all edges in the set to be allocated have been selected.

[0014] The steps of sorting edges according to their target weights and then allocating edges whose nodes at both ends are not yet assigned to empty disk pages allow nodes corresponding to edges with higher target weights to be placed in the same disk page. Since one disk page corresponds to one disk I / O operation, the number of I / O units used when querying specified data on the disk can be reduced, thereby improving query efficiency. Based on the total contribution value of the edges between the target node and its corresponding out-neighbor nodes, the out-neighbor nodes corresponding to the target node are stored in the target disk page, which improves the query relevance of nodes within a single target disk page, thus enhancing the query efficiency when querying the target disk.

[0015] Furthermore, after all target disk pages or all edges in the set to be allocated have been selected, the process also includes: when there are unallocated nodes and incomplete target disk pages, storing the unallocated nodes into the incomplete target disk pages in sequence.

[0016] Furthermore, the nodes in the target weighted undirected graph are allocated to the disk pages of the target disk, including: Clustering algorithms are used to divide the nodes in the weighted undirected graph of the target into multiple target clusters according to the distance between the nodes, and the induced subgraphs corresponding to each target cluster are obtained. For each node and edge in each induced subgraph, perform step S31 until all nodes corresponding to each target cluster are stored in the target disk; Merge full disk pages in the target disk, and combine the stored nodes and unallocated nodes in the non-full disk pages into a target cluster, and perform step S31 on the target cluster.

[0017] The method of using a clustering algorithm to divide nodes into several target clusters based on distance, and then storing each target cluster in parallel to the target disk using a two-stage decoupling approach, can effectively speed up the efficiency of storing each node in the target weighted undirected graph to the target disk.

[0018] Secondly, this application also provides a monotonic path-aware vector graph index disk layout optimization system for implementing any of the methods in the first aspect, including: The index graph construction module is used to construct a vector graph index based on the nearest neighbor relationships of the data vectors to be indexed. The weight calculation module is used to obtain the target weight of the edge corresponding to each node; The index graph acquisition module is used to obtain the target weighted undirected graph based on the vector graph index and the target weight of each edge. The disk layout optimization module is used to establish the target disk layout optimization function, and based on the target disk layout optimization function and the target weighted undirected graph, to allocate each node in the target weighted undirected graph to each disk page of the target disk.

[0019] Thirdly, this application also provides an electronic device, comprising: at least one memory for storing a program; and at least one processor for executing the program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to execute the method described in any possible implementation of the first aspect.

[0020] In summary, compared with the prior art, the above-described technical solutions conceived by this invention have the following beneficial effects: The method of this application prioritizes the allocation of target nodes at both ends of the edges with higher rankings to the same target disk page by sorting the target weights of the edges corresponding to each node. This improves the data locality when storing nodes in the vector graph index to the disk, thereby effectively reducing the number of I / O units used when querying specified data on the disk, thus improving the query efficiency when the disk is queried. Furthermore, by using the target disk layout optimization function and the target weights of the edges corresponding to each node to obtain the total contribution value of each edge, and based on the total contribution value of the edges corresponding to the outgoing neighbor nodes of each target node in the target disk page of the allocated target node, each outgoing neighbor node is sequentially allocated to the target disk page. This improves the query relevance and data locality of each node in a single target disk page, further enhancing the query efficiency when the target disk is queried. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in this application or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart illustrating a monotonic path-aware vector graph index disk layout optimization method provided in this application embodiment.

[0023] Figure 2 This is a schematic diagram of a vector graph index provided in an embodiment of this application.

[0024] Figure 3 This application provides a schematic diagram of the structure of a monotonic path-aware vector graph index disk layout optimization system.

[0025] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0026] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0027] In the following description, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The following description provides multiple embodiments of this application, which can be substituted or combined with each other. Therefore, this application can also be considered to include all possible combinations of the same and / or different embodiments described. Thus, if one embodiment includes features A, B, and C, and another embodiment includes features B and D, then this application should also be considered to include embodiments containing one or more other possible combinations of A, B, C, and D, even if such embodiments are not explicitly described in the following text.

[0028] The following description provides examples and does not limit the scope, applicability, or examples set forth in the claims. Changes may be made to the function and arrangement of the described elements without departing from the scope of this application. Various processes or components may be appropriately omitted, substituted, or added to the examples. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Furthermore, features described with respect to some examples may be combined into other examples.

[0029] Figure 1 This is a flowchart illustrating the monotonic path-aware vector graph index disk layout optimization method provided in this application embodiment, as follows: Figure 1 As shown, the method includes at least the following steps: S1. Construct a vector graph index based on the nearest neighbor relationship of the data vector to be indexed, and obtain the target weight of the edge corresponding to each node in the vector graph index. The node stores the number and number of the outgoing neighbor nodes corresponding to the node, and the edge is used to represent the connection relationship between the nodes.

[0030] In this embodiment, the purpose of this application is to optimize the disk layout during disk storage by assigning target weights to each node and edge in the vector graph index. Therefore, the execution entity of this method can be the controller of the storage unit corresponding to the disk.

[0031] In one possible implementation, a vector graph index is constructed based on the nearest neighbor relationships of the data vector to be indexed, including: An initial set of nodes is constructed based on the nearest neighbor relationships of the data vector to be indexed, and an initial empty graph is constructed that does not contain any nodes or edges. Select nodes sequentially from the initial node set, obtain the edges corresponding to the nodes, and insert the nodes and their corresponding edges into the initial empty graph to obtain the vector graph index.

[0032] In this embodiment, the purpose of constructing the graph index is twofold. First, it is to initially determine the nearest neighbor relationships between nodes within the vector graph index, thereby preparing for subsequent optimization of the disk storage layout. Second, it is to obtain the target weights of each edge during the construction process. For example... Figure 2 As shown, during data storage or retrieval, cross-modal data such as text, images, and videos are transformed into vectors through embedding techniques. The core of vector graph indexing is to use a graph structure to represent vectors and the nearest neighbor relationships between vectors, where nodes are p1 to p2 in the figure. 12 Each edge has a corresponding vector, and each edge represents the nearest neighbor relationship between the vectors of the connected nodes.

[0033] In one possible implementation, the edges corresponding to the nodes are obtained, and the nodes and their corresponding edges are inserted into the initial empty graph, including: S11. Obtain the neighborhood of a node based on the target search algorithm, and select the nearest outgoing neighbor node from the neighborhood as the first node. The neighborhood includes all outgoing neighbor nodes corresponding to the node. S12. Construct the first positive edge from node to the first node, and select nodes from the neighborhood whose distance to the node is greater than the distance to the first node as the second node; S13. Construct the first reverse edge from the first node to the node. If the number of first reverse edges constructed at the first node does not reach the out-degree limit of the first node, insert the node, the first reverse edge, and the first forward edge into the initial empty graph. S14. Remove the first node and the second node from the neighborhood of the node, and repeat step S11 until the number of out-neighbor nodes in the neighborhood is empty or the number of the first positive edge corresponding to the node reaches the upper limit of the out-degree of the node.

[0034] In this embodiment, the out-degree of a node is used to represent the number of outgoing edges of the node, and the number of outgoing edges has an upper limit, which is less than the number of outgoing neighbor nodes in the node's neighborhood. The selected target search algorithm is the ANN algorithm, and steps S11 to S14 belong to the incremental node insertion method. (The last sentence appears to be incomplete and possibly refers to a different implementation.) For example, this invention performs an ANN search on the current initial empty graph to obtain... neighborhood And then according to Distance traversal The nodes in. (Note: The original text appears to be incomplete and contains several typographical errors. A more accurate translation would require the full context.) Current distance The nearest node is , Corresponding to the first node in step S11, this embodiment will attempt to... and Add forward and reverse edges between them until... empty or The out-degree has reached the maximum out-degree limit. Add a forward edge. After success, Will from Removed from the middle, at the same time All of the above satisfy nodes Nodes will also be removed. This corresponds to the second node in step S12. If a reverse edge is added... back If the out-degree does not exceed the maximum out-degree limit, then the reverse edge has been successfully added; otherwise, The set of out-neighbor nodes and Together as neighborhood , re-for Perform the operation of adding forward edges once. By performing steps S11 to S14 on each node, the vector graph index can be obtained.

[0035] In one possible implementation, the target weights of the edges corresponding to each node in the vector graph index are obtained using the following formula: ; in, This indicates that the node can be reached monotonically. The number of edges, Represents a node To the node The corresponding edge, Representing an edge The number of nodes that can be monotonically reached. Representing an edge The target weight.

[0036] In this embodiment of the application, the process of obtaining the target weights of the edges corresponding to each node is carried out synchronously with the process of constructing the vector graph index.

[0037] This embodiment uses the edge A monotonic path For example, side The number of nodes that can be monotonically reached is... It can be used to represent edges The frequency with which the search path passes through the selected edge during the search process. With nodes , Able to arrive monotonously It means that there exists a containment And with A monotonic path ending at [a certain point]. And there exists [a certain path]. The necessary and sufficient condition for a monotonic path is that the edges are... There are corresponding blocking nodes. , i.e., node satisfy , Represents the distance function.

[0038] The calculation is performed in two groups: one for two-node monotonic paths and the other for multi-node monotonic paths. The results from both groups are then summed. For Corresponding two-node monotonic path ,side Only nodes that can be monotonically reached One, therefore the corresponding =1; for monotonic paths containing multiple nodes , If the number of monotonically reachable nodes equals the number of blocked nodes, then it is a monotonic path. corresponding The value represents the total number of blocked nodes.

[0039] The above Used to characterize nodes When it is the starting point of the search path, the search path passes through edges. The frequency of, and for If it is not the starting point of the search path, then it is necessary to To represent a node that can be monotonically reached. The number of edges, used to measure the number of nodes reached by the search path. The frequency of.

[0040] The calculation is performed in two groups: one for monotonic paths with two nodes and the other for monotonic paths with multiple nodes. The results from both groups are then summed. For reachable nodes... The two-node monotonic path, corresponding to The value is the node. The in-degree is the number of nodes that can be directly reached. The number of edges; for nodes that can be reached. In a multi-node monotonic path, each multi-node monotonic path has an edge that can block node p. The value is the number of edges that can block node p.

[0041] S2. Based on the vector graph index and the target weights of each edge, obtain the target weighted undirected graph.

[0042] In this embodiment, the purpose of obtaining the target weighted undirected graph is to use the target disk layout optimization function to obtain the data locality of each edge corresponding to the vector graph index. Thus, the disk layout optimization problem of the vector graph index can be equivalent to a weighted undirected graph under equilibrium constraints. Minimum cut problem.

[0043] S3. Establish the target disk layout optimization function, and based on the target disk layout optimization function and the target weighted undirected graph, allocate each node in the target weighted undirected graph to each disk page of the target disk.

[0044] In one possible implementation, the target disk layout optimization function is shown in the following formula:

[0045] in, Represents a node To the node The corresponding edge, Representing an edge The target weight, Represents all nodes within the vector graph index. To the node The corresponding set of edges, Represents a node Indicates the indicator function, in and When they exist in the same disk page, the value is represented by the default indicator value. This represents the total contribution value of all edges in E.

[0046] In one possible implementation, allocating nodes in the target weighted undirected graph to disk pages on the target disk includes: S31. Sort each edge according to its target weight, and add the edges whose nodes at both ends have not been assigned to the set to be assigned. S32. Select any empty disk page of the target disk as the target disk page, and select the nodes at both ends of the edge with the largest target weight from the set to be allocated as the target nodes and store them in the target disk page. S33. Based on the target disk layout optimization function, calculate the total contribution value of all edges between the target node and the corresponding out-neighbor nodes of the target node, and store the corresponding out-neighbor nodes of the target node into the target disk page in order of total contribution value from high to low, until the target disk page is full or all the corresponding out-neighbor nodes of the target node have been stored into the target disk page. S34. For the edges in the set to be allocated that have not been selected, execute step S32 until all target disk pages or all edges in the set to be allocated have been selected.

[0047] In this embodiment, the above steps are equivalent to dividing the nodes into several subsets within a disk page without repetition or omission, based on the target weights of each edge in the target weighted undirected graph, thereby maximizing the sum of the weights of edges whose endpoints are in the same subset. Since each disk page corresponds to one disk I / O operation, storing the nodes at both ends of an edge with a higher target weight into a single disk page reduces the number of times multiple disk pages are queried, thus reducing the number of I / O units used when the specified data on the disk is queried, thereby improving query efficiency. Furthermore, storing the out-neighbor nodes corresponding to the target node into the target disk page based on the total contribution value of all edges between the target node and its corresponding out-neighbor nodes aims to store nodes along the monotonic path corresponding to the same edge in the same disk page as much as possible. This improves the query relevance and data locality of nodes within a single target disk page, thereby enhancing query efficiency when the target disk is queried.

[0048] In one possible implementation, after all target disk pages or all edges in the set to be allocated have been selected, the process further includes: When there are unallocated nodes and incomplete target disk pages, the unallocated nodes are sequentially stored into the incomplete target disk pages. In this embodiment of the application, the unallocated nodes are sequentially stored into the unfilled target disk pages in order to reduce the number of unfilled disk pages and prevent omissions during node storage.

[0049] In one possible implementation, allocating nodes in the target weighted undirected graph to disk pages on the target disk includes: Clustering algorithms are used to divide the nodes in the weighted undirected graph of the target into multiple target clusters according to the distance between the nodes, and the induced subgraphs corresponding to each target cluster are obtained. For each node and edge in each induced subgraph, perform step S31 until all nodes corresponding to each target cluster are stored in the target disk; Merge full disk pages in the target disk, and combine the stored nodes and unallocated nodes in the non-full disk pages into a target cluster, and perform step S31 on the target cluster.

[0050] In the embodiments of this application, In vector graph indexing, since an edge represents the nearest neighbor relationship between two nodes, the probability of an edge existing between two nodes that are far apart is low.

[0051] Based on this observation, this invention uses a clustering algorithm to divide the vector into several clusters. This corresponds to pre-dividing the vector graph index into multiple subsets with coarse granularity. This classification method reduces the number of edges that are cut off and stored in different disk pages when each edge is stored on disk. The above-mentioned processing method for each target cluster is essentially a two-stage decoupling approach. In the first stage, step S31 is applied in parallel to each target cluster. In the second stage, the stored nodes and unallocated nodes in the incomplete disk pages are combined into a single target cluster, and step S31 is executed again. The first stage can use multi-threading technology to store multiple target clusters simultaneously, thus effectively accelerating the efficiency of storing each node on the target disk. The second stage merges the full disks and stores the stored nodes and unallocated nodes in the incomplete disk pages, achieving a single data traversal for each node. This optimizes the disk layout of the vector graph index, further improving the efficiency of storing each node on the target disk.

[0052] Figure 3 This is a schematic diagram of the structure of the monotonic path-aware vector graph index disk layout optimization system provided in the embodiments of this application, as shown below. Figure 3 As shown, the system includes at least: The index graph construction module is used to construct a vector graph index based on the nearest neighbor relationships of the data vectors to be indexed. The weight calculation module is used to obtain the target weight of the edge corresponding to each node; The index graph acquisition module is used to obtain the target weighted undirected graph based on the vector graph index and the target weight of each edge. The disk layout optimization module is used to establish the target disk layout optimization function, and based on the target disk layout optimization function and the target weighted undirected graph, to allocate each node in the target weighted undirected graph to each disk page of the target disk.

[0053] like Figure 4 As shown, Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: a processor 401, a communications interface 402, a memory 403, and a communication bus 404. The processor 401, communications interface 402, and memory 403 communicate with each other via the communication bus 404. The processor 401 can call software instructions in the memory 403 to execute the methods described in the above embodiments.

[0054] Furthermore, the logical instructions in the aforementioned memory 403 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application.

[0055] Based on the methods in the above embodiments, this application provides a computer-readable storage medium storing a computer program that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0056] Based on the methods in the above embodiments, this application provides a computer program product that, when run on a processor, causes the processor to execute the methods in the above embodiments.

[0057] It is understood that the processor in the embodiments of this application can be a CPU (Central Processing Unit), or other general-purpose processors, DSPs (Digital Signal Processors), ASICs (Application Specific Integrated Circuits), FPGAs (Field Programmable Gate Arrays), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. A general-purpose processor can be a microprocessor or any conventional processor.

[0058] The method steps in this application embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in random access memory (RAM), flash memory, ROM (Read-only Memory), PROM (Programmable ROM), EPROM (Erasable PROM), EEPROM (Electrically Erasable EPROM), registers, hard disks, portable hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC.

[0059] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line DSL) or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., SSD (Solid State Disk)).

[0060] It is understood that the various numerical designations used in the embodiments of this application are merely for the convenience of description and are not intended to limit the scope of the embodiments of this application.

[0061] Those skilled in the art will readily understand that the above are merely preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for optimizing disk layout using a monotonic path-aware vector graph index, characterized in that, include: S1. Construct a vector graph index based on the nearest neighbor relationship of the data vector to be indexed, and obtain the target weight of the edge corresponding to each node in the vector graph index. The node stores the number and number of the outgoing neighbor nodes corresponding to the node, and the edge is used to represent the connection relationship between the nodes. S2. Based on the vector graph index and the target weight of each edge, obtain the target weighted undirected graph; S3. Establish a target disk layout optimization function, and based on the target disk layout optimization function and the target weighted undirected graph, allocate each node in the target weighted undirected graph to each disk page of the target disk.

2. The monotonic path-aware vector graph index disk layout optimization method according to claim 1, characterized in that, The construction of the vector graph index based on the nearest neighbor relationships of the data vector to be indexed includes: An initial set of nodes is constructed based on the nearest neighbor relationships of the data vector to be indexed, and an initial empty graph is constructed that does not include any of the nodes and edges mentioned above. Select nodes sequentially from the initial node set, obtain the edges corresponding to the nodes, and insert the nodes and their corresponding edges into the initial empty graph to obtain the vector graph index.

3. The monotonic path-aware vector graph index disk layout optimization method according to claim 2, characterized in that, The step of obtaining the edge corresponding to the node and inserting the node and the edge corresponding to the node into the initial empty graph includes: S11. Obtain the neighborhood of the node based on the target search algorithm, and select the outgoing neighbor node closest to the node from the neighborhood as the first node. The neighborhood includes all the outgoing neighbor nodes corresponding to the node. S12. Construct a first positive edge from the node to the first node, and select a node from the neighborhood whose distance to the node is greater than the distance to the first node as the second node; S13. Construct a first reverse edge from the first node to the node, and if the number of the first reverse edges constructed at the first node does not reach the out-degree limit of the first node, insert the node, the first reverse edge, and the first forward edge into the initial empty graph. S14. Remove the first node and the second node from the neighborhood of the node, and repeat step S11 until the number of out-neighbor nodes in the neighborhood is empty or the number of the first positive edges corresponding to the node reaches the out-degree limit of the node.

4. The monotonic path-aware vector graph index disk layout optimization method according to claim 1, characterized in that, The target weights of the edges corresponding to each node in the vector graph index are obtained by the following formula: ; in, This indicates that the node can be reached monotonically. The number of edges, Represents a node To the node The corresponding edges, Representing an edge The number of nodes that can be monotonically reached. Representing an edge The target weight.

5. The monotonic path-aware vector graph index disk layout optimization method according to claim 4, characterized in that, The target disk layout optimization function is shown in the following formula: in, Represents a node To the node The corresponding edges, Representing an edge The target weight, Represents all nodes within the vector graph index. To the node The corresponding set of edges, Represents a node Indicates the indicator function, in and When they exist in the same disk page, the value is represented by the default indicator value. This represents the total contribution value of all edges in E.

6. The monotonic path-aware vector graph index disk layout optimization method according to claim 5, characterized in that, The step of allocating each node in the target weighted undirected graph to each disk page of the target disk includes: S31. Sort each edge according to the target weight corresponding to the edge, and put the edge whose nodes at both ends have not been assigned into the set to be assigned; S32. Select any empty disk page of the target disk as the target disk page, and select the nodes at both ends of the edge with the largest target weight from the set to be allocated as target nodes and store them in the target disk page; S33. Based on the target disk layout optimization function, calculate the total contribution value of all the edges between the target node and the corresponding out-neighbor nodes of the target node, and store the corresponding out-neighbor nodes of the target node into the target disk page in order of the total contribution value from high to low, until the target disk page is full or all the corresponding out-neighbor nodes of the target node have been stored into the target disk page. S34. For the edges that have not been selected in the set to be allocated, execute the steps of S32 until all target disk pages or all edges in the set to be allocated are selected.

7. The monotonic path-aware vector graph index disk layout optimization method according to claim 6, characterized in that, After all the target disk pages or all the edges in the set to be allocated have been selected, the process further includes: When there are unallocated nodes and incomplete target disk pages, the unallocated nodes are sequentially stored into the incomplete target disk pages.

8. The monotonic path-aware vector graph index disk layout optimization method according to claim 6, characterized in that, The step of allocating each node in the target weighted undirected graph to each disk page of the target disk includes: The nodes in the weighted undirected graph of the target are divided into multiple target clusters according to the distance between the nodes using a clustering algorithm, and the induced subgraphs corresponding to each target cluster are obtained. For each node and edge in each of the induced subgraphs, perform step S31 until each node corresponding to each target cluster is stored in the target disk; The full disk pages in the target disk are merged, and the stored nodes in the non-full disk pages are combined with the unallocated nodes to form a target cluster, and the step S31 is performed on the target cluster.

9. A monotonic path-aware vector graph index disk layout optimization system, used to implement the method as described in any one of claims 1-8, characterized in that, include: The index graph construction module is used to construct a vector graph index based on the nearest neighbor relationships of the data vectors to be indexed. The weight calculation module is used to obtain the target weight of the edge corresponding to each node; The index graph acquisition module is used to acquire a target weighted undirected graph based on the vector graph index and the target weight of each edge; The disk layout optimization module is used to establish a target disk layout optimization function, and based on the target disk layout optimization function and the target weighted undirected graph, to allocate each node in the target weighted undirected graph to each disk page of the target disk.

10. An electronic device, characterized in that, include: At least one memory for storing computer programs; At least one processor is configured to execute a program stored in the memory, wherein when the program stored in the memory is executed, the processor is configured to perform the method as described in any one of claims 1-8.