Vector graph index processing method and apparatus
By scanning and deleting the incoming node and its edges of the logically deleted node in the vector graph index, the real deletion of the vector graph index is achieved, solving the stability and efficiency problems caused by logical deletion, and improving the retrieval performance and recall rate.
Patent Information
- Application Number
- PCT/CN2024/134616
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-28
- Filing Date
- 2024-11-26
- Publication Date
- 2025-06-05
AI Technical Summary
In the prior art, when the vector graph index performs logical deletion node operation, the logical deletion node and its related edges still exist, resulting in poor stability of the vector graph index, reducing the efficiency and accuracy of vector retrieval, and occupying memory resources, requiring periodic full construction.
By scanning the vector graph index, the entry information of the logical delete node is obtained, and the entry node and its edges of the node are deleted based on this information, real deletion is achieved, memory resources are released, and the efficiency and accuracy of vector retrieval are improved.
It improves the stability of vector graph index, enhances the efficiency and accuracy of vector retrieval, ensures the adequacy of recall results, and reduces user operation and maintenance costs.
Smart Images

Figure CN2024134616_05062025_PF_FP_ABST
Abstract
Description
Vector image index processing method and device
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311615682.2 and application name “Method and device for processing vector image index”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of data retrieval technology, and in particular to a method and device for processing vector graph indexes. Background Art
[0003] Feature vectors can be extracted from massive amounts of unstructured data, such as images, text, video, and voice. This allows for analysis and retrieval of unstructured data through calculation and retrieval. In vertical search scenarios like large language models (LLMs) and news recommendations, large amounts of real-time data are continuously generated, and streaming vector data is constantly updated. Vector graph indexing, a vector index based on graph data structures for vector retrieval, offers the advantages of continuous updates, high precision, and high performance, fully meeting business needs.
[0004] In vector graph indexing, operations such as adding, deleting, and modifying nodes are performed based on real-time data. Regarding node deletion, related technologies use a flag bit for each node to indicate whether the node has been logically deleted. Setting this flag bit allows for logical deletion of a node. Edges connected to a logically deleted node (hereinafter referred to as a logically deleted node) still exist. Edges associated with the logically deleted node will be traversed during vector searches, but will not be included in the search result set.
[0005] However, when logically deleting nodes, the deleted nodes and their associated edges still exist in the graph. These holes lead to poor stability in the vector graph index, which in turn reduces the efficiency and accuracy of vector search and may result in insufficient search recall results. Furthermore, logically deleted nodes and their associated edges still occupy memory resources, so related technologies require periodic full rebuilds of the vector graph index, which increases the complexity of vector search. Summary of the Invention
[0006] The present application provides a method and device for processing a vector graph index, which solves the problem in related technologies that logically deleted nodes and related edges still exist in the graph, resulting in poor stability of the vector graph index. It can release memory resources occupied by the related edges of the logically deleted nodes, improve the efficiency and accuracy of vector retrieval, and ensure the adequacy of recall results.
[0007] In the first aspect, the present application provides a method for processing a vector graph index, which includes: scanning the vector graph index based on the number of nodes and the identification of the logically deleted node to obtain the in-degree information of the logically deleted node, the in-degree information indicating the in-degree node of the logically deleted node; based on the in-degree information of the logically deleted node, deleting the in-degree node of the logically deleted node and the edge of the logically deleted node.
[0008] Related technologies also offer a general Lambda architecture, Milvus. In this architecture, deleted nodes are first written to a cache. When the amount of data in the cache exceeds a specified threshold, the data in the cache is converted to read-only data, and the deleted node's ID is written to the disk in a delete (Del) file within the corresponding segment (seg) file. Only after the write is complete does the deleted node become invisible to users, effectively deleting the node.
[0009] Milvus periodically generates multiple seg files, which store the original vector information. Internal Del files mark whether each node has been deleted. Del files are loaded into a memory bitmap, and during vector retrieval, the deleted bits in the bitmap are used for simultaneous retrieval and filtering. For multiple seg files, users need to manually merge them using the compaction protocol and rebuild the vector graph index during the merge.
[0010] However, related technologies still rely on logical deletion, not actual node deletion. This also leads to poor stability of the vector graph index and memory usage. Furthermore, the vector graph index needs to be rebuilt during the seg file merge operation, which consumes a lot of processor power, leading to a complex multi-index state machine and poor search performance.
[0011] In an embodiment of the present application, the graph vector engine achieves high-performance true deletion through a scanning process when the adjacency matrix does not store in-degree information, converts logically deleted nodes into unreachable isolated nodes, and the nodes and related edges after true deletion do not occupy memory resources, thereby improving the stability of the vector graph index, thereby improving the efficiency and accuracy of vector retrieval and ensuring the adequacy of the recall results.
[0012] In one possible implementation, the method further includes: determining the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, the first in-degree node being any in-degree node of the logically deleted node, the out-degree node candidate set including at least one of the following: the updated neighbor node of the first in-degree node, the identifier of the second-order neighbor node of the target node in the updated neighbor node of the first in-degree node, and a target set; wherein, the updated neighbor node does not include the logically deleted node, the target node includes the n nodes in the updated neighbor node that are closest to the first in-degree node, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.
[0013] Its beneficial effect is to adaptively adjust the nodes and edges related to the deleted node by constructing an out-degree node candidate set, optimize the out-degree of the in-degree node of the deleted node and the in-degree of the out-degree node of the deleted node, thereby improving the efficiency and accuracy of vector retrieval and the index recall rate, while ensuring that the approximate nearest neighbor search (ANN) neighbor relationship of the points and edges related to the logically deleted node is still globally optimal.
[0014] In a possible implementation, the out-degree node candidate set includes a target set, and the method further includes: increasing the weight of the nodes in the target set in the out-degree node candidate set.
[0015] The beneficial effect is that the nodes in the neighbor list of the logically deleted node are more likely to be selected as neighbor nodes of the first in-degree node, that is, the in-degree is more likely to be increased by 1, thereby optimizing the structure of the vector graph index.
[0016] In a possible implementation, the out-degree node candidate set includes a target set, the optimal neighbor node of the first in-degree node includes the first node in the target set, and the method further includes: updating the target set by deleting the first node in the target set.
[0017] The beneficial effect is that it can ensure as much as possible that each node in the neighbor list of the logically deleted node is selected by other nodes as a neighbor node, thereby increasing the in-degree of each node in the neighbor list of the logically deleted node by 1 as much as possible.
[0018] In a possible implementation, after determining the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, the method further includes: performing a reverse edge connection on the first in-degree node.
[0019] The beneficial effect is to further optimize the structure of the local vector graph index by considering whether the first in-degree node can be used as a neighbor node of the optimal neighbor node.
[0020] In one possible implementation, the process of obtaining the in-degree information of the logically deleted node based on the number of nodes and the identification scanning vector graph index of the logically deleted node includes: obtaining the in-degree information of the logically deleted node based on the number of nodes and the identification scanning vector graph index of the logically deleted node through an asynchronous thread; the process of deleting the in-degree nodes of the logically deleted node and the edges of the logically deleted node based on the in-degree information of the logically deleted node includes: deleting the in-degree nodes of the logically deleted node and the edges of the logically deleted node based on the in-degree information of the logically deleted node through a main thread.
[0021] The beneficial effect is that the process of scanning for in-degree information is executed by an asynchronous thread, while the process of in-degree cleanup is performed by the main thread. This eliminates the need to add additional in-degree information to the vector graph index, effectively obtaining in-degree information while reducing the complexity and memory usage of the vector graph index. The main thread is used to perform processes such as adding, querying, modifying, and deleting vectors in the vector graph index, as well as vector retrieval. The scanning process can be carried out in parallel with the main thread without interfering with each other, obtaining the in-degree information of each logically deleted node without affecting the main thread. This ensures that memory does not expand during long-term additions, queries, modifications, and deletions. This allows for the creation of a one-stop, streaming, and out-of-the-box vector retrieval engine, improving the timeliness and accuracy of the recall system, eliminating the complex periodic index rebuild process, and reducing user operation and maintenance costs.
[0022] In one possible implementation, the process of scanning the vector graph index based on the number of nodes and the identification of logically deleted nodes through an asynchronous thread includes: when the second node is not locked, scanning the neighbor nodes of the second node through an asynchronous thread, and the second node is any node to be scanned.
[0023] In one possible implementation, the process of deleting the in-degree node of a logically deleted node and the edge of the logically deleted node based on the in-degree information of the logically deleted node through the main thread includes: when the second in-degree node is not locked, deleting the edge between the second in-degree node and the logically deleted node through the main thread, and the second in-degree node is any in-degree node of the logically deleted node.
[0024] In one possible implementation, the identifier of the logically deleted node is stored in a logically deleted pool. The method further includes: when the logically deleted pool meets the update condition, obtaining the number of nodes and the logically deleted pool in the vector graph index. The update condition includes: the data size stored in the logically deleted pool reaches a first threshold or the usage ratio of the logically deleted pool reaches a second threshold.
[0025] In a possible implementation, the method further includes: in response to a received deletion request of the third node, logically deleting the third node and adding an identifier of the third node to a logical deletion pool.
[0026] In Milvus, deletions are only visible after being written to disk, which can take a long time to take effect and potentially lead to inconsistent data. However, in this embodiment, received node deletion requests are directly converted into logical deletions. Logical deletions do not involve any adjustments to the nodes or edges in the vector graph index. The user interface returns immediately, and the deletion takes effect within seconds. Compared to related technologies, this reduces the time it takes for deletions to take effect.
[0027] In the second aspect, the present application provides a processing device for a vector graph index, which includes: a scanning module for scanning the vector graph index based on the number of nodes and the identification of the logically deleted node, and obtaining the in-degree information of the logically deleted node, wherein the in-degree information indicates the in-degree node of the logically deleted node; a cleaning module for deleting the in-degree node of the logically deleted node and the edge of the logically deleted node based on the in-degree information of the logically deleted node.
[0028] In one possible implementation, the device also includes: a determination module, used to determine the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, the first in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: the updated neighbor node of the first in-degree node, the identifier of the second-order neighbor node of the target node in the updated neighbor node of the first in-degree node, and the target set; wherein, the updated neighbor node does not include the logically deleted node, the target node includes the n nodes in the updated neighbor node that are closest to the first in-degree node, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.
[0029] In a possible implementation, the out-degree node candidate set includes a target set, and the determination module is further configured to increase the weight of nodes in the target set in the out-degree node candidate set.
[0030] In a possible implementation, the out-degree node candidate set includes a target set, the optimal neighbor node of the first in-degree node includes the first node in the target set, and the determination module is further configured to update the target set by deleting the first node in the target set.
[0031] In a possible implementation, the determination module is further configured to perform a reverse edge connection on the first in-degree node after determining the best neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node.
[0032] In one possible implementation, the scanning module is specifically used to scan the vector graph index based on the number of nodes and the identification of the logically deleted node through an asynchronous thread to obtain the in-degree information of the logically deleted node; the cleaning module is specifically used to delete the in-degree nodes of the logically deleted node and the edges of the logically deleted node based on the in-degree information of the logically deleted node through a main thread.
[0033] In a possible implementation, the scanning module is specifically configured to scan neighbor nodes of the second node through an asynchronous thread when the second node is not locked, where the second node is any node to be scanned.
[0034] In a possible implementation, the cleaning module is specifically configured to delete the edge between the second in-degree node and the logically deleted node through the main thread when the second in-degree node is not locked, where the second in-degree node is any in-degree node of the logically deleted node.
[0035] In one possible implementation, the identifier of the logical deletion node is stored in the logical deletion pool, and the determination module is specifically used to obtain the number of nodes and the logical deletion pool in the vector graph index when the logical deletion pool meets the update conditions. The update conditions include: the data size stored in the logical deletion pool reaches a first threshold or the usage ratio of the logical deletion pool reaches a second threshold.
[0036] In a possible implementation, the device further includes: a logical deletion module, configured to, in response to a received deletion request of the third node, logically delete the third node and add an identifier of the third node to a logical deletion pool.
[0037] In a third aspect, the present application provides a vector graph index processing device, which includes: one or more processors; a memory for storing one or more computer programs or instructions; when the one or more computer programs or instructions are executed by one or more processors, the one or more processors implement a method as described in any one of the first aspects.
[0038] In a fourth aspect, the present application provides a vector graph index processing device, comprising a processor for executing the method as described in any one of the first aspects.
[0039] In a fifth aspect, the present application provides a vector graph index processing device, which includes: a processing circuit and an interface circuit; wherein the interface circuit is used to couple with a memory outside the vector graph index processing device and provide a communication interface for the processing circuit to access the memory; the processing circuit is used to execute program instructions in the memory to implement a method as described in any one of the first aspects.
[0040] In a specific implementation, the vector graph index processing device may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be a transistor, a gate circuit, a trigger, and various logic circuits. The input signal received by the input circuit may be, for example, but not limited to, received and input by a receiver, and the signal output by the output circuit may be, for example, but not limited to, output to a transmitter and transmitted by the transmitter, and the input circuit and the output circuit may be the same circuit, which is used as an input circuit and an output circuit at different times. The embodiments of the present application do not limit the specific implementation of the processor and various circuits.
[0041] In one implementation, the processing device of the vector map index can be a wireless communication device, that is, a computer device that supports wireless communication functions. Specifically, the wireless communication device can be a terminal such as a smart phone, or a wireless access network device such as a base station. The network chip can also be called a system on chip (SoC), or simply a SoC chip. The communication chip may include a baseband processing chip and a radio frequency processing chip. The baseband processing chip is sometimes also called a modem or baseband chip. The radio frequency processing chip is sometimes also called a radio frequency transceiver or radio frequency chip. In physical implementation, some or all chips in the communication chip can be integrated inside the SoC chip. For example, the baseband processing chip is integrated in the SoC chip, and the radio frequency processing chip is not integrated with the SoC chip. The interface circuit can be the radio frequency processing chip in the wireless communication device, and the processing circuit can be the baseband processing chip in the wireless communication device.
[0042] In another implementation, the vector map index processing device may be a component within a wireless communication device, such as an integrated circuit product such as a network chip or a communication chip. The interface circuit may be an input / output interface, interface circuit, output circuit, input circuit, pin, or related circuit on the chip or chip network. The processor may also be embodied as a processing circuit or a logic circuit.
[0043] In a sixth aspect, the present application provides a computer-readable storage medium, in which program code is stored. When the program code is executed by a processor, the method as described in any one of the first aspects is implemented.
[0044] In a seventh aspect, the present application provides a chip, comprising: at least one processor. The at least one processor is configured to execute the method according to any one of the first aspects.
[0045] Optionally, the chip further includes a memory, and at least one processor is configured to execute code in the memory. When the at least one processor executes the code, the chip implements the method as described in any one of the first aspects.
[0046] Optionally, the chip may also be an integrated circuit.
[0047] In an eighth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to implement the method as described in any one of the first aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] FIG1 is a schematic diagram of an LLM scenario provided in an embodiment of the present application;
[0049] FIG2 is a flow chart of a method for processing a vector graph index according to an embodiment of the present application;
[0050] FIG3 is a schematic diagram of a data structure of a vector graph index provided in an embodiment of the present application;
[0051] FIG4 is a schematic diagram of an in-degree cleanup process for a logically deleted node provided in an embodiment of the present application;
[0052] FIG5 is a flowchart of a method for processing a vector graph index provided by an embodiment of the present application;
[0053] FIG6 is a flow chart of another method for processing vector graph indexes provided in an embodiment of the present application;
[0054] FIG7 is a schematic diagram of an optimization process of a vector graph index provided in an embodiment of the present application;
[0055] FIG8 is a schematic diagram showing the effect of a method for processing a vector graph index provided by an embodiment of the present application;
[0056] FIG9 is a block diagram of a vector image index processing device provided by an embodiment of the present application;
[0057] FIG10 is a block diagram of another apparatus for processing vector graph indexes provided by an embodiment of the present application;
[0058] FIG11 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;
[0059] FIG12 is a schematic structural diagram of a vector graph index processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.
[0061] The terms "first," "second," and the like in the description, embodiments, claims, and drawings of this application are used solely for descriptive purposes and are not to be construed as indicating or implying relative importance or order. Furthermore, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions, such as, for example, inclusion of a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0062] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0063] Vector search is increasingly used in business scenarios such as recommendation systems, image recognition, natural language processing, voiceprint matching, audio retrieval, and file deduplication. Vector search involves retrieving K vectors from a given vector dataset that are similar to the query vector using a metric (e.g., L2, inner product, etc.). The essence of vector search is the ANN algorithm, which is primarily implemented based on four data structures: tree, graph, quantization, and hashing.
[0064] Graph-based vector indexing (called vector graph indexing) enables online construction of vector indexes from 0 to 1 without offline model training. Vector graph index neighbor retrieval uses a global, directional, greedy traversal starting from a fixed node to determine the nearest neighbors of a query request in real-time insertion scenarios. It supports dynamic node addition.
[0065] The embodiment of the present application is about processing vector graph indexes, so the following first explains the various terms used in the embodiment of the present application:
[0066] Vector graph index: uses the structure of a graph adjacency matrix to describe the neighbor relationship between nodes, while saving the original vector of each node for distance calculation;
[0067] Adjacency matrix: A two-dimensional array storage structure of a graph. One dimension stores the node (also called vertex) information in the graph, and the other dimension stores the edge information in the graph. The edge information is also the neighbor list of each node. The neighbor list is used to store the identifiers of the node's neighbor nodes (out-degree neighbor nodes);
[0068] Out-degree: The number of arcs starting from a node with a node as the arc tail is called the out-degree of the node;
[0069] In-degree: The number of arcs that end at a node with a node as the arc head is called the in-degree of the node.
[0070] The embodiments of this application can be applied to vertical search business scenarios such as LLM and news recommendation. Taking LLM as an example, due to factors such as the training cycle cost of LLM, LLM features are static. The graph vector engine can serve as a search enhancement plug-in for LLM, assisting LLM in providing dynamic knowledge supplementation capabilities and supporting vector streaming to continuously serve arbitrary additions, deletions, modifications and queries.
[0071] For example, please refer to Figure 1, which is a schematic diagram of an LLM scenario provided in an embodiment of the present application. Figure 1 shows three business scenarios when an LLM-based system (i.e., LLM), a feature analysis module, and a graph vector engine are applied to LLM. The three business scenarios include: 1. The graph vector engine serves as the generative pre-trained transformer cache (GPTCache) of the LLM; 2. The graph vector engine serves as the personal history memory / private domain knowledge base of the LLM; 3. The graph vector engine serves as the vertical domain knowledge base (knowledge base) of the LLM. Personal history memory / private domain knowledge base is a long-term memory.
[0072] As shown in Figure 1, for GPTCache, when a query is frequently requested, it can be inserted into the vector graph index online. This allows for neighbor retrieval using the vector graph index without requiring the LLM to perform a search. This improves the LLM's inference response speed and reduces its inference cost.
[0073] For specific domains such as personal historical memories, private domain knowledge bases, and vertical domain knowledge bases, LLMs have low descriptive accuracy. Therefore, vector graph indexing can be used to independently index knowledge in specific domains and provide accurate ANNs. As shown in Figure 1, the feature analysis module encodes the user's streaming unstructured data into vector data, and the graph vector engine constructs the vector data index online, enabling plug-and-check visibility at the second level. The user's streaming unstructured data can include questions and answers (QAs) and documents (Docs).
[0074] The present application provides a method for processing vector graph indexes. This method can be applied to a graph vector engine (e.g., a Poisson vector engine service) to implement true deletion of vector graph indexes. Please refer to Figure 2, which is a flow chart of a method for processing vector graph indexes provided by the present application. Vector graph indexes can support online streaming add, query, modify, and delete operations. Add, query, modify, and delete operations refer to add / query / update / delete (CRUD). The method may include the following steps:
[0075] 101. Scan the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain in-degree information of the logically deleted node, where the in-degree information indicates the in-degree node of the logically deleted node.
[0076] A logically deleted node refers to a node that has been logically deleted. Upon receiving a delete request for a first node, the graph vector engine logically deletes the first node in response to the delete request. The first node refers to any node to be deleted, and this process can be performed upon receiving a delete request for any node.
[0077] The graph vector engine converts received node deletion requests directly into logical deletions. Logical deletions do not affect any node or edge adjustments in the graph vector index. The user interface returns immediately, achieving a high transaction per second (TPS) rate, making deleted nodes invisible to users within seconds. This process can be executed by the main thread of the graph vector engine.
[0078] For example, a deletion flag can be set for each node in the data structure of the vector graph index. In response to a deletion request from the first node, the main thread sets the deletion flag of the first node to "logically deleted." For example, the deletion flag can be a Del bit, where a Del bit of 0x00 indicates that the node is normal, 0x01 indicates that the node is logically deleted, and 0x11 indicates that the node is actually deleted. On this basis, an is_delete function is also provided. The is_delete function can be used to determine whether the Del bit of a node is 0x01 to determine whether the node has been logically deleted.
[0079] In the embodiments of the present application, an algorithmic design is implemented to make logically deleted nodes invisible. When inserting new nodes, logically deleted nodes are not considered, and their in-degree remains unchanged. When a search process traverses to a logically deleted node, the edges associated with the logically deleted node are accessible, but the final search results do not include the logically deleted node.
[0080] Optionally, the graph vector engine may be pre-set with a logical deletion pool, and after logically deleting the first node, the identity (ID) of the first node is added to the logical deletion pool. The logical deletion pool may be a std::vector <int32>logic_delete_pool, std means (standard), vector means (vector), int32 means (32-bit signed integer), and logic_delete_pool means (logical delete pool).
[0081] The graph vector engine can periodically monitor the storage status of the logical deletion pool. When the logical deletion pool meets the update conditions, it starts to convert the logical deletion node to the real deletion node. First, the number of nodes and the logical deletion pool in the vector graph index at the current moment are obtained. Among them, the update conditions include but are not limited to: the size of the data stored in the logical deletion pool reaches a first threshold or the usage ratio of the logical deletion pool reaches a second threshold. The update conditions can be customized, for example, it can also be a preset time interval since the last execution of the process, which is not limited in the embodiment of the present application. The process can be executed by an asynchronous thread in the graph vector engine, that is, the process is executed in the background.
[0082] In-degree information can include the identifier of the in-degree node of the logically deleted node. The in-degree node of the logically deleted node refers to the node connected to the other end of the in-edge of the logically deleted node. The graph vector engine can scan the vector graph index based on the number of nodes and the entire logically deleted pool. The number of nodes is used to limit the scan range. The graph vector engine scans this number of nodes, effectively avoiding scanning nodes added after the scan starts.
[0083] As can be seen from the preceding description, the graph vector index's data structure records each node's neighbor list. A node's neighbor list is used to record the identifiers of its out-degree nodes. A node's out-degree node refers to the node to which the other end of its outgoing edge connects. The graph vector engine scans each node's neighbor list. If a node's neighbor list contains the identifier of a logically deleted node, the node's identifier is added to the logically deleted node's in-degree information.
[0084] This process can be executed by an asynchronous thread, eliminating the need to add in-degree information to the vector graph index, reducing its complexity and memory usage. The main thread is used to perform CRUD and vector retrieval operations on the vector graph index. The scanning process can run in parallel with the main thread without interfering with each other, allowing the in-degree information of each logically deleted node to be obtained without affecting the main thread.
[0085] In an embodiment of the present application, an independent lock can be set for each node in the data structure of the vector graph index to ensure the thread safety of the node neighbor list, effectively avoid mutual interference between the main thread and the asynchronous thread, and achieve high-performance parallel scanning.
[0086] When the main thread is manipulating a node's neighbor list (e.g., CRUD or retrieval), it locks the node. When the asynchronous thread is scanning the vector graph index, it first checks whether the node's lock is held before scanning the node's neighbor list. If the lock is not held, the neighbor list is scanned. If the lock is held, the scan is blocked. The blocking period is typically 1 millisecond.
[0087] Similarly, while scanning a node's neighbor list, the asynchronous thread also locks the node. Before operating on a node's neighbor list, the main thread first checks whether the node's lock is occupied. If the node's lock is not occupied, the operation is performed. If the node's lock is occupied, the current operation is blocked.
[0088] For example, please refer to Figure 3, which is a schematic diagram of the data structure of a vector graph index provided in an embodiment of the present application. Figure 3 further illustrates the parallel process of the aforementioned main thread and asynchronous thread using four nodes A, B, O, and C as examples. Figure 3 shows the lock bit, Del bit, ID, and neighbor list of four nodes, where the identifiers of these four nodes are A, B, O, and C, respectively. It should be noted that the identifiers of the nodes in the neighbor list shown in Figure 3 are for illustrative purposes only and do not constitute any limitation.
[0089] Assuming that the main thread receives a delete instruction for the O node at time T1, the Del position of the O node is set to 0x01, marking the O node as a logically deleted node, and the ID of the O node is added to the logic_delete_pool. At this time, the O node is invisible to both the user side and the algorithm layer. If the main thread adds the C node to the vector graph index after time T1, the main thread traverses each node in the vector graph index to globally search for M neighboring nodes. When traversing each node, the main thread first determines whether the node is locked through the lock bit. When the node is not locked, it determines whether the node is logically deleted through the is_delete function. When a node is logically deleted, the node will not be added to the neighbor list of the C node. Therefore, the in-degree of the O node remains unchanged after time T1.
[0090] To ensure accurate in-degree information obtained through scanning, the asynchronous thread first checks whether each node is locked before scanning its neighbor list in parallel. If a node is locked, the scan is blocked. If a reverse edge is applied to node C during the addition process, causing a change in node A's neighbor list, the main thread locks node A. If the asynchronous thread begins scanning node A's neighbor list and the node is still locked, the scan is blocked.
[0091] As shown in Figure 3, the asynchronous thread scans the identifier of node O in the neighbor lists of nodes A and B. Therefore, nodes A and B are in-degree nodes of node O. The in-degree information of node O includes: A and B. In addition, the neighbor list of the newly added node C does not include O.
[0092] When the amount of data in the vector graph index is large, if the TPS of the main thread's CRUD or retrieval operations is normal, the probability of collision between the main thread and the asynchronous thread is small, which can basically ensure that the main thread and the asynchronous thread do not interfere with each other.
[0093] For example, the identifier of the logical delete node whose in-degree information has been determined may be deleted from the logical delete pool, that is, the logical delete pool may be continuously updated.
[0094] 102. Based on the in-degree information of the logically deleted node, delete the in-degree nodes of the logically deleted node and the edges of the logically deleted node.
[0095] Optionally, after determining the in-degree information of all logically deleted nodes, for each logically deleted node, each edge between the in-degree node of the logically deleted node and the logically deleted node may be deleted. Alternatively, each time the in-degree information of a logically deleted node is determined, each edge between the in-degree node of the logically deleted node and the logically deleted node may be deleted. This embodiment of the present application is not limited to this.
[0096] As can be seen from the preceding description, a node's neighbor list records the identities of its out-degree nodes. That is, the neighbor list is used to store the edges associated with each node in the vector graph index. The graph vector engine can delete the edges between a tombstone node's in-degree nodes and the tombstone node by removing the tombstone node's identifier from its in-degree neighbor list.
[0097] By executing process 102, the logically deleted node is converted into an unreachable isolated node, achieving true node deletion. The memory / disk area occupied by the node can be reused. This solves the problems of low retrieval efficiency and accuracy and insufficient recall results caused by the invalid edges of the logically deleted node.
[0098] This process can be executed by the main thread. After the asynchronous thread determines the in-degree information of the logically deleted node, the main thread first determines whether the first in-degree node of the logically deleted node is locked before deleting the edge between the first in-degree node and the logically deleted node. If the first in-degree node is not locked, the first in-degree node is locked, and then the edge between the first in-degree node and the logically deleted node is deleted, thereby quickly completing the in-degree cleanup of the logically deleted node. When the first in-degree node is locked, the in-degree cleanup process is blocked. The first in-degree node is any in-degree node of the logically deleted node.
[0099] In the embodiment of the present application, since the asynchronous thread has already determined the in-degree information of the logically deleted node, the main thread only needs to directly read the in-degree information of the logically deleted node determined by the asynchronous thread, lock the in-degree node of the logically deleted node, and delete the edge between the in-degree node of the logically deleted node and the logically deleted node. The complexity of clearing the in-degree of the logically deleted node is O1, which shortens the locking time and improves the deletion efficiency. Compared with related technologies, it improves the overall queries per second (QPS) performance of CRUD.
[0100] For example, please refer to Figure 4, which is a schematic diagram of an in-degree cleanup process for logically deleting a node provided in an embodiment of the present application. Figure 4 uses the in-degree scan results shown in Figure 3 as an example to illustrate the in-degree cleanup process, and uses the neighbor list as an example to illustrate the in-degree cleanup process. Figure 4 includes two parts (a) and (b). (a) shows the vector graph before in-degree cleanup and the neighbor lists of nodes A, B, and O. (b) shows the vector graph after in-degree cleanup and the neighbor lists of nodes A, B, and O.
[0101] As shown in Figure 4(a), the neighbor lists of nodes A and B include node O. After indegree cleaning, as shown in Figure 4(b), node O is deleted from node A's neighbor list and node O is deleted from node B's neighbor list. Simultaneously, in the vector graph, the edge from node A to node O is deleted, and the edge from node B to node O is deleted, making node O unreachable.
[0102] For example, you can also add a real delete pool on the vector map index data structure std::vector <int32>real_delete_pool, where real means real. After the logically deleted node's identifier is removed from the neighbor list of each in-degree node of the logically deleted node, the logically deleted node becomes a real deleted node. At this point, the node's identifier can be added to the real deletion pool, and the node's Del bit is set to 0x11.
[0103] The following describes the flow of the aforementioned processes 101 to 102 with the main thread and the asynchronous thread as the execution subjects. Please refer to Figure 5, which is a flowchart of a method for processing a vector graph index provided in an embodiment of the present application. As shown in ①, the main thread responds to a deletion request received for a certain node, adds the ID of the node to the logic_delete_pool, and sets the Del bit of the node to 0x01. The main thread will then feedback a deletion response, which indicates that the logical deletion is successful or the logical deletion fails. As shown in ②, the logically deleted node is invisible when a new node is inserted, and the logically deleted node is filtered out when a retrieval is performed.
[0104] As shown in ③, the asynchronous thread monitors the state of the logical delete pool to determine whether the threshold has been reached (i.e., whether the update condition has been met). The measurement elements for whether the threshold has been reached include: data size, ratio, or time. For data size, it is determined whether the size of the data stored in the logical delete pool has reached the first threshold. For ratio, it is determined whether the usage ratio of the logical delete pool has reached the second threshold. For time, it is determined whether the time interval with the most recent update time has reached the third threshold. As shown in ④, based on the number of vector index points at the time of the snapshot scan indegree and the logical delete pool, an indegree scan is performed to obtain the indegree map of the logical delete nodes. This process is parallel with the processes of adding nodes, searching, and logical deletion (parallel with [add][search][logic delete]).
[0105] As shown in step ⑤, after the main thread performs indegree cleaning, it adds the node ID to the real_delete_pool, and the memory occupied by the node and related edges becomes reusable.
[0106] To summarize, the method for processing a vector graph index provided in an embodiment of the present application first scans the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain the in-degree information of the logically deleted node, where the in-degree information indicates the in-degree node of the logically deleted node. Then, based on the in-degree information of the logically deleted node, the in-degree node of the logically deleted node and the edges of the logically deleted node are deleted. The graph vector engine directly converts the received node deletion request into a logical deletion. The logical deletion does not involve the adjustment of any nodes and edges in the vector graph index. The user interface returns immediately, and the deletion can take effect in seconds. Through the scanning process, high-performance real deletion is achieved when the adjacency matrix does not store the in-degree information. The logically deleted node is converted into an unreachable isolated node. The nodes and related edges after real deletion will not occupy memory resources, thereby improving the stability of the vector graph index, thereby improving the efficiency and accuracy of vector retrieval and ensuring the adequacy of the recall results.
[0107] Furthermore, in embodiments of the present application, the process of scanning for in-degree information can be executed by an asynchronous thread, while the process of in-degree cleanup can be executed by the main thread. This eliminates the need to add additional in-degree information to the vector graph index, effectively obtaining in-degree information while reducing the complexity and memory usage of the vector graph index. The main thread is used to execute processes such as CRUD and vector retrieval of the vector graph index. The scanning process can run in parallel with the main thread without interfering with each other, obtaining the in-degree information of each logically deleted node without affecting the main thread. This ensures that memory does not expand during long-term CRUD operations, enabling the creation of a one-stop, streaming, uninterrupted, and ready-to-use vector retrieval engine. This improves the timeliness and accuracy of the recall system, eliminates the complex periodic index rebuild process, and reduces user operation and maintenance costs.
[0108] When the embodiment of the present application is applied to an LLM scenario (such as the scenario shown in FIG1 ), it can not only construct vector data online through real-time vector graph indexing, achieving plug-and-check visibility in seconds for users, but also can actually delete or update data points when there is expired or to-be-updated vector data in the vector graph index, and can continuously receive user requests for addition, deletion, and modification. The vector graph index memory will not expand, and the retrieval accuracy and retrieval performance will not be reduced, and users do not need to periodically perform full offline construction.
[0109] Based on the above embodiment, the structure of the vector graph index can be further optimized. Please refer to Figure 6, which is a flow chart of another method for processing vector graph indexes provided by an embodiment of the present application. This method can be applied to a graph vector engine (e.g., a Poisson vector engine service). This method may include the following steps:
[0110] 201. Scan the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain in-degree information of the logically deleted node, where the in-degree information indicates the in-degree node of the logically deleted node.
[0111] This process can refer to the aforementioned process 101, and will not be described in detail in this embodiment of the present application.
[0112] 202. Based on the in-degree information of the logically deleted node, delete the in-degree nodes of the logically deleted node and the edges of the logically deleted node.
[0113] This process can refer to the aforementioned process 102, and will not be described in detail in this embodiment of the present application.
[0114] Through the aforementioned process 201 to process 202, the out-degree of the in-degree node of the logically deleted node is reduced by 1. Since the logically deleted node is unreachable, the in-degree of the out-degree node of the logically deleted node (that is, the node in the neighbor list of the logically deleted node) is reduced by 1.
[0115] 203. Determine the optimal neighbor node of the second in-degree node from the out-degree node candidate set of the second in-degree node, where the second in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: an updated neighbor node of the second in-degree node, an identifier of a second-order neighbor node of the target node among the updated neighbor nodes of the second in-degree node, and a target set.
[0116] The updated neighbor nodes do not include the identifier of the logically deleted node, the target nodes include the n nodes closest to the second in-degree node in the updated neighbor nodes, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.
[0117] The updated neighbor nodes of the second in-degree node can be nodes in the neighbor list after the identifier of the logically deleted node is deleted. The second-order neighbors of the target node refer to the neighbor nodes (also called out-degree neighbor nodes) of the neighbor nodes of the target node (also called out-degree neighbor nodes). n is a custom parameter, for example, n can be 2 or 3, etc., and this embodiment of the application does not limit this.
[0118] For example, if the out-degree node candidate set includes the target set, the weight of the nodes in the target set in the out-degree node candidate set can be increased. The weight of a node represents the probability of the node being selected as a neighbor node. Since the in-degree of the neighbor nodes of the logically deleted node is reduced by 1 after the aforementioned processes 201 to 202, this process can make it easier for the nodes in the neighbor list of the logically deleted node to be selected as the neighbor nodes of the second in-degree node, that is, more likely to have their in-degree increased by 1, thereby optimizing the structure of the vector graph index.
[0119] If the weights of the nodes in the target set are increased in the out-degree node candidate set, the process is to determine the optimal neighbor node of the second in-degree node from the out-degree node candidate set after the weights of the nodes in the target set are increased.
[0120] For example, the optimal neighbor node of the second in-degree node can be determined from the out-degree node candidate set through a heuristic edge selection strategy. Heuristic edge selection is an edge selection strategy introduced in the hierarchical navigable small world (HNSW) vector graph index. It means that when a newly added node traverses EfConstruction (a variable in HNSW) nearest neighbor nodes in a global search, and the maximum number of neighbors of the newly added node is M (EfConstruction>M), M neighbor nodes will be optimally selected from the EfConstruction nearest neighbor nodes. In this process, not only the distance size but also the directionality is considered to prevent the edge aggregation retrieval from falling into the local optimum.
[0121] When the out-degree node candidate set includes the target set, and the second in-degree node's optimal neighbor node includes the second node in the target set, the target set can be dynamically adjusted. For example, the target set can be updated by deleting the second node in the target set. This ensures that each node in the logically deleted node's neighbor list is selected as a neighbor node by other nodes, thereby increasing the in-degree of each node in the logically deleted node's neighbor list by 1 as much as possible.
[0122] For example, when the target set is cleared, the target set may be restored, and the nodes in the neighbor list of the logically deleted node may be added back into the target set.
[0123] After determining the optimal neighbor node of the second in-degree node from the out-degree node candidate set, we can also perform a reverse edge connection on the second in-degree node. That is, we consider whether the second in-degree node can be a neighbor node of the optimal neighbor node, thereby further optimizing the structure of the local vector graph index.
[0124] It should be noted that the aforementioned process 203 is described by taking the second in-degree node of a logically deleted node as an example. The relevant process of any in-degree node of any logically deleted node can refer to process 203, and the embodiment of the present application will not be repeated here.
[0125] Please refer to Figure 7, which is a schematic diagram of the optimization process of a vector graph index provided in an embodiment of the present application. Figure 7 uses a neighbor list as an example for illustration. Figure 7 includes two parts (a) and (b), where (a) represents the local vector graph index before optimization, and (b) represents the local vector graph index after optimization. Figure 7 shows multiple nodes and the neighbor lists of node A, node B, and node O. The following explanation uses node O as a logically deleted node and node A as the second in-degree node as an example.
[0126] As shown in Figure 7(a), the neighbor lists of node O's in-degree nodes (e.g., nodes A and B) have vacancies, and their out-degrees are reduced by 1. Node O is unreachable, so the in-degrees of each node in node O's neighbor list (e.g., nodes C and D) are reduced by 1. First, determine the candidate set of out-degree nodes V. For example, the candidate set of out-degree nodes V consists of: node A's updated neighbor list, the identifiers of the target node's second-order neighbor nodes in node A's updated neighbor list, and the target set S.
[0127] Then, a heuristic weighted edge selection strategy is used to determine the optimal M neighbor nodes of node A. First, the weight of each node in set S in set V is increased, and then the heuristic edge selection strategy is applied to the weighted set V to obtain the optimal M neighbor nodes of node A.
[0128] Next, set V is dynamically adjusted and reversed. If a subset Y in set S is selected as the optimal M neighbor nodes of node A, subset Y is removed from set S. If set S is empty, set S is reset to the neighbor list of node O. Reverse edges are then applied to node A. For example, if A is connected to the optimal neighbor node C, it is necessary to determine whether node C should be connected to node A, further optimizing the local graph structure.
[0129] The neighbor node optimization process of node B can refer to node A, and the embodiment of this application will not be repeated here. As shown in Figure 7, the optimal neighbor node of node A includes node C, and the optimal neighbor node of node B includes node D.
[0130] The deletion and self-repair process in the aforementioned vector graph index has a complexity of O(2M), where M represents the maximum number of neighbor nodes for each node, typically 32. The self-update time for truly deleted nodes is short, resulting in high TPS performance and minimal impact on the CRUD process of the vector graph index. Furthermore, the entire process is executed based on the full vector graph index, without involving other temporary index structures, ensuring high reliability.
[0131] Please refer to Figure 8, which is a schematic diagram illustrating the effects of a method for processing a vector graph index provided in an embodiment of the present application. Figure 8 illustrates the changes in the vector graph index after a logically deleted node is converted into a truly deleted node using the method provided in an embodiment of the present application. As can be seen from Figure 8, the vector graph index does not contain logically deleted nodes and their associated edges, and the structure of the vector graph index has been adaptively adjusted, maintaining the globally optimal ANN neighbor relationship for the points and edges associated with the logically deleted nodes.
[0132] In summary, the method for processing a vector graph index provided by an embodiment of the present application first scans the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain the in-degree information of the logically deleted node, the in-degree information indicates the in-degree node of the logically deleted node, and then based on the in-degree information of the logically deleted node, deletes the in-degree node of the logically deleted node and the edge of the logically deleted node, and finally determines the optimal neighbor node of the second in-degree node from the out-degree node candidate set of the second in-degree node, the second in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: the updated neighbor node of the second in-degree node, the identifier of the second-order neighbor node of the target node in the updated neighbor node of the second in-degree node, and the target set, and the graph vector engine directly converts the received node deletion request into Logical deletion does not involve any adjustment of nodes and edges in the vector graph index. The user interface returns immediately and the deletion can take effect in seconds. Through the scanning process, high-performance real deletion is achieved when the adjacency matrix does not store in-degree information. The logically deleted nodes are converted into unreachable isolated nodes. The nodes and related edges after real deletion will not occupy memory resources, which improves the stability of the vector graph index. By constructing an out-degree node candidate set, the related nodes and edges of the deleted nodes are adaptively adjusted, and the out-degree of the in-degree node of the deleted node and the in-degree of the out-degree node of the deleted node are optimized. This improves the efficiency and accuracy of vector retrieval and ensures the adequacy of the recall results. At the same time, it ensures that the ANN neighbor relationship of the points and edges related to the logically deleted nodes is still the global optimal.
[0133] The order of the methods provided in the embodiments of the present application can be adjusted appropriately, and the process can be increased or decreased accordingly. Any method that can be easily thought of by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application, and the embodiments of the present application do not limit this.
[0134] The above mainly introduces the method for processing vector graph indexes provided by the embodiment of the present application from the perspective of the device. It can be understood that in order to realize the above functions, the device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0135] The embodiment of the present application can divide the functional modules of the device according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing subsystem. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0136] Figure 9 is a block diagram of a vector image index processing device provided in an embodiment of the present application. When functional modules are divided according to their functions, the vector image index processing device 300 may include a scanning module 301 and a cleaning module 302. For example, the vector image index processing device may be a graph vector engine, a chip therein, or other combined device or component having the functions of the vector image index processing device described above. The functions of the various modules of the device are as follows:
[0137] The scanning module 301 is used to scan the vector graph index based on the number of nodes and the identification of the logically deleted node to obtain the in-degree information of the logically deleted node, and the in-degree information indicates the in-degree node of the logically deleted node; the cleaning module 302 is used to delete the in-degree node of the logically deleted node and the edge of the logically deleted node based on the in-degree information of the logically deleted node.
[0138] In combination with the above scheme, please refer to Figure 10. Figure 10 is a block diagram of another vector graph index processing device provided in an embodiment of the present application. Based on Figure 9, the device also includes: a determination module 303, used to determine the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, the first in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following items: the updated neighbor node of the first in-degree node, the identifier of the second-order neighbor node of the target node in the updated neighbor node of the first in-degree node, and the target set; wherein, the updated neighbor node does not include the logically deleted node, and the target node includes the n nodes in the updated neighbor node that are closest to the first in-degree node, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.
[0139] In combination with the above solution, the out-degree node candidate set includes the target set, and the determination module 303 is further configured to increase the weight of the nodes in the target set in the out-degree node candidate set.
[0140] In combination with the above solution, the out-degree node candidate set includes the target set, the optimal neighbor node of the first in-degree node includes the first node in the target set, and the determination module 303 is further used to update the target set by deleting the first node in the target set.
[0141] In combination with the above solution, the determination module 303 is further configured to perform a reverse edge connection on the first in-degree node after determining the best neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node.
[0142] In combination with the above scheme, the scanning module 301 is specifically used to scan the vector graph index based on the number of nodes and the identification of the logically deleted node through an asynchronous thread to obtain the in-degree information of the logically deleted node; the cleaning module 302 is specifically used to delete the in-degree nodes of the logically deleted node and the edges of the logically deleted node based on the in-degree information of the logically deleted node through the main thread.
[0143] In combination with the above solution, the scanning module 301 is specifically configured to scan neighboring nodes of the second node through an asynchronous thread when the second node is not locked, where the second node is any node to be scanned.
[0144] In combination with the above solution, the cleaning module 302 is specifically configured to delete the edge between the second in-degree node and the logically deleted node through the main thread when the second in-degree node is not locked. The second in-degree node is any in-degree node of the logically deleted node.
[0145] In combination with the above scheme, the identifier of the logical deletion node is stored in the logical deletion pool, and the determination module 303 is specifically used to obtain the number of nodes and the logical deletion pool in the vector graph index when the logical deletion pool meets the update conditions. The update conditions include: the data size stored in the logical deletion pool reaches a first threshold or the usage ratio of the logical deletion pool reaches a second threshold.
[0146] In combination with the above solution, as shown in FIG10 , the apparatus further includes: a logical deletion module 304 for logically deleting the third node in response to a received deletion request of the third node and adding an identifier of the third node to the logical deletion pool.
[0147] FIG11 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 400 may be a chip or functional module in a graph vector engine. As shown in FIG11 , the electronic device 400 includes a processor 401 , a transceiver 402 , and a communication line 403 .
[0148] The processor 401 is used to execute any step in the method embodiments shown in Figures 2 and 6, and when executing processes such as receiving a node deletion request, it can choose to call the transceiver 402 and the communication line 403 to complete the corresponding operation.
[0149] Furthermore, the electronic device 400 may further include a memory 404 , wherein the processor 401 , the memory 404 and the transceiver 402 may be connected via a communication line 403 .
[0150] Transceiver 402 is used to communicate with other devices or other communication networks, such as Ethernet, radio access networks (RAN), wireless local area networks (WLAN), etc. Transceiver 402 can be a module, circuit, transceiver, or any device capable of implementing communication.
[0151] The transceiver 402 is mainly used for sending and receiving requests, etc., and may include a transmitter and a receiver for sending and receiving requests, etc. respectively; operations other than sending and receiving requests, etc. are implemented by the processor, such as determining the number of nodes in the vector graph index and the logical deletion pool, etc.
[0152] The communication line 403 is used to transmit information between the components included in the electronic device 400.
[0153] In one design, the processor can be considered as the logic circuit and the transceiver as the interface circuit.
[0154] The memory 404 is used to store instructions, where the instructions may be computer programs.
[0155] It should be noted that memory 404 can exist independently of processor 401 or can be integrated with processor 401. Memory 404 can be used to store instructions, program code, or some data. Memory 404 can be located within electronic device 400 or outside of electronic device 400, without limitation. Processor 401 is configured to execute instructions stored in memory 404 to implement the methods provided in the above embodiments of this application.
[0156] In one example, processor 401 may include one or more processors, such as processor 0 and processor 1 in FIG. 11 .
[0157] As an optional implementation, the electronic device 400 includes multiple processors. For example, in addition to the processor 401 in FIG. 11 , it may also include a processor 407 .
[0158] As an optional implementation, the electronic device 400 further includes an output device 405 and an input device 406. For example, the input device 406 is a keyboard, a mouse, a microphone, a joystick, or the like, and the output device 405 is a display screen, a speaker, or the like.
[0159] It should be pointed out that the electronic device 400 can be a chip system or a device with a similar structure as shown in Figure 11. Among them, the chip system can be composed of chips, or it can include chips and other discrete devices. The actions, terms, etc. involved in the various embodiments of this application can refer to each other without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are only an example. Other names can also be used in the specific implementation without limitation. In addition, the component structure shown in Figure 11 does not constitute a limitation on the electronic device 400. In addition to the components shown in Figure 11, the electronic device 400 may include more or fewer components than those shown in Figure 11, or combine certain components, or arrange the components differently.
[0160] The processor and transceiver described in this application can be implemented on an integrated circuit (IC), an analog IC, a radio frequency integrated circuit, a mixed-signal IC, an application specific integrated circuit (ASIC), a printed circuit board (PCB), an electronic device, etc. The processor and transceiver can also be manufactured using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), N-type metal oxide semiconductor (NMOS), P-type metal oxide semiconductor (positive channel metal oxide semiconductor, PMOS), bipolar junction transistor (BJT), bipolar CMOS (BiCMOS), silicon germanium (SiGe), gallium arsenide (GaAs), etc.
[0161] Figure 12 is a structural diagram of a vector image index processing device provided in an embodiment of the present application. The vector image index processing device can be applied to the scenario shown in the above method embodiment. For ease of explanation, Figure 12 only shows the main components of the vector image index processing device, including a processor, a memory, a control circuit, and an input-output device. The processor is mainly used to process communication protocols and communication data, execute software programs, and process software program data. The memory is mainly used to store software programs and data. The control circuit is mainly used for power supply and transmission of various electrical signals. The input-output device is mainly used to receive data input by the user and output data to the user.
[0162] When the processing device of the vector map index is a map vector engine, the control circuit can be a mainboard, the memory includes a hard disk, RAM, ROM and other media with storage functions, the processor can include a baseband processor and a central processing unit, the baseband processor is mainly used to process the communication protocol and communication data, the central processing unit is mainly used to control the entire vector map index processing device, execute software programs, process software program data, and the input and output devices include a display screen, keyboard and mouse, etc.; the control circuit can further include or be connected to a transceiver circuit or transceiver, such as a network cable interface, etc., for sending or receiving data or signals, such as data transmission and communication with other devices. Furthermore, it can also include an antenna for sending and receiving requests, and for data / request transmission with other devices.
[0163] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes computer program code. When the computer program code runs on a computer, it enables the computer to execute any of the methods described in the embodiments of the present application.
[0164] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above-mentioned method embodiments can be completed by a computer or a device with the processing capability of a vector image index to execute a computer program or instruction to control the relevant hardware. The computer program or the group of instructions can be stored in the above-mentioned computer-readable storage medium. When executed, the computer program or the group of instructions may include the processes of the above-mentioned method embodiments. The computer-readable storage medium can be the internal storage unit of the image vector engine of any of the above-mentioned embodiments, such as the hard disk or memory of the image vector engine. The above-mentioned computer-readable storage medium can also be an external storage device of the above-mentioned image vector engine, such as a plug-in hard disk equipped on the above-mentioned image vector engine, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. Further, the above-mentioned computer-readable storage medium can also include both the internal storage unit of the above-mentioned image vector engine and an external storage device. The above-mentioned computer-readable storage medium is used to store the above-mentioned computer program or instruction and other programs and data required by the above-mentioned image vector engine. The above-mentioned computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0165] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0166] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0167] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0168] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0169] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0170] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
[0171] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited to this. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims. The above mainly introduces the vector map index processing method provided by the embodiment of the present application from the perspective of the device. It can be understood that in order to achieve the above functions, the device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
Claims
1. A method for processing a vector graph index, characterized in that: The method comprises: Scan the vector graph index based on the number of nodes and the identification of the logically deleted node to obtain the in-degree information of the logically deleted node, wherein the in-degree information indicates the in-degree node of the logically deleted node; Based on the in-degree information of the logically deleted node, the in-degree nodes of the logically deleted node and the edges of the logically deleted node are deleted.
2. The method according to claim 1, characterized in that The method further comprises: Determine the best neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, the first in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: an updated neighbor node of the first in-degree node, an identifier of a second-order neighbor node of a target node in the updated neighbor nodes of the first in-degree node, and a target set; Among them, the updated neighbor nodes do not include the logically deleted node, the target nodes include n nodes among the updated neighbor nodes that are closest to the first in-degree node, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.
3. The method according to claim 2, characterized in that The out-degree node candidate set includes the target set, and the method further includes: Increase the weight of the nodes in the target set in the out-degree node candidate set.
4. The method according to claim 2 or 3, characterized in that: The out-degree node candidate set includes the target set, the best neighbor node of the first in-degree node includes the first node in the target set, and the method further includes: The target set is updated by deleting the first node in the target set.
5. The method according to any one of claims 2 to 4, characterized in that After determining the best neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, the method further includes: Perform a reverse edge connection on the first in-degree node.
6. The method according to any one of claims 1 to 5, characterized in that The step of scanning the vector graph index based on the number of nodes and the identification of the logically deleted node to obtain the in-degree information of the logically deleted node includes: Scanning the vector graph index based on the number of nodes and the identification of the logically deleted node through an asynchronous thread to obtain the in-degree information of the logically deleted node; The deleting the in-degree node of the logically deleted node and the edge of the logically deleted node based on the in-degree information of the logically deleted node comprises: Based on the in-degree information of the logically deleted node, the in-degree nodes of the logically deleted node and the edges of the logically deleted node are deleted through the main thread.
7. The method according to claim 6, characterized in that Scanning the vector graph index based on the number of nodes and the identification of logically deleted nodes through an asynchronous thread includes: When the second node is not locked, the neighboring nodes of the second node are scanned by the asynchronous thread, and the second node is any node to be scanned.
8. The method according to claim 6 or 7, characterized in that: The deleting the in-degree node of the logically deleted node and the edge of the logically deleted node based on the in-degree information of the logically deleted node through the main thread includes: When the second in-degree node is not locked, the edge between the second in-degree node and the logical deletion node is deleted by the main thread, and the second in-degree node is any in-degree node of the logical deletion node.
9. The method according to any one of claims 1 to 8, characterized in that The identification of the logical deletion node is stored in a logical deletion pool, and the method further comprises: When the logical delete pool meets the update condition, the number of nodes in the vector graph index and the logical delete pool are obtained, and the update condition includes: the size of data stored in the logical delete pool reaches a first threshold or the usage ratio of the logical delete pool reaches a second threshold.
10. The method according to claim 9, characterized in that The method further comprises: In response to the received deletion request of the third node, the third node is logically deleted and an identifier of the third node is added to the logical deletion pool.
11. A vector graph index processing device, characterized in that: The device comprises: A scanning module, configured to scan the vector graph index based on the number of nodes and the identification of the logically deleted node, and obtain the in-degree information of the logically deleted node, wherein the in-degree information indicates the in-degree node of the logically deleted node; A cleaning module is used to delete the in-degree nodes of the logically deleted node and the edges of the logically deleted node based on the in-degree information of the logically deleted node.
12. A vector graph index processing device, characterized in that: The device comprises: one or more processors; a memory for storing one or more computer programs or instructions; When one or more computer programs or instructions are executed by one or more processors, the one or more processors implement the method according to any one of claims 1 to 10.
13. A vector graph index processing device, characterized in that: The device comprises: Processing circuits and interface circuits; Wherein, the interface circuit is used to couple with a memory outside the communication device and provide a communication interface for the processing circuit to access the memory; The processing circuit is used to execute program instructions in the memory to implement the method according to any one of claims 1 to 10.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores program codes, and when the program codes are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Task adjustment method applied to task engine, related device and storage medium
CN113377348A
Commodity redundant data updating method and device, equipment and medium
CN114969067A
Proximity graph maintenance for fast online nearest neighbor search
CN115935013A
Cited By
Vector index construction method and device and storage medium
CN121166697A