Vector graph index processing method and device

By scanning and deleting the in-degree node and its edges of the logically deleted node in the vector graph index, the vector graph index instability problem caused by the logically deleted node is solved, and the efficiency and accuracy of vector retrieval is improved.

CN120067094APending Publication Date: 2025-05-30HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311615682.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, when the vector graph index performs logical deletion node operation, the logical deletion node and its related edges still exist, resulting in poor stability of the vector graph index, affecting the efficiency and accuracy of vector retrieval.

Method used

By scanning the vector graph index, the entry information of the logical delete node is obtained, and the entry node and its edges of the node are deleted based on this information, real deletion is achieved and memory resources are released.

Benefits of technology

It improves the stability of vector graph index, enhances the efficiency and accuracy of vector retrieval, ensures the adequacy of recall results, and reduces user operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067094A_ABST
    Figure CN120067094A_ABST
Patent Text Reader

Abstract

The invention provides a vector diagram index processing method and device, and belongs to the technical field of data retrieval, and the method comprises the steps that a vector diagram index is scanned based on the number of nodes and identifiers of logic deletion nodes, in-degree information of the logic deletion nodes is obtained, and the in-degree information indicates in-degree nodes of the logic deletion nodes; and deleting the in-degree node of the logic deletion node and the edge of the logic deletion node based on the in-degree information of the logic deletion node. According to the method, the memory resources occupied by the related edges of the logic deletion nodes can be released, the vector retrieval efficiency and accuracy are improved, and the sufficiency of recall results is ensured. The method and the device are used for deleting the vector graph index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data retrieval, and in particular, to a method and apparatus for processing a vector graph index. Background Art

[0002] For massive amounts of unstructured data such as pictures, texts, videos, and voices, feature vectors can be extracted, and the analysis and retrieval of unstructured data can be achieved through the calculation and retrieval of feature vectors. In vertical search business scenarios such as large language models (LLMs) and news recommendations, a large amount of real-time data is continuously generated, and streaming vector data is also constantly updated. A vector graph index is a vector index for vector retrieval based on a graph data structure, which has the advantages of sustainable update, high precision, and high performance, and can fully meet business requirements.

[0003] In a vector graph index, operations such as adding nodes, deleting nodes, and modifying nodes need to be performed according to real-time data. For the delete node operation, in related technologies, each node has a flag bit, which is used to indicate whether the node is logically deleted, and the node is logically deleted by setting this flag bit. The edges connected to the logically deleted node (hereinafter referred to as the logically deleted node) still exist, and the edges related to the logically deleted node will be traversed during vector retrieval, but will not be included in the retrieval result set.

[0004] However, in the method of logically deleting nodes, the logically deleted nodes and related edges still exist in the graph. These holes lead to poor stability of the vector graph index, which in turn reduces the efficiency and accuracy of vector retrieval, and may lead to insufficient retrieval recall results. In addition, the logically deleted nodes and related edges still occupy memory resources. Therefore, in related technologies, it is necessary to periodically perform a full-scale construction of the vector graph index, and the full-scale construction process results in a high complexity of vector retrieval. Summary of the Invention

[0005] This application provides a method and apparatus for processing a vector graph index, which solves the problem that the logically deleted nodes and related edges still exist in the graph in related technologies, resulting in poor stability of the vector graph index, can release the memory resources occupied by the related edges of the logically deleted nodes, improve the efficiency and accuracy of vector retrieval, and ensure the sufficiency of recall results.

[0006] In a first aspect, this application provides a method for processing a vector graph index, the method includes: scanning the vector graph index based on the number of nodes and the identifiers of the logically deleted nodes to obtain the in-degree information of the logically deleted nodes, where the in-degree information indicates the in-degree nodes of the logically deleted nodes; based on the in-degree information of the logically deleted nodes, deleting the edges between the in-degree nodes of the logically deleted nodes and the logically deleted nodes.

[0007] The related art also provides a general Lambda architecture Milvus. In this architecture, for a deleted node, it is first written to the buffer area. When the amount of data in the buffer area is greater than the specified threshold, the data in the buffer area is converted into read-only data, and the ID of the deleted node is written to disk and written into the delete (Del) file inside the corresponding segment (seg) file. Only after the disk writing is completed is the deleted node invisible to the user and the deletion takes effect.

[0008] Milvus periodically generates multiple seg files. The seg files are used to store the original vector information, and the internal Del file is used to mark whether each node is deleted. The Del file is loaded into the in-memory bitmap. During vector retrieval, filtering is performed while retrieving through the deletion bits of the bitmap. For multiple seg files, the user needs to manually call the compact interface to merge the multiple seg files and reconstruct the vector graph index during the merging.

[0009] However, in the related art, it is still a logical deletion and does not achieve the true deletion of nodes, which will also cause problems such as poor stability of the vector graph index and occupation of memory resources. In addition, during the process of the user merging the seg files, the vector graph index is reconstructed, and the reconstruction process consumes a high amount of energy for the processor, resulting in a complex multi-index state machine and low retrieval performance.

[0010] In the embodiments of the present application, the graph vector engine realizes high-performance true deletion when the in-degree information is not stored in the adjacency matrix through the scanning process, converts the logically deleted nodes into unreachable isolated nodes, and the nodes and related edges after true deletion do not occupy memory resources, improving the stability of the vector graph index, thereby improving the efficiency and accuracy of vector retrieval and ensuring the sufficiency of the recall results.

[0011] In a possible implementation manner, the method further includes: determining the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, where the first in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: the updated neighbor node of the first in-degree node, the identifier of the second-order neighbor node of the target node in the updated neighbor node of the first in-degree node, and the target set; where the updated neighbor node does not include the logically deleted node, the target node includes the n nodes closest to the first in-degree node in the updated neighbor node, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.

[0012] The beneficial effect is to adaptively adjust the relevant nodes and edges of the deleted node by constructing an out-degree node candidate set, optimize the out-degree of the in-degree nodes of the deleted node and the in-degree of the out-degree nodes of the deleted node, so as to improve the efficiency and accuracy of vector retrieval and the index recall rate, while ensuring that the approximate nearest neighbor search (ANN) neighbor relationship of the points and edges related to the logically deleted node is still globally optimal.

[0013] In a possible implementation, the out-degree node candidate set includes a target set, and the method further includes: increasing the weights of the nodes in the target set in the out-degree node candidate set.

[0014] The beneficial effect is that the nodes in the neighbor list of the logically deleted node are more likely to be selected as the neighbor nodes of the first in-degree node, that is, the in-degree is more likely to be incremented by 1, thereby optimizing the structure of the vector graph index.

[0015] In a possible implementation, the out-degree node candidate set includes a target set, and the first node in the target set is included in the optimal neighbor nodes of the first in-degree node. The method further includes: updating the target set by deleting the first node in the target set.

[0016] The beneficial effect is that it can ensure as much as possible that each node in the neighbor list of the logically deleted node is selected by other nodes as a neighbor node, so that the in-degree of each node in the neighbor list of the logically deleted node is incremented by 1 as much as possible.

[0017] In a possible implementation, after determining the optimal neighbor nodes of the first in-degree node from the out-degree node candidate set of the first in-degree node, the method further includes: performing a reverse edge connection on the first in-degree node.

[0018] The beneficial effect is to further optimize the structure of the local vector graph index by considering whether the first in-degree node can be a neighbor node of the optimal neighbor node.

[0019] In a possible implementation, the process of scanning the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain the in-degree information of the logically deleted node includes: scanning the vector graph index based on the number of nodes and the identifier of the logically deleted node through an asynchronous thread to obtain the in-degree information of the logically deleted node; the process of deleting the edges between the in-degree nodes of the logically deleted node and the logically deleted node based on the in-degree information of the logically deleted node includes: deleting the edges between the in-degree nodes of the logically deleted node and the logically deleted node through the main thread based on the in-degree information of the logically deleted node.

[0020] The beneficial effects are that the process of scanning to obtain in-degree information is executed by an asynchronous thread, and the process of in-degree cleaning is executed by the main thread. In this way, there is no need to additionally add in-degree information to the vector graph index, reducing the complexity and memory occupancy of the vector graph index while efficiently obtaining in-degree information. The main thread is used to execute processes such as adding, querying, modifying, deleting, and vector retrieval of the vector graph index. The scanning process can be parallel to the main thread without interference, obtaining the in-degree information of each logically deleted node without affecting the main thread. Thus, it can ensure that the memory does not expand during long-term addition, query, modification, and deletion operations, enabling the creation of a one-stop streaming continuous service and out-of-the-box vector retrieval engine, improving the timeliness and accuracy of the recall system, eliminating the cumbersome periodic index reconstruction process, and reducing the user's usage and maintenance costs.

[0021] In a possible implementation, the process of scanning the vector graph index by an asynchronous thread based on the number of nodes and the identifiers of logically deleted nodes includes: when the second node is not locked, scanning the neighbor nodes of the second node by the asynchronous thread, where the second node is any node to be scanned.

[0022] In a possible implementation, the process of deleting the edges between the in-degree nodes of the logically deleted node and the logically deleted node by the main thread based on the in-degree information of the logically deleted node includes: when the second in-degree node is not locked, deleting the edge between the second in-degree node and the logically deleted node by the main thread, where the second in-degree node is any in-degree node of the logically deleted node.

[0023] In a possible implementation, the identifiers of the logically deleted nodes are stored in the logical deletion pool, and the method further includes: when the logical deletion pool meets the update condition, obtaining the number of nodes in the vector graph index and the logical deletion pool, where the update condition includes: the data size stored in the logical deletion pool reaches the first threshold or the usage ratio of the logical deletion pool reaches the second threshold.

[0024] In a possible implementation, the method further includes: in response to a deletion request for a third node received, logically deleting the third node and adding the identifier of the third node to the logical deletion pool.

[0025] In Milvus, the deletion operation is visible only after it is written to disk, and the deletion takes effect after a long time, which may cause data inconsistency. In the embodiments of the present application, the deletion request for the received node is directly converted into a logical deletion. The logical deletion does not involve any adjustment of nodes and edges in the vector graph index, and the user interface immediately returns, and the deletion can take effect in seconds. Compared with the related technology, the deletion effective time is reduced.

[0026] In a second aspect, the present application provides a processing device for vector graph indexing. The device includes: a scanning module configured to scan the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain the in-degree information of the logically deleted node, where the in-degree information indicates the in-degree nodes of the logically deleted node; a cleaning module configured to delete the edges between the in-degree nodes of the logically deleted node and the logically deleted node based on the in-degree information of the logically deleted node.

[0027] In a possible implementation, the device further includes: a determination module configured to determine the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, where the first in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: the updated neighbor nodes of the first in-degree node, the identifiers of the second-order neighbor nodes of the target nodes in the updated neighbor nodes of the first in-degree node, and the target set; where the updated neighbor nodes do not include the logically deleted node, the target nodes include the n nodes closest to the first in-degree node in the updated neighbor nodes, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.

[0028] In a possible implementation, the out-degree node candidate set includes the target set, and the determination module is further configured to increase the weights of the nodes in the target set in the out-degree node candidate set.

[0029] In a possible implementation, the out-degree node candidate set includes the target set, and the first node in the target set is included in the optimal neighbor node of the first in-degree node. The determination module is further configured to update the target set by deleting the first node in the target set.

[0030] In a possible implementation, the determination module is further configured to perform reverse edge connection on the first in-degree node after determining the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node.

[0031] In a possible implementation, the scanning module is specifically configured to scan the vector graph index based on the number of nodes and the identifier of the logically deleted node through an asynchronous thread to obtain the in-degree information of the logically deleted node; the cleaning module is specifically configured to delete the edges between the in-degree nodes of the logically deleted node and the logically deleted node through the main thread based on the in-degree information of the logically deleted node.

[0032] In a possible implementation, the scanning module is specifically configured to scan the neighbor nodes of the second node through an asynchronous thread when the second node is not locked, where the second node is any node to be scanned.

[0033] In a possible implementation manner, the cleaning module is specifically configured to delete the edge between the second in-degree node and the logically deleted node through the main thread when the second in-degree node is not locked, where the second in-degree node is any in-degree node of the logically deleted node.

[0034] In a possible implementation manner, the identifier of the logically deleted node is stored in the logical deletion pool. The determination module is specifically configured to obtain the number of nodes in the vector graph index and the logical deletion pool when the logical deletion pool meets the update condition, where the update condition includes: the data size stored in the logical deletion pool reaches the first threshold or the usage ratio of the logical deletion pool reaches the second threshold.

[0035] In a possible implementation manner, the apparatus further includes: a logical deletion module, configured to perform logical deletion on the third node and add the identifier of the third node to the logical deletion pool in response to a deletion request of the third node received.

[0036] In a third aspect, the present application provides a processing apparatus for a vector graph index. The apparatus includes: one or more processors; a memory for storing one or more computer programs or instructions; when the one or more computer programs or instructions are executed by the one or more processors, the one or more processors implement the method according to any one of the first aspect.

[0037] In a fourth aspect, the present application provides a processing apparatus for a vector graph index, including a processor for executing the method according to any one of the first aspect.

[0038] In a fifth aspect, the present application provides a processing apparatus for a vector graph index. The apparatus includes: a processing circuit and an interface circuit; wherein, the interface circuit is used to couple with a memory external to the processing apparatus for the vector graph index and provide a communication interface for the processing circuit to access the memory; the processing circuit is used to execute program instructions in the memory to implement the method according to any one of the first aspect.

[0039] In the specific implementation process, the processing apparatus for the vector graph index may be a chip, the input circuit may be an input pin, the output circuit may be an output pin, and the processing circuit may be transistors, gate circuits, flip-flops, and various logic circuits, etc. The input signal received by the input circuit may be received and input by, for example, but not limited to, a receiver. The output signal output by the output circuit may be output to, for example, but not limited to, a transmitter and transmitted by the transmitter, and the input circuit and the output circuit may be the same circuit, which is used as the input circuit and the output circuit at different times respectively. The embodiments of the present application do not limit the specific implementation manners of the processor and various circuits.

[0040] In one implementation, the processing device for the vector graph index can be a wireless communication device, that is, a computer device supporting wireless communication functions. Specifically, the wireless communication device can be a terminal such as a smart phone, or a radio access network device such as a base station. A network chip can also be referred to as a system on chip (SoC), or simply as an SoC chip for short. The communication chip can include a baseband processing chip and a radio frequency processing chip. The baseband processing chip is sometimes also referred to as a modem or a baseband chip. The radio frequency processing chip is sometimes also referred to as a radio frequency transceiver or a radio frequency chip. In a physical implementation, some or all of the chips in the communication chip can be integrated inside the SoC chip. For example, the baseband processing chip is integrated in the SoC chip, and the radio frequency processing chip is not integrated with the SoC chip. The interface circuit can be the radio frequency processing chip in the wireless communication device, and the processing circuit can be the baseband processing chip in the wireless communication device.

[0041] In another implementation, the processing device for the vector graph index can be some components in a wireless communication device, such as integrated circuit products like network chips or communication chips. The interface circuit can be an input / output interface, an interface circuit, an output circuit, an input circuit, a pin, or a related circuit, etc. on the chip or chip network. The processor can also be embodied as a processing circuit or a logic circuit.

[0042] In a sixth aspect, the present application provides a computer-readable storage medium, in which program code is stored. When the program code is executed by a processor, the method described in any item of the first aspect is implemented.

[0043] In a seventh aspect, the present application provides a chip, including: at least one processor. The at least one processor is used to execute the method described in any item of the first aspect.

[0044] Optionally, the chip further includes a memory. The at least one processor is used to execute the code in the memory. When the at least one processor executes the code, the chip implements the method described in any item of the first aspect.

[0045] Optionally, the above chip can also be an integrated circuit.

[0046] In an eighth aspect, the present application provides a computer program product containing instructions. When it runs on a computer, the computer implements the method described in any item of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a schematic diagram of an LLM scenario provided by an embodiment of the present application;

[0048] Figure 2Schematic flowchart of a method for processing vector graph indexes provided by an embodiment of the present application;

[0049] Figure 3 Schematic diagram of a data structure of a vector graph index provided by an embodiment of the present application;

[0050] Figure 4 Schematic diagram of the in-degree cleaning process of a logically deleted node provided by an embodiment of the present application;

[0051] Figure 5 Flowchart of a method for processing vector graph indexes provided by an embodiment of the present application;

[0052] Figure 6 Schematic flowchart of another method for processing vector graph indexes provided by an embodiment of the present application;

[0053] Figure 7 Schematic diagram of the optimization process of a vector graph index provided by an embodiment of the present application;

[0054] Figure 8 Schematic diagram of the effect of a method for processing vector graph indexes provided by an embodiment of the present application;

[0055] Figure 9 Block diagram of a device for processing vector graph indexes provided by an embodiment of the present application;

[0056] Figure 10 Block diagram of another device for processing vector graph indexes provided by an embodiment of the present application;

[0057] Figure 11 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application;

[0058] Figure 12 Schematic diagram of the structure of a device for processing vector graph indexes provided by an embodiment of the present application. Detailed implementation manners

[0059] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application fall within the scope of protection of the present application.

[0060] In the embodiments of the specification, claims and drawings of the present application, terms such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0061] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0062] In business scenarios such as recommendation systems, image recognition, natural language processing, voiceprint matching, audio retrieval, and file deduplication, the application of vector retrieval is becoming more and more widespread. Vector retrieval refers to retrieving the K vectors closest to the query vector in a given vector dataset according to a certain metric (L2, inner product...). The essence of vector retrieval is the ANN algorithm, and the ANN algorithm is mainly implemented based on four types of data structures: trees, graphs, quantization, and hashing.

[0063] Among them, the graph-based vector index (referred to as vector graph index) can realize the online construction of vector index from 0 to 1 without offline model training. The nearest neighbor retrieval of the vector graph index is to perform a global directional greedy traversal starting from a fixed node in a real-time insertion scenario to determine the nearest neighbor of the query request, and it supports dynamic addition of nodes.

[0064] The embodiments of the present application are for the processing of vector graph indexes. Therefore, the following first explains each noun used in the embodiments of the present application:

[0065] Vector graph index: It uses the structure of a graph adjacency matrix to describe the nearest neighbor relationship between nodes, and at the same time saves the original vector of each node for distance calculation;

[0066] Adjacency matrix: A two-dimensional array storage structure of a graph. One dimension stores the node (also known as vertex) information in the graph, and the other dimension stores the edge information in the graph. The edge information is also the neighbor list of each node, and the neighbor list is used to store the identifiers of the neighbor nodes (out-degree neighbor nodes) of the node;

[0067] Out-degree: Taking a certain node as the tail of the arc, the number of arcs starting from this node is called the out-degree of this node;

[0068] In-degree: Taking a certain node as the head of the arc, the number of arcs ending at this node is called the in-degree of this node.

[0069] The embodiments of this application can be applied to vertical search business scenarios such as LLM and news recommendation. Taking LLM as an example, due to factors such as the training cycle cost of LLM, the features of LLM are static. The graph vector engine can be used as a retrieval enhancement plugin for LLM, assisting LLM to provide LLM with dynamic knowledge supplementation capabilities, supporting vector streaming to continuously serve arbitrary addition, deletion, modification, and query.

[0070] Exemplarily, please refer to Figure 1 , Figure 1 which is a schematic diagram of an LLM scenario provided by the embodiments of this application, Figure 1 showing three business scenarios when the LLM-based system (i.e., LLM), the feature analysis module, and the graph vector engine are applied to LLM. These three business scenarios respectively include: 1. The graph vector engine serves as the generative pre-trained transformer cache (GPTCache) of LLM; 2. The graph vector engine serves as the personal historical memory / private domain knowledge base of LLM; 3. The graph vector engine serves as the and vertical domain knowledge base (knowledge base) of LLM. The personal historical memory / private domain knowledge base is a long-term memory.

[0071] As Figure 1 shown, for GPTCache, when the frequency of a query is high, this query can be inserted online into the vector graph index. In this way, approximate nearest neighbor retrieval can be performed through the vector graph index, without requesting LLM for retrieval, which can improve the inference response speed of LLM and at the same time reduce the inference cost of LLM.

[0072] For specific domains such as the personal historical memory / private domain knowledge base and the vertical domain knowledge base, the description accuracy of LLM is relatively low. Therefore, the knowledge of specific domains can be independently indexed through the vector graph index and accurate ANN can be provided. As Figure 1As shown, the feature analysis module encodes the user's streaming unstructured data to generate vector data, and the graph vector engine constructs an index of the vector data online, so as to achieve second (s)-level visibility for the user's plug-and-play query. The user's streaming unstructured data may include question & answer (QAs) and documents (Docs), etc.

[0073] An embodiment of the present application provides a method for processing a vector graph index. This method can be applied to a graph vector engine (such as a Poisson vector engine service) to implement the real deletion of the vector graph index. Please refer to Figure 2 , Figure 2 is a schematic flowchart of a method for processing a vector graph index provided by an embodiment of the present application. The vector graph index can support online streaming create / retrieve / update / delete (CRUD). This method may include the following processes:

[0074] 101. Scan the vector graph index based on the number of nodes and the identifiers of the logically deleted nodes to obtain the in-degree information of the logically deleted nodes. The in-degree information indicates the in-degree nodes of the logically deleted nodes.

[0075] A logically deleted node refers to a node that has been logically deleted. When the graph vector engine receives a deletion request for a first node, in response to the deletion request for the first node, the first node is logically deleted. The first node refers to any node to be deleted, and this process can be executed when a deletion request for any node is received.

[0076] The graph vector engine directly converts the received deletion request for a node into a logical deletion. The logical deletion does not involve any adjustment of nodes and edges in the vector graph index. The user interface immediately returns, and the transaction per second (TPS) for deletion is high, so that the deleted nodes can be made invisible to the user in seconds. This process can be executed by the main thread in the graph vector engine.

[0077] Exemplarily, a deletion flag bit can be set for each node in the data structure of the vector graph index. The main thread, in response to the deletion request for the first node, sets the deletion flag bit of the first node to "logical deletion". For example, the deletion flag bit can be the Del bit. When the Del bit is 0x00, it indicates that the node is normal; when the Del bit is 0x01, it indicates that the node has been logically deleted; when the Del bit is 0x11, it indicates that the node has been really deleted. On this basis, there is also a function named is_delete. It can be determined whether the node has been logically deleted by determining whether the Del bit of the node is 0x01 through the is_delete function.

[0078] In the embodiments of the present application, the implementation logic for logically deleting nodes to make them invisible is designed at the algorithm level. When inserting a new node, the logically deleted nodes are not considered, and the in-degree of the logically deleted nodes remains unchanged. When the retrieval process traverses a logically deleted node, the edges associated with the logically deleted node can be accessed, but the final retrieval result does not include the logically deleted node.

[0079] Optionally, the graph vector engine may be preset with a logical deletion pool, and after logically deleting the first node, the identity (ID) of the first node is added to the logical deletion pool. The logical deletion pool can be std::vector <int32>The logic_delete_pool, the std means (standard), the vector means (vector), the int32 means (32-bit signed integer), and the logic_delete_pool means (logical deletion pool).

[0080] The graph vector engine can periodically monitor the storage status of the logical deletion pool. When the logical deletion pool meets the update conditions, it starts to convert logical deletion nodes to real deletion nodes. First, obtain the number of nodes in the vector graph index and the logical deletion pool at the current moment. Among them, the update conditions include but are not limited to: the size of the data stored in the logical deletion pool reaches the first threshold or the usage ratio of the logical deletion pool reaches the second threshold. The update conditions can be customized. For example, it can also be that a preset time interval has passed since the last execution of this process. The embodiments of the present application do not limit this. This process can be executed by an asynchronous thread in the graph vector engine, that is, this process is executed in the background.

[0081] The in-degree information can include the identifiers of the in-degree nodes of the logical deletion nodes. The in-degree nodes of the logical deletion nodes refer to the nodes connected to the other end of the in-edge of the logical deletion nodes. The graph vector engine can perform a full scan of the vector graph index based on the number of nodes and the logical deletion pool. The number of nodes is used to limit the scanning range. The graph vector engine scans these numbers of nodes, which can effectively avoid scanning the nodes newly added after the start scanning time.

[0082] As can be seen from the foregoing description, the neighbor list of each node is recorded in the data structure of the vector graph index. The neighbor list of the node is used to record the identifiers of the out-degree nodes of the node. The out-degree nodes of the node refer to the nodes connected to the other end of the out-edge of the node. The graph vector engine scans the neighbor list of each node. If the identifier of a certain logical deletion node is recorded in the neighbor list of a certain node, the identifier of this node is added to the in-degree information of this logical deletion node.

[0083] This process can be executed by an asynchronous thread, so there is no need to additionally add in-degree information to the vector graph index, which can reduce the complexity and memory occupancy of the vector graph index. The main thread is used to execute the CRUD of the vector graph index and vector retrieval and other processes. The scanning process can be parallel to the main thread and does not interfere with each other, so as to obtain the in-degree information of each logical deletion node without affecting the main thread.

[0084] In the embodiments of the present application, an independent lock can be set for each node on the data structure of the vector graph index to ensure the thread safety of the node neighbor list, effectively avoid interference between the main thread and the asynchronous thread, and achieve high-performance parallel scanning.

[0085] When the main thread operates on the neighbor list of a certain node (such as CRUD or retrieval), it will lock the node. During the process of the asynchronous thread scanning the vector graph index, it determines whether the lock of a certain node is occupied before scanning the neighbor list of the node. When the lock of the node is not occupied, it scans the neighbor list of the node. When the lock of the node is occupied, the scanning process is blocked. Usually, the blocking time is 1 millisecond (ms).

[0086] Similarly, when the asynchronous thread scans the neighbor list of a certain node, it will also lock the node. The main thread determines whether the lock of the node is occupied before operating on the neighbor list of the node. When the lock of the node is not occupied, it operates on the node. When the lock of the node is occupied, the current operation is blocked.

[0087] For example, please refer to Figure 3 , Figure 3 which is a schematic diagram of the data structure of a vector graph index provided by an embodiment of this application. Figure 3 Taking four nodes A, B, O, and C as an example, the process of parallelism between the main thread and the asynchronous thread is further described. Figure 3 It shows the lock bit, Del bit, ID, and neighbor list of the four nodes. The identifiers of these four nodes are A, B, O, and C respectively. It should be noted that Figure 3 the identifiers of the nodes shown in the neighbor list are only for illustrative purposes and do not constitute any limitation.

[0088] Suppose that at time T1, the main thread receives a delete instruction for node O, then sets the Del bit of node O to 0x01, marks node O as a logically deleted node, and adds the ID of node O to the logic_delete_pool. At this time, node O is invisible both on the user side and in the algorithm layer. If the main thread adds node C to the vector graph index after time T1, the main thread traverses each node in the vector graph index to globally find M nearest neighbor nodes. When the main thread traverses each node, it first determines whether the node is locked through the lock bit. When the node is not locked, it then determines whether the node is logically deleted through the is_delete function. When a certain node is logically deleted, it will not add the node to the neighbor list of node C. Therefore, the in-degree of node O remains unchanged after time T1.

[0089] To ensure the accuracy of the in-degree information obtained by scanning, the asynchronous thread determines whether a node is locked before scanning the neighbor list of each node in parallel. If a node is locked, the scanning will be blocked. When adding node C and performing reverse edge connection to node C, which causes a change in the neighbor list of node A, the main thread will lock node A. If the asynchronous thread is ready to start scanning the neighbor nodes of node A but node A is still locked, the asynchronous thread will block the scanning.

[0090] As Figure 3 shown, the asynchronous thread scans the identifier of node O in the neighbor lists of both node A and node B. Therefore, node A and node B are the in-degree nodes of node O, and the in-degree information of node O includes: A and B. And the neighbor list of the newly added node C does not include O.

[0091] When the data volume of the vector graph index is large, if the TPS of the CRUD or retrieval operations of the main thread is normal, the collision probability between the main thread and the asynchronous thread is small, and it can basically ensure that the main thread and the asynchronous thread do not interfere with each other.

[0092] Exemplarily, the identifier of the logically deleted node whose in-degree information has been determined can be deleted from the logical deletion pool, that is, the logical deletion pool can be continuously updated.

[0093] 102. Based on the in-degree information of the logically deleted node, delete the edge between the in-degree node of the logically deleted node and the logically deleted node.

[0094] Optionally, after determining the in-degree information of all logically deleted nodes, for each logically deleted node, delete the edge between each in-degree node of the logically deleted node and the logically deleted node. Or, for each logically deleted node whose in-degree information is determined, delete the edge between each in-degree node of the logically deleted node and the logically deleted node. The embodiments of the present application do not make any limitations in this regard.

[0095] As can be seen from the foregoing description, the neighbor list of a node is used to record the identifiers of the out-degree nodes of the node, that is, the neighbor list is used to store the edges of each node in the vector graph index. The graph vector engine can delete the edge between the in-degree node of the logically deleted node and the logically deleted node by deleting the identifier of the logically deleted node from the neighbor list of the in-degree node of the logically deleted node.

[0096] By executing this process 102, the logically deleted node is transformed into an unreachable isolated node, realizing the true deletion of the node, and the memory / disk occupied area of the node can be reused. Thus, problems such as low retrieval efficiency and accuracy and insufficient recall results caused by the invalid edges of the logically deleted nodes are solved.

[0097] This process can be executed by the main thread. After the asynchronous thread determines the in-degree information of the logically deleted node, before the main thread deletes the edge between the first in-degree node of the logically deleted node and the logically deleted node, it first determines whether the first in-degree node is locked. When the first in-degree node is not locked, it locks the first in-degree node and then deletes the edge between the first in-degree node and the logically deleted node, thus quickly completing the in-degree cleaning of the logically deleted node. When the first in-degree node is locked, the in-degree cleaning process is blocked. The first in-degree node is any in-degree node of the logically deleted node.

[0098] In the embodiment of the present application, since the asynchronous thread has determined the in-degree information of the logically deleted node, the main thread only needs to directly read the in-degree information of the logically deleted node determined by the asynchronous thread, lock the in-degree nodes of the logically deleted node, and delete the edges between the in-degree nodes of the logically deleted node and the logically deleted node. The complexity of the in-degree cleaning of the logically deleted node is O1, which makes the locking time shorter and the deletion efficiency higher. Compared with the related technology, the comprehensive queries per second (QPS) performance of CRUD is improved.

[0099] Exemplarily, please refer to Figure 4 , Figure 4 which is a schematic diagram of the in-degree cleaning process of a logically deleted node provided by the embodiment of the present application. Figure 4 It is described by taking Figure 3 the in-degree scan result shown as an example, and the in-degree cleaning process is described by taking the neighbor list as an example. Figure 4 It includes two parts (a) and (b). (a) shows the vector graph and the neighbor lists of node A, node B, and node O before in-degree cleaning. (b) shows the vector graph and the neighbor lists of the nodes, node B, and node O after in-degree cleaning.

[0100] As shown in (a) of Figure 4 , the neighbor lists of node A and node B include node O. After in-degree cleaning, as shown in (b) of Figure 4 , O in the neighbor list of node A is deleted, and O in the neighbor list of node B is deleted. At the same time, in the vector graph, the edge from node A to node O is deleted, and the edge from node B to node O is deleted, and node O is unreachable.

[0101] Exemplarily, a real deletion pool std::vector can also be added to the vector graph index data structure. <int32>real_delete_pool. The meaning of "real" is actual. After deleting the identifier of the logically deleted node from the neighbor list of each in-degree node of the logically deleted node, the logically deleted node becomes a real deleted node. At this time, the identifier of the node can be added to the real delete pool, and the Del position of the node is set to 0x11.

[0102] The following uses the main thread and the asynchronous thread as the execution entities to illustrate the processes 101 to 102. Please refer to Figure 5 , Figure 5 This is a flowchart of a method for processing a vector graph index provided by an embodiment of the present application. As shown in ①, in response to a deletion request for a certain node received, the main thread adds the ID of the node to the logic_delete_pool and sets the Del bit of the node to 0x01. Then the main thread will feedback a deletion response, and the deletion response indicates that the logical deletion is successful or the logical deletion fails. As shown in ②, the logically deleted node is invisible when inserting a new node, and the logically deleted node is filtered out when performing a search.

[0103] As shown in ③, the asynchronous thread monitors the status of the logic delete pool to determine whether a threshold (i.e., whether the update condition is met) is reached. The measurement elements for whether the threshold is reached include: data size, ratio, or time. For the data size, it is determined that the data size stored in the logic delete pool reaches the first threshold. For the ratio, it is determined whether the usage ratio of the logic delete pool reaches the second threshold. For the time, it is determined whether the time interval from the most recent update time reaches the third threshold. As shown in ④, based on the number of vector index points at the snapshot scan indegree moment and the logic delete pool, an in-degree scan is performed to obtain the indegree map of the logically deleted node. This process is parallel with the processes of adding a new node, retrieving, and logical deletion (parallel with [add][search][logic delete]).

[0104] As shown in ⑤, after the main thread performs indegree cleaning, it adds the ID of the node to the real_delete_pool, and the memory occupied by the node and the related edges can be reused.

[0105] In summary, for the method for processing a vector graph index provided in the embodiments of the present application, first, based on the number of nodes and the identifiers of logically deleted nodes, the vector graph index is scanned to obtain the in-degree information of the logically deleted nodes. The in-degree information indicates the in-degree nodes of the logically deleted nodes. Then, based on the in-degree information of the logically deleted nodes, the edges between the in-degree nodes of the logically deleted nodes and the logically deleted nodes are deleted. The graph vector engine directly converts the received node deletion request into a logical deletion. The logical deletion does not involve any adjustment of nodes and edges in the vector graph index, and the user interface immediately returns. The deletion can take effect in seconds. Through the scanning process, high-performance real deletion is achieved when the in-degree information is not stored in the adjacency matrix. The logically deleted nodes are converted into unreachable isolated nodes. The nodes and related edges after real deletion do not occupy memory resources, improving the stability of the vector graph index, thereby improving the efficiency and accuracy of vector retrieval and ensuring the sufficiency of recall results.

[0106] Moreover, in the embodiments of the present application, the process of scanning to obtain the in-degree information can be executed by an asynchronous thread, and the process of in-degree cleaning can be executed by the main thread. In this way, there is no need to additionally add in-degree information to the vector graph index, reducing the complexity and memory occupancy of the vector graph index while efficiently obtaining the in-degree information. The main thread is used to execute processes such as CRUD of the vector graph index and vector retrieval. The scanning process can be parallel to the main thread and does not interfere with each other. Without affecting the main thread, the in-degree information of each logically deleted node can be obtained, so that the memory does not expand under long-term CRUD, and a one-stop streaming and continuously available vector retrieval engine can be built, as well as improving the timeliness and accuracy of the recall system, eliminating the complicated periodic index reconstruction process, and reducing the user's usage and operation and maintenance costs.

[0107] When the embodiments of the present application are applied to the LLM scenario (such as Figure 1 the scenario shown), not only can the vector graph index be used to construct vector data online in real time, achieving second-level visibility for users to plug and query, but also when there are expired or to-be-updated vector data in the vector graph index, the data points can be truly deleted or updated, and the user's add, delete, and modify requests can be continuously received. The memory of the vector graph index will not expand, and the retrieval accuracy and retrieval performance will not decrease, eliminating the need for users to perform periodic offline full-scale construction.

[0108] Based on the foregoing embodiments, the structure of the vector graph index can be further optimized. Please refer to Figure 6 , Figure 6 which is a schematic flowchart of another method for processing a vector graph index provided in the embodiments of the present application. This method can be applied to a graph vector engine (such as the Poisson vector engine service). This method can include the following processes:

[0109] 201. Based on the number of nodes and the identification scan vector graph index of the logically deleted nodes, obtain the in-degree information of the logically deleted nodes, where the in-degree information indicates the in-degree nodes of the logically deleted nodes.

[0110] This process can refer to the foregoing process 101, and the embodiments of this application will not elaborate here.

[0111] 202. Based on the in-degree information of the logically deleted nodes, delete the edges between the in-degree nodes of the logically deleted nodes and the logically deleted nodes.

[0112] This process can refer to the foregoing process 102, and the embodiments of this application will not elaborate here.

[0113] Through the foregoing processes 201 to 202, the out-degree of the in-degree nodes of the logically deleted nodes is decreased by 1. Since the logically deleted nodes are unreachable, the in-degree of the out-degree nodes of the logically deleted nodes (i.e., the nodes in the neighbor list of the logically deleted nodes) is decreased by 1.

[0114] 203. Determine the optimal neighbor node of the second in-degree node from the out-degree node candidate set of the second in-degree node. The second in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: the updated neighbor nodes of the second in-degree node, the identifiers of the second-order neighbor nodes of the target nodes in the updated neighbor nodes of the second in-degree node, and the target set.

[0115] Among them, the updated neighbor nodes do not include the identifier of the logically deleted node. The target nodes include the n nodes closest to the second in-degree node in the updated neighbor nodes, where n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.

[0116] The updated neighbor nodes of the second in-degree node can be the nodes in the neighbor list after deleting the identifier of the logically deleted node. The second-order neighbor of the target node refers to the neighbor node (also called the out-degree neighbor node) of the neighbor node (also called the out-degree neighbor node) of the target node. n is a custom parameter, and n can be, for example, 2 or 3, etc. The embodiments of this application do not limit this.

[0117] Exemplarily, if the out-degree node candidate set includes the target set, the weights of the nodes in the target set in the out-degree node candidate set can be increased. The weight of a node represents the probability that the node is selected as a neighbor node. Since the in-degree of the neighbor nodes of the logically deleted nodes is decreased by 1 after the foregoing processes 201 to 202, through this process, the nodes in the neighbor list of the logically deleted nodes are more likely to be selected as the neighbor nodes of the second in-degree node, that is, the in-degree is more likely to be increased by 1, thereby optimizing the structure of the vector graph index.

[0118] If the weight of a node in the target set is increased in the out-degree node candidate set, then in this process, the optimal neighbor node of the second in-degree node is determined from the out-degree node candidate set in which the weight of the node in the target set is increased.

[0119] Exemplarily, the optimal neighbor node of the second in-degree node can be determined from the out-degree node candidate set through a heuristic edge selection strategy. Heuristic edge selection is an edge selection strategy introduced in the vector graph index of the hierarchical navigable small world (HNSW) algorithm. It means that when a newly added node retrieves and traverses EfConstruction (referring to a variable in HNSW) nearest neighbor nodes globally, and the maximum number of neighbors of the newly added node is M (EfConstruction > M), M neighbor nodes will be optimally selected from the EfConstruction nearest neighbor nodes. In this process, both the distance and directionality are considered to prevent the edge aggregation retrieval from falling into a local optimum.

[0120] When the out-degree node candidate set includes the target set and the second node in the target set is included in the optimal neighbor nodes of the second in-degree node, the target set can be dynamically adjusted. Exemplarily, the target set can be updated by deleting the second node in the target set. This can ensure that as much as possible, each node in the neighbor list of the logically deleted node is selected as a neighbor node by other nodes, so that the in-degree of each node in the neighbor list of the logically deleted node is increased by 1 as much as possible.

[0121] Exemplarily, when the target set is emptied, the target set can be restored, and the nodes in the neighbor list of the logically deleted node are added back to the target set.

[0122] After determining the optimal neighbor node of the second in-degree node from the out-degree node candidate set, the second in-degree node can also be reversely connected. That is, consider whether the second in-degree node can be a neighbor node of the optimal neighbor node, so as to further optimize the structure of the local vector graph index.

[0123] It should be noted that the foregoing process 203 is described by taking the second in-degree node of the logically deleted node as an example. The relevant processes of any in-degree node of any logically deleted node can refer to process 203, and the embodiments of the present application will not be elaborated herein.

[0124] Please refer to Figure 7 , Figure 7 which is a schematic diagram of an optimization process of a vector graph index provided by an embodiment of the present application, Figure 7 and is described by taking the neighbor list as an example. Figure 7 It includes two parts, (a) and (b). (a) represents the local vector graph index before optimization, and (b) represents the local vector graph index after optimization. Figure 7 Multiple nodes and the neighbor lists of node A, node B, and node O are shown in Figure 7 . Hereinafter, taking node O as the logically deleted node and node A as the second in-degree node as an example for illustration.

[0125] As Figure 7 shown in (a) of Figure 7 , the neighbor lists of the in-degree nodes of node O (such as node A and node B, etc.) have vacancies, and the out-degree is reduced by 1. Node O is unreachable, so the in-degree of each node (such as node C and node D) in the neighbor list of node O is reduced by 1. First, determine the out-degree node candidate set V. Taking the out-degree node candidate set V including: the updated neighbor list of node A, the identifiers of the second-order neighbor nodes of the target nodes in the updated neighbor list of node A, and the target set S as an example.

[0126] After that, use the heuristic weighted edge selection strategy to determine the optimal M neighbor nodes of node A. First, increase the weights of the nodes in set S in set V, and then use the heuristic edge selection strategy for the set V after increasing the weights to obtain the optimal M neighbor nodes of node A.

[0127] Then, perform dynamic adjustment of set V and reverse connection. If there is a subset Y in set S that is selected as the optimal M neighbor nodes of node A, then delete subset Y from set S. If set S is an empty set, then set set S to the neighbor list of node O again. And perform a reverse connection for node A. For example, if A is connected to the optimal neighbor node C, it is necessary to determine whether to connect node C to node A to further optimize the local graph structure.

[0128] The optimization process of the neighbor nodes of node B can refer to that of node A, and the embodiments of the present application will not elaborate here. As Figure 7 shown, the optimal neighbor nodes of node A include node C, and the optimal neighbor nodes of node B include node D.

[0129] The complexity of the deletion and self-repair processes in the aforementioned vector graph index is equal to O(2M), where M represents the maximum number of neighbor nodes of each node, usually 32. The self-update time of the actual deleted node is short, the deletion TPS performance is high, and the impact on the CRUD process of the vector graph index is small. And the entire process is executed based on the full-volume vector graph index without involving other temporary index structures, so the reliability is relatively high.

[0130] Please refer to Figure 8 , Figure 8 which is a schematic diagram of the effect of a method for processing a vector graph index provided by the embodiments of the present application. Figure 8 shows the change of the vector graph index after the logical deletion node is converted into a real deletion node by the method provided in the embodiments of the present application. It can be seen from Figure 8 that there are no logical deletion nodes and related edges in the vector graph index, and the structure of the vector graph index is adaptively adjusted, and the ANN neighbor relationship of the point edges related to the logical deletion nodes is still globally optimal.

[0131] In summary, for the method for processing the vector graph index provided by the embodiments of the present application, first, based on the number of nodes and the identifiers of the logical deletion nodes, the vector graph index is scanned to obtain the in-degree information of the logical deletion nodes, and the in-degree information indicates the in-degree nodes of the logical deletion nodes. Then, based on the in-degree information of the logical deletion nodes, the edges between the in-degree nodes of the logical deletion nodes and the logical deletion nodes are deleted. Finally, the optimal neighbor node of the second in-degree node is determined from the out-degree node candidate set of the second in-degree node. The second in-degree node is any in-degree node of the logical deletion node, and the out-degree node candidate set includes at least one of the following: the updated neighbor node of the second in-degree node, the identifier of the second-order neighbor node of the target node in the updated neighbor node of the second in-degree node, and the target set. The graph vector engine directly converts the received node deletion request into a logical deletion. The logical deletion does not involve any adjustment of the nodes and edges in the vector graph index, and the user interface immediately returns. The deletion can take effect in seconds. Through the scanning process, high-performance real deletion is achieved when the in-degree information is not stored in the adjacency matrix. The logical deletion node is converted into an unreachable isolated node. The nodes and related edges after real deletion do not occupy memory resources, improving the stability of the vector graph index. And by constructing the out-degree node candidate set, the relevant nodes and edges of the deleted node are adaptively adjusted, optimizing the out-degree of the in-degree node of the deleted node and the in-degree of the out-degree node of the deleted node. Thus, while improving the efficiency and accuracy of vector retrieval and ensuring the sufficiency of the recall results, it is ensured that the ANN neighbor relationship of the point edges related to the logical deletion nodes is still globally optimal.

[0132] The sequence of the methods provided by the embodiments of the present application can be appropriately adjusted, and the process can also be increased or decreased accordingly according to the situation. Any person skilled in the art can easily think of a changed method within the technical scope disclosed in the present application, and it should be covered by the protection scope of the present application. The embodiments of the present application do not make any limitations in this regard.

[0133] The above mainly introduced the processing method of the vector graph index provided by the embodiments of the present application from the perspective of the device. It can be understood that in order for the device to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that in combination with the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0134] The embodiments of the present application can divide the functions of the device according to the above method examples. For example, each function module can be divided corresponding to each function, or two or more functions can be integrated into a processing subsystem. The above integrated module can be implemented in the form of hardware or in the form of a software function module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there can be other division methods in actual implementation.

[0135] Figure 9 It is a block diagram of a processing device for a vector graph index provided by an embodiment of the present application. In the case of dividing each function module corresponding to each function, the processing device 300 for the vector graph index may include: a scanning module 301 and a cleaning module 302. Exemplarily, the processing device for the vector graph index may be a graph vector engine, or a chip therein, or other combined devices, components, etc. with the functions of the above-mentioned processing device for the vector graph index. The functions of each module of the device are as follows:

[0136] The scanning module 301 is configured to scan the vector graph index based on the number of nodes and the identifier of the logically deleted nodes, and obtain the in-degree information of the logically deleted nodes, where the in-degree information indicates the in-degree nodes of the logically deleted nodes; the cleaning module 302 is configured to delete the edges between the in-degree nodes of the logically deleted nodes and the logically deleted nodes based on the in-degree information of the logically deleted nodes.

[0137] Combined with the above solution, please refer to Figure 10 , Figure 10 It is a block diagram of another processing device for a vector graph index provided by an embodiment of the present application. In Figure 9 Based on the above, the device further includes: a determination module 303, configured to determine an optimal neighbor node of a first in-degree node from an out-degree node candidate set of the first in-degree node, where the first in-degree node is any in-degree node of a logically deleted node, and the out-degree node candidate set includes at least one of the following: an updated neighbor node of the first in-degree node, an identifier of a second-order neighbor node of a target node in the updated neighbor node of the first in-degree node, and a target set; wherein, the updated neighbor node does not include the logically deleted node, the target node includes the n nodes closest to the first in-degree node in the updated neighbor node, n is a positive integer, and the target set includes neighbor nodes of the logically deleted node.

[0138] Combined with the above solution, the out-degree node candidate set includes the target set, and the determination module 303 is further configured to increase the weights of the nodes in the target set in the out-degree node candidate set.

[0139] Combined with the above solution, the out-degree node candidate set includes the target set, and a first node in the target set is included in the optimal neighbor node of the first in-degree node. The determination module 303 is further configured to update the target set by deleting the first node in the target set.

[0140] Combined with the above solution, the determination module 303 is further configured to perform a reverse edge connection on the first in-degree node after determining the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node.

[0141] Combined with the above solution, the scanning module 301 is specifically configured to scan the vector graph index based on the number of nodes and the identifier of the logically deleted node through an asynchronous thread to obtain the in-degree information of the logically deleted node; the cleaning module 302 is specifically configured to delete the edge between the in-degree node of the logically deleted node and the logically deleted node based on the in-degree information of the logically deleted node through the main thread.

[0142] Combined with the above solution, the scanning module 301 is specifically configured to scan the neighbor nodes of a second node through an asynchronous thread when the second node is not locked, and the second node is any node to be scanned.

[0143] Combined with the above solution, the cleaning module 302 is specifically configured to delete the edge between the second in-degree node and the logically deleted node through the main thread when the second in-degree node is not locked, and the second in-degree node is any in-degree node of the logically deleted node.

[0144] Combined with the above solution, the identifier of the logically deleted node is stored in the logical deletion pool. The determination module 303 is specifically configured to obtain the number of nodes and the logical deletion pool in the vector graph index when the logical deletion pool meets the update condition, and the update condition includes: the data size stored in the logical deletion pool reaches a first threshold or the usage ratio of the logical deletion pool reaches a second threshold.

[0145] Combined with the above solution, as Figure 10 shown, the device further includes: a logical deletion module 304, configured to logically delete a third node in response to a deletion request received from a third node, and add an identifier of the third node to a logical deletion pool.

[0146] Figure 11 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 400 may be a chip or a functional module in a graph vector engine. As Figure 11 shown, the electronic device 400 includes a processor 401, a transceiver 402, and a communication line 403.

[0147] Among them, the processor 401 is configured to execute any step in the method embodiments as Figure 2 and Figure 6 shown, and when executing processes such as receiving a deletion request of a node, it can optionally call the transceiver 402 and the communication line 403 to complete corresponding operations.

[0148] Further, the electronic device 400 may further include a memory 404. Among them, the processor 401, the memory 404, and the transceiver 402 may be connected through the communication line 403.

[0149] The transceiver 402 is configured to communicate with other devices or other communication networks. The other communication networks may be an Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc. The transceiver 402 may be a module, a circuit, a transceiver, or any device capable of implementing communication.

[0150] The transceiver 402 is mainly used for sending and receiving requests, etc., and may include a transmitter and a receiver, which are respectively used for sending and receiving requests, etc.; operations other than sending and receiving requests, etc. are implemented by the processor, such as determining the number of nodes in the vector graph index and the logical deletion pool, etc.

[0151] The communication line 403 is configured to transmit information between the components included in the electronic device 400.

[0152] In one design, the processor can be regarded as a logic circuit, and the transceiver can be regarded as an interface circuit.

[0153] The memory 404 is configured to store instructions. Among them, the instructions may be computer programs.

[0154] It should be noted that the memory 404 can exist independently of the processor 401 or be integrated with the processor 401. The memory 404 can be used to store instructions, program codes, or some data, etc. The memory 404 can be located inside the electronic device 400 or outside the electronic device 400, without limitation. The processor 401 is configured to execute the instructions stored in the memory 404 to implement the method provided in the foregoing embodiments of the present application.

[0155] In one example, the processor 401 may include one or more processors, such as Figure 11 the processor 0 and the processor 1 in

[0156] As an alternative implementation, the electronic device 400 includes multiple processors. For example, in addition to Figure 11 the processor 401 in

[0157] As an alternative implementation, the electronic device 400 further includes an output device 405 and an input device 406. Exemplarily, the input device 406 is a device such as a keyboard, a mouse, a microphone, or a joystick, and the output device 405 is a device such as a display screen or a speaker.

[0158] It should be noted that the electronic device 400 can be a chip system or a device with a Figure 11 similar structure in Figure 11 Among them, the chip system can be composed of chips or can include chips and other discrete devices. Actions, terms, etc. involved among the embodiments of the present application can be referred to each other without limitation. The message names or parameter names in the messages exchanged between the devices in the embodiments of the present application are only examples, and other names can also be adopted in specific implementations without limitation. In addition, Figure 11 the shown composition structure does not limit the electronic device 400. In addition to Figure 11 the shown components, the electronic device 400 can include more or fewer components than

[0159] The processors and transceivers described in this application can be implemented on integrated circuits (ICs), analog ICs, radio frequency integrated circuits, mixed-signal ICs, application specific integrated circuits (ASICs), printed circuit boards (PCBs), electronic devices, etc. The processors and transceivers can also be fabricated using various IC process technologies, such as complementary metal oxide semiconductor (CMOS), n-metal-oxide-semiconductor (NMOS), p-channel metal oxide semiconductor (PMOS), bipolar junction transistor (BJT), BiCMOS, silicon germanium (SiGe), gallium arsenide (GaAs), etc.

[0160] Figure 12 FIG. 4 is a schematic structural diagram of a processing device for vector graph indexing provided for an embodiment of this application. The processing device for vector graph indexing can be applicable to the scenarios shown in the above method embodiments. For ease of explanation, Figure 12 only the main components of the processing device for vector graph indexing are shown, including a processor, a memory, a control circuit, and an input / output device. The processor is mainly used to process communication protocols and communication data, execute software programs, and process the data of software programs. The memory is mainly used to store software programs and data. The control circuit is mainly used for power supply and the transmission of various electrical signals. The input / output device is mainly used to receive data input by the user and output data to the user.

[0161] When the processing device for vector graph indexing is a graphic vector engine, the control circuit can be a motherboard, the memory includes storage media with storage functions such as hard disks, RAMs, ROMs, etc., the processor can include a baseband processor and a central processor. The baseband processor is mainly used to process communication protocols and communication data, and the central processor is mainly used to control the entire processing device for vector graph indexing, execute software programs, and process the data of software programs. The input / output device includes a display screen, a keyboard, a mouse, etc.; the control circuit can further include or be connected to a transceiver circuit or a transceiver, such as: a network cable interface, etc., for sending or receiving data or signals, such as for data transmission and communication with other devices. Further, an antenna can also be included, for the transceiver of requests, for data / request transmission with other devices.

[0162] According to the method provided by the embodiments of the present application, the present application also provides a computer program product, which includes computer program code. When the computer program code runs on a computer, it causes the computer to execute any of the methods described in the embodiments of the present application.

[0163] The embodiments of the present application also provide a computer-readable storage medium. All or part of the processes in the above method embodiments can be executed by a computer or a device with the processing ability of a vector graph index by computer programs or instructions to control the relevant hardware to complete. The computer programs or the set of instructions can be stored in the above computer-readable storage medium. When the computer programs or the set of instructions are executed, they can include the processes of the above method embodiments. The computer-readable storage medium can be an internal storage unit of the graph vector engine in any of the foregoing embodiments, such as the hard disk or memory of the graph vector engine. The above computer-readable storage medium can also be an external storage device of the above graph vector engine, such as a plug-in hard disk equipped on the above graph vector engine, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the above computer-readable storage medium can also include both the internal storage unit of the above graph vector engine and the external storage device. The above computer-readable storage medium is used to store the above computer programs or instructions and other programs and data required by the above graph vector engine. The above computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.

[0164] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0165] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the foregoing method embodiments and will not be described herein again.

[0166] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0167] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0168] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0169] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs and other various media that can store program codes.

[0170] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims. The above mainly introduces the processing method of the vector map index provided by the embodiments of the present application from the perspective of the device. It can be understood that in order for the device to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that in combination with the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

Claims

1. A method for processing a vector graph index, characterized in that, the method includes: scanning the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain the in-degree information of the logically deleted node, where the in-degree information indicates the in-degree nodes of the logically deleted node; deleting the edges between the in-degree nodes of the logically deleted node and the logically deleted node based on the in-degree information of the logically deleted node.

2. The method according to claim 1, characterized in that, the method further includes: determining the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, where the first in-degree node is any in-degree node of the logically deleted node, and the out-degree node candidate set includes at least one of the following: the updated neighbor nodes of the first in-degree node, the identifiers of the second-order neighbor nodes of the target nodes in the updated neighbor nodes of the first in-degree node, and the target set; wherein, the updated neighbor nodes do not include the logically deleted node, the target nodes include the n nodes closest to the first in-degree node in the updated neighbor nodes, n is a positive integer, and the target set includes the neighbor nodes of the logically deleted node.

3. The method according to claim 2, characterized in that, the out-degree node candidate set includes the target set, and the method further includes: increasing the weights of the nodes in the target set in the out-degree node candidate set.

4. The method according to claim 2 or 3, characterized in that, the out-degree node candidate set includes the target set, and the optimal neighbor node of the first in-degree node includes the first node in the target set, and the method further includes: updating the target set by deleting the first node in the target set.

5. The method according to any one of claims 2 to 4, characterized in that, after determining the optimal neighbor node of the first in-degree node from the out-degree node candidate set of the first in-degree node, the method further includes: performing a reverse connection edge on the first in-degree node.

6. The method according to any one of claims 1 to 5, characterized in that, the scanning the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain the in-degree information of the logically deleted node includes: scanning the vector graph index based on the number of nodes and the identifier of the logically deleted node through an asynchronous thread to obtain the in-degree information of the logically deleted node; the deleting the edges between the in-degree nodes of the logically deleted node and the logically deleted node based on the in-degree information of the logically deleted node includes: deleting the edges between the in-degree nodes of the logically deleted node and the logically deleted node through the main thread based on the in-degree information of the logically deleted node.

7. The method according to claim 6, characterized in that, the scanning the vector graph index through an asynchronous thread based on the number of nodes and the identifier of the logically deleted node includes: when the second node is not locked, scanning the neighbor nodes of the second node through the asynchronous thread, where the second node is any node to be scanned.

8. The method according to claim 6 or 7, characterized in that, The main thread deletes the in-degree information of the logically deleted node and deletes the edge between the in-degree node of the logically deleted node and the logically deleted node, including: When the second in-degree node is not locked, the main thread deletes the edge between the second in-degree node and the logically deleted node, where the second in-degree node is any in-degree node of the logically deleted node.

9. The method according to any one of claims 1 to 8, characterized in that the identifier of the logically deleted node is stored in a logical deletion pool, and the method further includes: When the logical deletion pool meets the update condition, obtain the number of nodes in the vector graph index and the logical deletion pool, where the update condition includes: the data size stored in the logical deletion pool reaches a first threshold or the usage ratio of the logical deletion pool reaches a second threshold.

10. The method according to claim 9, characterized in that the method further includes: In response to a deletion request for a third node, logically delete the third node and add the identifier of the third node to the logical deletion pool.

11. A processing device for a vector graph index, characterized in that the device includes: A scanning module for scanning the vector graph index based on the number of nodes and the identifier of the logically deleted node to obtain the in-degree information of the logically deleted node, where the in-degree information indicates the in-degree nodes of the logically deleted node; A cleaning module for deleting the edge between the in-degree node of the logically deleted node and the logically deleted node based on the in-degree information of the logically deleted node.

12. A processing device for a vector graph index, characterized in that the device includes: One or more processors; A memory for storing one or more computer programs or instructions; When one or more computer programs or instructions are executed by one or more processors, one or more processors implement the method according to any one of claims 1 to 10.

13. A processing device for a vector graph index, characterized in that the device includes: A processing circuit and an interface circuit; Wherein, the interface circuit is used to couple with a memory external to the communication device and provide a communication interface for the processing circuit to access the memory; The processing circuit is used to execute program instructions in the memory to implement the method according to any one of claims 1 to 10.

14. A computer-readable storage medium, characterized in that Program code is stored in the computer-readable storage medium, and when the program code is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Vector index construction method and device and storage medium

    CN121166697A