Query processing method and apparatus in graph database
By constructing a vector indexing framework in a graph database, the problem that traditional graph databases cannot handle vector data is solved. This enables the storage and retrieval of vector data and similarity search, expands the application scope of graph databases, and improves data processing efficiency and accuracy.
Patent Information
- Application Number
- CN202511013888.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-07-22
AI Technical Summary
Traditional graph databases cannot effectively process vector data, which means they cannot support similarity searches of high-dimensional embedded vectors, limiting their application areas that require vector similarity searches.
Add support for vector data to graph databases by building a vector indexing framework, including a vector index manager, a vector index library, and a vector indexing algorithm, to support efficient vector similarity queries.
It implements the function of storing and retrieving vector data in graph databases, supports similarity search based on vector attributes, expands the applicability of graph databases, improves the efficiency and accuracy of data processing, and reduces operating costs.
Smart Images

Figure CN120523982B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] One or more embodiments of the present specification relate to the field of data storage, and in particular, to a query processing method and apparatus in a graph database. BACKGROUND
[0002] As a database system specially used for storing and querying graph-structured data, the graph database has gradually risen in the field of data processing in recent years. The graph database represents entities and their relationships in the form of nodes and relationship edges, providing an efficient solution for complex data correlation analysis. Through the graph database, users can intuitively see the correlation between data, quickly locate key information, and conduct in-depth analysis. This intuitive and in-depth analysis capability makes the graph database occupy an important position in big data processing.
[0003] In the fields of deep learning, natural language processing, etc., it is often necessary to convert text, image, etc. data into high-dimensional embedding vectors in order to conduct more accurate semantic analysis and similarity comparison. However, the traditional graph database cannot effectively process vector data, which in turn leads to the inability to effectively support similarity search for high-dimensional embedding vectors, which greatly reduces the efficiency of the graph database when vector similarity search is required. Therefore, a method is needed to enable the graph database to support processing of vector data while supporting similarity search based on vectors. SUMMARY
[0004] One or more embodiments of the present specification describe a query processing method and apparatus in a graph database, which adds support for vectors and vector indexes in the graph database by designing a special data structure, thereby improving the efficiency of the graph database in the field of machine learning.
[0005] In a first aspect, a query processing method in a graph database is provided, the graph database storing attribute data of nodes and edges, the attribute data including attribute values in the form of vectors; the method comprising:
[0006] obtaining a target attribute value in the form of a vector according to a query request, the target attribute value belonging to a target attribute item;
[0007] determining a target vector index, the target vector index being constructed according to each vector under the target attribute item;
[0008] determining a number of nodes or edges with a high ranking of similarity to the target attribute value according to the target attribute value and the target vector index.
[0009] In some possible implementations, the query request includes the target attribute item and the target attribute value.
[0010] In some possible implementation manners, the first identifier and the target attribute item are included in the query request; and the target attribute value in vector form is obtained according to the query request, including:
[0011] The attribute data is queried according to the first identifier, and an attribute value of a node or an edge corresponding to the first identifier for the target attribute item is obtained as the target attribute value.
[0012] In some possible implementation manners, the attribute data is stored in the form of a key-value pair; the attribute data is queried according to the first identifier, and an attribute value of a node or an edge corresponding to the first identifier for the target attribute item is obtained, including:
[0013] The attribute data is queried with the first identifier as a key, and first attribute value data corresponding to a node or an edge with the first identifier is obtained, where the first attribute value data includes a plurality of attribute values corresponding to a plurality of attribute items;
[0014] The attribute value for the target attribute item is determined in the first attribute value data according to the target attribute item.
[0015] In some possible implementation manners, the graph database further stores type data of nodes / edges; the type data shows a plurality of attribute items included in each node type / each edge type; and the attribute value for the target attribute item is determined in the first attribute value data according to the target attribute item, including:
[0016] The node type / edge type corresponding to the first identifier is determined according to the first identifier, and a plurality of first attribute items corresponding to the node type / edge type are obtained by querying the type data;
[0017] The target attribute value corresponding to the first position is determined in the first attribute value data according to a first position of the target attribute item in the plurality of first attribute items; the plurality of first attribute items and the first attribute value data are in an ordered sequence and one-to-one correspondence.
[0018] In some possible implementation manners, the attribute data further includes attribute values in a non-vector form; the attribute values in the vector form are stored in the attribute data in a first form, and the attribute values in the non-vector form are stored in the attribute data in a second form; and the target attribute value corresponding to the first position is determined in the first attribute value data, including:
[0019] A first value located at the first position in the first attribute value data is determined, and the first value is read in a first manner as the target attribute value; the first manner is used to read the attribute value stored in the first form.
[0020] In some possible implementation manners, the graph database further stores index information data, in which construction manner information of a plurality of vector indexes is recorded; the target vector index is determined, including:
[0021] target index information corresponding to the target attribute item is queried from the index information data, and construction manner information of a vector index corresponding to each vector under the target attribute item is shown;
[0022] the target vector index is constructed according to the target index information.
[0023] In some possible implementation manners, the index information data is stored in the form of a key-value pair; the target index information corresponding to the target attribute item is queried from the index information data, including:
[0024] the index information data is queried with the target attribute item as the key, and the target index information is obtained.
[0025] In some possible implementation manners, the target vector index is constructed according to the target index information, including:
[0026] a plurality of nodes or edges related to a plurality of vectors under the target attribute item are retrieved from the attribute data;
[0027] the target vector index is constructed based on the plurality of vectors according to the target index information.
[0028] In some possible implementation manners, after the target vector index is constructed, the method further includes:
[0029] the target vector index is stored corresponding to the target attribute item.
[0030] In some possible implementation manners, the target index information includes: vector dimension, vector distance type, vector index category, and vector index parameter.
[0031] In some possible implementation manners, the method further includes:
[0032] a creation request is acquired, and the creation request includes the target attribute item and the target index information;
[0033] the target attribute item and the target index information are stored in the index information data.
[0034] In some possible implementation manners, the plurality of nodes or edges include a plurality of nodes; the query request further includes a first edge type; and the method further includes:
[0035] a plurality of target nodes connected to the plurality of nodes by the first edge type are determined in the graph database.
[0036] In a second aspect, a query processing apparatus in a graph database is provided. The graph database stores attribute data of nodes and edges, and the attribute data includes attribute values in vector form. The apparatus includes:
[0037] An obtaining unit configured to obtain a target attribute value in vector form according to a query request, the target attribute value belonging to a target attribute item;
[0038] A vector index determining unit configured to determine a target vector index, the target vector index being constructed according to each vector under the target attribute item;
[0039] A determining unit configured to determine a plurality of nodes or edges with a high ranking of similarity to the target attribute value according to the target attribute value and the target vector index.
[0040] In a third aspect, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. When the computer program is executed in a computer, the computer program causes the computer to execute the method in the first aspect.
[0041] In a fourth aspect, a computing device is provided, and the computing device includes a memory and a processor. The memory stores executable code, and the processor executes the executable code to implement the method in the first aspect.
[0042] The query processing method and apparatus in a graph database provided by the embodiments of the present disclosure can realize the access function of attribute values in vector form in the graph database, and support the similarity search based on the attribute values in vector form in the graph database, thereby enhancing the function of the graph database and expanding the application range of the graph database.
[0043] The query processing method and apparatus in a graph database provided by the embodiments of the present disclosure can realize the access function of attribute values in vector form in the graph database, and support the similarity search based on the attribute values in vector form in the graph database, thereby enhancing the function of the graph database and expanding the application range of the graph database. BRIEF DESCRIPTION OF DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in the present specification, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only a part of the embodiments disclosed in the present specification, and other drawings can be obtained by those skilled in the art without creative effort.
[0045] Figure 1 An architecture diagram of a graph database according to one embodiment is shown.
[0046] Figure 2 A data structure diagram of index information data according to one embodiment is shown.
[0047] Figure 3 A flow chart of a query processing method in a graph database according to one embodiment is shown.
[0048] Figure 4 A schematic block diagram of a query processing apparatus in a graph database according to one embodiment is shown. DETAILED DESCRIPTION
[0049] The schemes provided in the present specification will be described below with reference to the accompanying drawings.
[0050] A conventional graph database uses nodes to represent entities, attributes on the nodes to record specific data of the entities, and relationship edges between the nodes to record relationships between the entities, and attributes on the relationship edges to record specific data of the relationships, when saving graph data and graph structures. For example, a user node can have attributes such as a frequently purchased commodity type and an active time, and two user nodes can have a relationship edge representing a friend relationship, and the friend relationship can have attributes such as a friend-adding date. Such a graph structure containing node attributes and relationship edge attributes (hereinafter also referred to simply as “edges”) is generally referred to as a property graph, and the data types of the attributes therein are generally conventional data types such as a string type, a Boolean type, an integer type, and a floating-point number type, but do not support vector-type node attributes / relationship edge attributes.
[0051] However, for many complex application scenarios, it is often not enough to rely only on the information of the graph structure in the conventional graph database. For example, in a recommendation system, if a target user is to be recommended a commodity, a page, or an activity, it is often necessary to find other users similar to the target user based on the interest preferences and behavior patterns of the target user, and then make corresponding recommendations to the target user according to the frequently purchased commodities of the other users. Information such as interest preferences and behavior patterns is difficult to represent in conventional data types, and is usually better represented in the form of embedded vectors. Therefore, if a graph database cannot support vector data, it will greatly limit its application in these fields.
[0052] In the related art, two sets of databases are often used, one for storing graph structure data and the other for storing vector data, and then joint analysis is performed based on the two sets of databases. However, such a method has the difficulty of data migration and synchronization between databases, and the use of two sets of databases also increases the operating cost. In addition, from the perspective of data analysis, storing graph structure and vector data separately is not conducive to comprehensive analysis. In many cases, graph structure and vector data are related to each other, and separating them for processing may cause information loss or inaccurate analysis results.
[0053] To solve the above problems, the embodiments of the present specification propose a query processing method in a graph database, which adds support for vector data to the graph database, meets the needs of high-dimensional embedding vector similarity search for the graph database, and further meets the needs of users for comprehensive data analysis containing vector data and non-vector data, improves the efficiency and accuracy of data processing, and also reduces the operating cost.
[0054] First, the architecture of the graph database extended by the embodiments of the present specification to support vector attributes is introduced. Figure 1 The architecture of the graph database according to one embodiment is shown. As shown in Figure 1 The graph database at least includes a computing engine and a storage layer, the computing engine includes a plurality of program API interfaces, and a vector index framework. The storage layer at least includes a vector index manager, a vector index library, and stores graph structure data, attribute data, type data, and index information data in the data storage layer.
[0055] Vector indexing refers to a specific data structure constructed on a vector set to support efficient nearest neighbor (Nearest Neighbors, NN) or approximate nearest neighbor (Approximate Nearest Neighbor, ANN) queries. Vector indexing is based on vector indexing algorithms, such as IVF_FLAT algorithm (inverted file flat algorithm), HNSW algorithm (Hierarchical Navigable Small World, hierarchical navigable small world algorithm), etc. The algorithm constructs a corresponding data structure, i.e. a corresponding vector index, for multiple vectors in the vector set. By using vector indexing, similar vectors of a specific vector can be quickly queried.
[0056] The program API interfaces of the computing engine can include a plurality of API interfaces related to operating node attributes / edge attributes and operating vector indexes. For example, the creation, deletion, modification, and query interfaces for node attributes / edge attributes, and the creation, deletion, query, and use interfaces for vector indexes. These API interfaces can be directly called by users.
[0057] The vector indexing framework contains multiple interfaces that implement various functions for operating vector indexes, including adding, deleting, constructing, saving, loading, and using vector indexes.
[0058] The vector index manager in the storage layer is responsible for controlling the vector indexes, directly manipulating the vector indexes in the vector index library, and ensuring the real-time performance of the vector indexes.
[0059] The vector index library is used to store the pre-built vector indexes corresponding to each node attribute / edge attribute.
[0060] Graph data structures are used to store the topological data of a graph, including the connections between nodes and edges. Graph data structures can be stored based on forms such as adjacency lists or adjacency matrices; this is not a limitation.
[0061] Attribute data is used to store the specific attribute data of nodes / edges, including attribute data in vector format and attribute data in non-vector format.
[0062] Type data is used to store the pre-configured node / edge types and a list of attributes for each node / edge type. In the attribute graph, node / edge types are represented by type labels, and each type label contains several node / edge attributes. For example, in a transaction graph, node types can include user types and product types. User type nodes can have attributes such as interest preference, behavior pattern, and number of historical orders; product type nodes can have attributes such as name, price, and brand. Edge types can include purchase types and browsing types. Purchase type edges can have attributes such as purchase time and payment method; browsing type edges can have attributes such as browsing time and dwell time.
[0063] Index information data includes information used to define or describe how the vector index is constructed, including vector dimensions, vector distance type, vector index category, vector index parameters, and so on.
[0064] based on Figure 1 The database architecture shown can add support for vector attributes and vector indexes for nodes / edges in the graph database.
[0065] The following describes the specific methods for adding support for vector attributes of nodes / edges in a graph database.
[0066] For each type of node / edge in the graph data, firstly, create a type label corresponding to the node / edge type and its contained attributes, which may include vector type attributes. Then, jointly store the type label and its contained attributes in the type data.
[0067] In one embodiment, the data structure of the type data is a key-value pair, which stores the type label and its contained attributes in the type data jointly, including: taking the label identification of the type label as the key, taking its contained attributes as the value, constructing the key-value pair, and storing it in the type data.
[0068] For example, for the node of the commodity type, taking its type label "commodity" as the key, taking its contained attributes (commodity name, size, style, price, region, whether participating in the discount) as the value, constructing the key-value pair: "commodity"-(commodity name, size, style, price, region, whether participating in the discount), and storing the key-value pair in the type data.
[0069] The specific attribute values of each specific node / edge are stored in the attribute data.
[0070] In one embodiment, the data structure of the attribute data is a key-value pair. When constructing the attribute data, the identification (ID) of the node / edge can be taken as the key, the attribute value of the corresponding node / edge (which includes the attribute value of the node / edge under each attribute) can be taken as the value, the key-value pair can be constructed, and stored in the attribute data.
[0071] In the attribute value of the node / edge, the vector form attribute value and the non-vector form attribute value can be stored in different forms. Specifically, the vector form attribute value can be stored in a first form, and the non-vector form attribute value can be stored in a second form. Correspondingly, when reading the attribute value of the node / edge, the vector attribute value stored in the first form is read in a first manner, and the non-vector attribute value stored in the second form is read in a second manner.
[0072] In a more specific embodiment, the attribute value of the node / edge can be stored in the form surrounded by parentheses (). Among them, the non-vector form attribute value can be directly stored, and the vector form attribute value can be stored in the form surrounded by square brackets [], and each attribute value is separated by a comma.
[0073] For example, a node for a product type might have the following attribute list: (product name, size, style, price, location, whether it participates in a discount). Style and location can be represented as embedding vectors, i.e., they are vector-based attributes, while the other attributes are non-vector-based. Correspondingly, the node attribute value Att-N1 of a specific product node N1 could be ("chair", 30.5, [1, 0, 0, 1], 20, [1, 1, 0, 0], 1). The non-vector-based attribute values include the string value "chair" (the attribute value for the "product name" attribute), the floating-point value 30.5 (the attribute value for the "size" attribute), the integer value 20 (the attribute value for the "price" attribute), and the boolean value 1 (the attribute value for the "whether it participates in a discount" attribute). The vector-based attribute values include the style embedding vector [1, 0, 0, 1] and the location embedding vector [1, 1, 0, 0].
[0074] Reading the vector attribute value stored in the first form in the first manner can be done by reading the vector jointly determined by the values enclosed in square brackets [] in the node attribute value.
[0075] The above describes the specific methods for adding support for vector attributes of nodes / edges in a graph database. Next, we will describe the specific methods for adding support for vector indexes in a graph database.
[0076] A vector index is a data structure built for the attributes of nodes / edges of a specific vector type. Since nodes / edges of different types may have attributes with the same name—for example, a node of type "product" and a node of type "user" might both have an attribute named "style"—but the style of a product and the style of a user actually represent different meanings and cannot be used interchangeably. Therefore, when constructing a vector index, both the corresponding type label and attribute name must be explicitly specified to ensure uniqueness. Any combination of type label and attribute name for a node / edge is referred to as an attribute item in the embodiments of this specification. For example, the aforementioned "product-style" can constitute an attribute item, and "user-style" can also constitute an attribute item. An attribute item can uniquely correspond to a specific attribute of a node / edge of a specific type.
[0077] For constructing a vector index, the system first receives index information used to build the index, including vector dimensions, vector distance type, vector index category, vector index parameters, etc., and stores the index information and the aforementioned attribute items together in the index information data. The index information can be set by the user according to their actual needs.
[0078] In one embodiment, the data structure of the index information data is a key-value pair, taking the attribute item to which the vector index is directed as the key, and taking the index information as the value, constructing the key-value pair and storing it in the index information data. Figure 2 A data structure diagram of index information data according to one embodiment is shown. The storage form of the key-value pair corresponding to any index information can be: "attribute item" - "vector dimension, vector distance type, vector index category, vector index parameter".
[0079] Among them, the vector dimension can be the dimension of each vector in the vector set corresponding to the vector index; the vector distance type can be the measurement method used to measure the distance between vectors, such as Euclidean distance (L2 distance), inner product distance, etc.; the vector index category can be the corresponding vector index algorithm, such as IVF_FLAT algorithm, HNSW algorithm, etc.; the vector index parameter can be the parameter required when using the related vector index algorithm. For example, for the IVF_FLAT algorithm, the vector index parameter can include the number of partitions created using the k-means algorithm ivf_flat_nlist; for the HNSW algorithm, the vector index parameter can include the maximum number of edges or connections that each node in the data structure can have at each level of the hierarchy hnsw_m, and the number of candidate nodes considered in the index construction process hnsw_ef_construction.
[0080] For example, for the "style" attribute of the above-mentioned "commodity" type node, the combination constitutes the attribute item "commodity-style", and the key-value pair corresponding to one index information can be: "commodity-style" - "128, L2, hnsw, {'hnsw_m':32, 'hnsw_ef_construction': 200}".
[0081] The data structure of the above-mentioned type data, attribute data and index information data can also use other forms, such as based on relational data and stored in a relational database, such as stored in a SQL database.
[0082] In constructing the vector index, first, the attribute item corresponding to the vector index needs to be determined, and at the same time, the index information used to construct the vector index is also determined. The index information can be determined based on the corresponding API interface accepting user input on site, or can be obtained by querying the above-mentioned index information data based on the type label and attribute of the user input attribute item.
[0083] Then, the nodes / edges of the node / edge type belonging to the type label are obtained by querying the attribute data, and the attribute values of each node / edge about the attribute are obtained, each attribute value being in vector form.
[0084] Next, for the vector set composed of each attribute value queried, the vector index corresponding to the vector index category recorded in the index information is used to construct the corresponding vector index using the vector index construction algorithm. The vector index can be used to query a number of similar nodes or edges in the graph database with respect to the node / edge attribute of the to-be-queried node / edge belonging to the node / edge category, for the attribute value of the to-be-queried node / edge.
[0085] For example, for the "style" node attribute of the "commodity" type label, first, the node type "commodity" of each node is queried in the attribute data, and then the attribute value of the "style" attribute of each "commodity" node is queried. Each attribute value is in vector form and can form a vector set. Next, according to the vector set composed of each attribute value, the HNSW algorithm recorded in the index information is used to construct the corresponding vector index based on the vector dimension, vector distance type and vector index parameter setting of the HNSW algorithm in the index information.
[0086] After the vector index is constructed, the vector index and its corresponding attribute item can also be stored in the vector index library.
[0087] In an embodiment, the data structure of the vector index library is a key-value pair, the attribute item is taken as the key, the vector index is taken as the value, the key-value pair is constructed, and stored in the vector index library. It can be directly read during subsequent use without re-construction.
[0088] In addition, the index information and its corresponding attribute item can also be stored in the index information data.
[0089] The above describes the construction and storage process of the vector index in the graph database. In some embodiments, the corresponding vector index can also be queried and deleted in the vector index library according to the attribute item, which is not described here.
[0090] Corresponding to the above vector index, the program API interface of the computing engine can include at least the following interfaces for operating the vector index: creating a vector index addVectorIndex(), deleting a vector index deleteVectorIndex(), querying a vector index ShowVectorIndex(), limiting the number of similarity retrieval vertexVectorKnnSearch(), and limiting the distance of similarity retrieval vertexVectorRangeSearch().
[0091] The parameters received by the create vector index addVectorIndex() include a type label label name, an attribute field name, a vector dimension dimension, a vector distance type distance type, a vector index type index type, and a vector index parameter index spec. The create vector index addVectorIndex() creates a corresponding vector index according to the received parameters.
[0092] The parameters received by the delete vector index deleteVectorIndex() include a type label label name, an attribute field name, a vector index type index type, and a vector dimension dimension. The delete vector index deleteVectorIndex() deletes a corresponding vector index according to the received parameters.
[0093] The show vector index ShowVectorIndex() is used to return a list of currently existing vector indexes.
[0094] The parameters received by the limited number of similarity retrieval vertexVectorKnnSearch() include a type label label name, an attribute field name, a to-be-queried attribute vector vector, a number of returned results topk, and a vector index parameter query spec. The limited number of similarity retrieval vertexVectorKnnSearch() is used to query topk nodes or edges with a high similarity ranking to the to-be-queried attribute vector vector.
[0095] The parameters received by the limited distance similarity retrieval vertexVectorRangeSearch() include a type label label name, an attribute field name, a to-be-queried attribute vector vector, a similar distance range radius, and a vector index parameter query spec. The limited distance similarity retrieval vertexVectorRangeSearch() is used to query each node or edge with a distance less than the radius to the to-be-queried attribute vector vector.
[0096] By using the graph database extended to support vector attributes, the access to node / edge vector attributes and the similarity node / edge query based on vector indexes can be met.
[0097] The following describes the specific implementation steps of the query processing method in the graph database in combination with specific embodiments.
[0098] Figure 3 A flowchart of a query processing method in a graph database according to an embodiment is shown. The execution subject of the method can be any platform or server or device cluster with computing and processing capabilities. As shown in FIG. 1, the method includes the following steps. Figure 3As shown, the graph database stores attribute data of nodes and edges, the attribute data including attribute values in vector form; the method comprises at least: step S302, obtaining a target attribute value in vector form according to a query request, the target attribute value belonging to a target attribute item; step S304, determining a target vector index, the target vector index being constructed according to each vector under the target attribute item; step S306, determining a number of nodes or edges with a top ranking similarity with the target attribute value according to the target attribute value and the target vector index.
[0099] The graph data can be related data of multiple fields. In an embodiment, the graph data is related data of a recommendation system, nodes in the graph data including user nodes and commodity nodes, and edges in the graph data including friend relationships between user nodes, browsing relationships and purchase relationships between user nodes and commodity nodes, and same-category commodity relationships between commodity nodes. Attributes of the user nodes at least include an interest preference attribute, and the interest preference attribute is a vector attribute; attributes of the commodity nodes at least include a commodity image attribute, and the commodity image attribute is a vector attribute. Edge types can include a purchase type and a browsing type, edges of the purchase type can have a purchase time attribute, a payment method attribute, and the like; edges of the browsing type can have a browsing time attribute, a stay time attribute, and a potential purchase willingness attribute, and the potential purchase willingness attribute is a vector attribute.
[0100] The specific execution processes of the above steps are described below.
[0101] First, in step S302, a target attribute value in vector form is obtained according to a query request, the target attribute value belonging to a target attribute item.
[0102] The target attribute item can be an attribute item of a node or an attribute item of an edge, which is not limited herein. The target attribute item can include a type label and an attribute. The target attribute value can be an attribute value of a target attribute in the target attribute item.
[0103] For example, the target attribute item can be “commodity-style” above, and the target attribute value can be an attribute value of the “style” type.
[0104] In an embodiment, the query request in step S302 includes the target attribute item and the target attribute value.
[0105] In the embodiment, the target attribute item and the target attribute value can be directly included in the query request, for example, input from a user. In a specific embodiment, the user inputs the target attribute item to be queried and the original attribute value of the target attribute item, which is in a non-vector form. The input original attribute value is first encoded by a pre-trained encoder to obtain a vector type target attribute value. Each vector under an attribute item of any vector attribute in the graph database comes from the encoding result of the same encoder, so that each vector obtained by encoding is located in the same embedding space.
[0106] In another embodiment, the query request in step S302 includes the first identifier and the target attribute item. Step 302 specifically includes: querying the attribute data according to the first identifier to obtain the attribute value of the node or edge corresponding to the first identifier for the target attribute item as the target attribute value.
[0107] The first identifier can be the identifier of the node or edge to be queried. The attribute data includes the specific attribute value data of each node / edge in the graph database. The attribute value data corresponding to the node or edge can be found in the attribute data through the first identifier, and the target attribute value corresponding to the target attribute item can be obtained by querying.
[0108] In an embodiment, the attribute data in step S302 is stored as a key-value pair; and step S302 specifically includes steps 11 and 12.
[0109] In step 11, the attribute data is queried with the first identifier as the key to obtain the first attribute value data corresponding to the node or edge with the first identifier, which includes a plurality of attribute values corresponding to a plurality of attribute items.
[0110] The first identifier and the first attribute value data of the node or edge to be queried are stored in the attribute data in the form of a key-value pair. According to the first identifier, the corresponding first attribute value data can be queried from the attribute data, which includes the attribute values corresponding to each attribute item of the node or edge.
[0111] For example, continuing the above example, the object to be queried is a node, which has a first identifier N1. According to the first identifier, the first attribute value data (chair, 30.5, [1, 0, 0, 1], 20, [1, 1, 0, 0], 1) can be obtained from the attribute data.
[0112] Then, in step 12, the attribute value for the target attribute item is determined in the first attribute value data according to the target attribute item.
[0113] The attribute values of the node or edge to be queried can be arranged in a specific order and stored in the first attribute value data. Thus, the corresponding attribute value can be queried according to the position of the target attribute item in the first attribute value.
[0114] In a more specific embodiment, the graph database further stores type data of nodes / edges; the type data shows a plurality of attribute items included in each node type / edge type, wherein the first identification identifies each attribute item included in the corresponding node or edge. In this embodiment, step 12 includes step 121 and step 122.
[0115] In step 121, the node type / edge type corresponding to the first identification is determined according to the first identification, and then the corresponding first attribute items are queried in the type data.
[0116] The node type / edge type corresponding to the first identification can be determined in various ways. In one embodiment, the node type / edge type can be written into the first node identification, and then the node type / edge type can be directly read from the first node identification. In another embodiment, the node type / edge type corresponding to each node / edge can be stored in a specific key-value pair data structure, and the node type / edge type corresponding to the first identification can be queried from the key-value pair data structure, which is not limited herein.
[0117] According to the node type / edge type corresponding to the first identification, the corresponding first attribute items can be queried in the type data. In one embodiment, the type data is stored as a key-value pair, wherein each node type / edge type and the attribute items included therein are saved. According to the node type / edge type corresponding to the first identification, the corresponding first attribute items can be queried in the type data.
[0118] For example, continuing the above example, the node type of the node to be queried can be "goods". According to the node type, the first attribute items of the node to be queried are queried from the type data, including (goods name, size, style, price, region, and whether to participate in the preferential treatment).
[0119] Then, in step 122, the target attribute value corresponding to the first position is determined in the first attribute value data according to the first position of the target attribute item in the first attribute items; the first attribute items and the first attribute value data are in an ordered sequence and one-to-one correspondence.
[0120] Since each specific attribute value in the first attribute items and the first attribute value data has a one-to-one correspondence, the target attribute value at the corresponding first position can be read in the first attribute value data according to the first position of the target attribute item in the first attribute items.
[0121] In some possible implementations, the first attribute value data can include attribute values in non-vector form in addition to attribute values in vector form such as the target attribute item. Both are stored in different forms in the first attribute value data and need to be read in different ways.
[0122] Based on this, in one embodiment, the attribute data further includes attribute values in non-vector form; the attribute values in vector form are stored in the attribute data in a first form, and the attribute values in non-vector form are stored in the attribute data in a second form. In this embodiment, the step 122 of determining the target attribute value corresponding to the first position in the first attribute value data includes:
[0123] determining a first value located at the first position in the first attribute value data, and reading the first value in a first way as the target attribute value; the first way is used to read attribute values stored in the first form.
[0124] The first form of storage can be to store the attribute values in vector form in square brackets, and the second form of storage can be to directly store the attribute values in non-vector form. Correspondingly, the first way of reading can be to read the data surrounded by square brackets at the first position.
[0125] For example, continuing the above example, the first position of the target attribute item in the first attribute item is the third position, and the target attribute value is obtained by reading the data surrounded by square brackets at the third position in the first attribute value data (“chair”, 30.5, [1, 0, 0, 1], 20, [1, 1, 0, 0], 1).
[0126] In other embodiments, the attribute data can also be stored as relational data, for example, can be stored in a relational database. In this embodiment, a corresponding relational database query language can be used to query the target attribute value corresponding to the first identifier in the attribute data.
[0127] In other embodiments, the type data can also be stored as relational data, for example, can be stored in a relational database. In this embodiment, a corresponding relational database query language can be used to query the first attribute item corresponding to the node type / edge type in the type data.
[0128] After obtaining the target attribute item and the target attribute value, next, in step S304, a target vector index is determined, which is constructed according to each vector under the target attribute item.
[0129] Each vector under the target attribute item can be a vector form of an attribute value of each node / edge having the target attribute item in the graph data with respect to the target attribute item.
[0130] Each vector under the target attribute item can be a vector form of an attribute value of each node / edge having the target attribute item in the graph data with respect to the target attribute item.
[0131] In an embodiment, the graph database further stores index information data recording construction manner information of a plurality of vector indexes. In this embodiment, step S304 comprises steps 21 to 22.
[0132] In step 21, the target index information corresponding to the target attribute item is queried from the index information data, which shows the construction manner information of the vector indexes corresponding to each vector under the target attribute item.
[0133] In an embodiment, the index information data is stored as a key-value pair. In this embodiment, step 21 specifically comprises: querying the index information data with the target attribute item as the key to obtain the target index information.
[0134] Then, in step 22, the target vector index is constructed according to the target index information.
[0135] For example, continuing the above example, the target attribute item can be “commodity-style”, and the corresponding target index information can be “128, L2, hnsw, {'hnsw_m': 32, 'hnsw_ef_construction': 200}”.
[0136] The target vector index corresponding to the target attribute item can be constructed according to the target index information corresponding to the target attribute item and each vector under the target attribute item.
[0137] Specifically, step 22 comprises steps 221 and 222.
[0138] In step 221, a plurality of vectors of each node or edge with respect to the target attribute item are retrieved from the attribute data.
[0139] In step 222, the target vector index is constructed based on the plurality of vectors according to the target index information.
[0140] The target vector index can be constructed using a preset vector index algorithm according to the vector set determined by each node or edge with respect to the plurality of vectors under the target attribute item. Various types of vector index algorithms can be used, such as IVF_FLAT algorithm, HNSW algorithm, etc., which are not limited here.
[0141] In some embodiments, after step 222, the step 22 further includes:
[0142] Step 223, store the target vector index corresponding to the target attribute item.
[0143] The storage of step 223 can be based on key-value pairs or based on a relational database, which is not limited here.
[0144] After storing the target vector index, the corresponding vector index can be read directly based on the attribute item subsequently.
[0145] Based on this, in another embodiment, step S304 includes: reading the pre-constructed target vector index according to the target attribute item.
[0146] In one embodiment, the target index information of the foregoing various embodiments includes: vector dimension, vector distance type, vector index category, and vector index parameter.
[0147] In some possible implementations, the method further includes:
[0148] Obtaining a creation request, which includes a target attribute item and target index information; and storing the target attribute item and target index information into index information data.
[0149] After pre-storing the target index information, the target index information can be read from the index information data according to the target attribute item subsequently.
[0150] Returning to Figure 3 , finally, in step S306, a number of nodes or edges with a high similarity ranking are determined according to the target attribute value and the target vector index.
[0151] The number of nodes or edges with a high similarity ranking of step S306 can be k nodes or k edges with a preset number, which are determined by the similarity search vertexVectorKnnSearch() with a limited number of the foregoing interface; or each node or each edge with a distance less than a preset distance radius, which are determined by the similarity search vertexVectorRangeSearch() with a limited distance.
[0152] Through the above steps S302 to S306, a number of similar nodes / similar edges of the node / edge to be queried can be determined in the graph database supporting vector attributes and vector indexes. The overall system is improved, and the operation and maintenance cost is reduced.
[0153] In some possible implementation, when the query is for several similar nodes of the to-be-queried node, the specific nodes connected by each of the similar nodes of the to-be-queried node can also be queried. In this embodiment, the first edge type is also included in the query request. After step S306, the method further includes:
[0154] determining, in the graph database, several target nodes each connected to the to-be-queried node by the first edge type.
[0155] The above steps can be performed simultaneously by using the graph topology data and the vector attribute data in the graph database to perform corresponding comprehensive analysis.
[0156] For example, in one embodiment, the to-be-queried node can be a specific user, and the target attribute item can be an interest preference attribute. Each similar node can be a plurality of users having similar interest preferences to the specific user. The relationship corresponding to the first edge type can be a purchase relationship, and the several target nodes can be goods or services purchased by each user. Further, the goods or services recommended to the specific user can be determined according to each target node.
[0157] According to the query processing method in the graph database of the embodiments of the present specification, the vector data and the graph database are stored in the same system, which can reduce the architecture complexity and operation and maintenance cost. And support local joint query and calculation of vector and graph structure at the same time, which can make full use of the index and query optimization mechanism of the graph database, significantly improve the performance of the composite query involving vector retrieval and graph traversal. And by sharing the same storage engine for vector data and graph data, update operations can be executed within the same transaction scope, so as to more easily ensure data consistency and atomicity. In addition, since the vector data and the graph data are processed in the same platform, the developer can more flexibly optimize the query mode, for example, directly embedding vector similarity calculation in the graph traversal process, or dynamically adjusting the vector retrieval strategy according to the graph structure.
[0158] According to another aspect, an embodiment of a query processing apparatus in a graph database is also provided. Figure 4 A schematic block diagram of a query processing apparatus in a graph database according to an embodiment is shown, which can be deployed in any device, platform or device cluster with computing and processing capability. The graph database stores attribute data of nodes and edges, and the attribute data includes attribute values in the form of vectors. As shown in the figure, Figure 4 The apparatus 400 is deployed in the computing engine and includes:
[0159] The obtaining unit 402 is configured to obtain, according to a query request, a target attribute value in the form of a vector, which belongs to a target attribute item;
[0160] The vector index determining unit 404 is configured to determine a target vector index, the target vector index being constructed according to each vector under the target attribute item.
[0161] The determining unit 406 is configured to determine, according to the target attribute value and the target vector index, a number of nodes or edges with a high similarity to the target attribute value.
[0162] In some possible implementation manners, attribute data of nodes and edges are stored in the graph database, and the attribute data includes attribute values in the form of vectors. The apparatus 400 further includes:
[0163] The second determining unit is configured to determine, in the graph database, a number of target nodes each connected to the number of nodes by the first edge type.
[0164] According to another aspect, an embodiment also provides a computer-readable storage medium having stored thereon a computer program, which, when executed in a computer, causes the computer to perform the method described in any of the embodiments.
[0165] According to another aspect, an embodiment also provides a computing device including a memory and a processor, wherein the memory has stored thereon executable code that, when executed by the processor, implements the method described in any of the embodiments.
[0166] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts of each of the embodiments can be referred to each other. Each of the embodiments mainly describes the difference from other embodiments. In particular, for the apparatus embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant part can be referred to the part of the method embodiment.
[0167] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous.
[0168] It is to be noted that, in the present text, the relative terms such as first and second, and the like are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between such entities or operations. Moreover, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or apparatus including the element.
[0169] Those of ordinary skill in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware or by programs instructing relevant hardware, and the programs can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0170] The above detailed description of the specific embodiments of the present application has further detailed the purposes, technical solutions and beneficial effects of the present application. It should be understood that the above detailed description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A query processing method in a graph database, wherein the graph database stores attribute data of nodes and edges, and the attribute data includes attribute values in vector form; The graph database also stores index information data, which records information on the construction method of several vector indexes; the method includes: Based on the query request, obtain the target attribute values in vector form, which belong to the target attribute items; A target vector index is determined, which is constructed based on the target index information corresponding to each vector under the target attribute item and the target attribute item; the target index information is obtained from the index information data, and it shows the construction method information of the vector index corresponding to each vector under the target attribute item. Based on the target attribute value and the target vector index, determine a number of nodes or edges that rank highly similar to the target attribute value.
2. The method according to claim 1, wherein, The query request includes the target attribute item and the target attribute value.
3. The method according to claim 1, wherein, The query request includes a first identifier and the target attribute item; Based on the query request, obtain the target attribute values in vector form, including: The attribute data is queried based on the first identifier to obtain the attribute value of the node or edge corresponding to the first identifier for the target attribute item, which is then used as the target attribute value.
4. The method according to claim 3, wherein, The attribute data is stored as key-value pairs; Querying the attribute data based on the first identifier to obtain the attribute value of the node or edge corresponding to the first identifier for the target attribute item includes: Using the first identifier as the key, query the attribute data to obtain the first attribute value data corresponding to the node or edge with the first identifier, which includes multiple attribute values corresponding to multiple attribute items; Based on the target attribute item, determine the attribute value for the target attribute item in the first attribute value data.
5. The method according to claim 4, wherein, The graph database also stores node / edge type data; the type data shows several attribute items included in each node type / edge type; Based on the target attribute item, determining the attribute value for the target attribute item in the first attribute value data includes: Based on the first identifier, determine the node type / edge type corresponding to the first identifier, and then query the corresponding first attribute items in the type data; Based on the first position of the target attribute item among the plurality of first attribute items, the target attribute value corresponding to the first position is determined in the first attribute value data; the plurality of first attribute items and the first attribute value data are ordered sequences and correspond one-to-one.
6. The method according to claim 5, wherein, The attribute data also includes attribute values in non-vector form; the attribute values in vector form are stored in the attribute data in a first form, and the attribute values in non-vector form are stored in the attribute data in a second form. Determining the target attribute value corresponding to the first position in the first attribute value data includes: Determine the first value located at the first position in the first attribute value data, and read the first value in a first manner as the target attribute value; The first method is used to read attribute values stored in a first form.
7. The method according to claim 1, wherein, Determine the target vector index, including: The target index information corresponding to the target attribute item is obtained by querying the index information data; Based on the target index information, construct the target vector index.
8. The method according to claim 7, wherein, The index information data is stored as key-value pairs; The target index information corresponding to the target attribute item is obtained by querying the index information data, including: Using the target attribute item as the key, query the index information data to obtain the target index information.
9. The method according to claim 7, wherein, Based on the target index information, a target vector index is constructed, including: Multiple vectors of multiple nodes or edges under the target attribute item are retrieved from the attribute data; The target vector index is constructed based on the target index information and the multiple vectors.
10. The method according to claim 7, wherein, After constructing the target vector index, the method further includes: The target vector index is stored in correspondence with the target attribute item.
11. The method according to claim 7, wherein, The target index information includes: vector dimension, vector distance type, vector index category, and vector index parameters.
12. The method of claim 7, further comprising: Retrieve the creation request, which includes target attribute items and target index information; The target attribute items and target index information are stored in the index information data.
13. The method according to claim 1, wherein, The plurality of nodes or edges includes a plurality of nodes; the query request also includes a first edge type; the method further includes: In the graph database, determine several target nodes that are each connected to the first edge type by the several nodes.
14. A query processing apparatus for a graph database, wherein the graph database stores attribute data of nodes and edges, and the attribute data includes attribute values in vector form; The graph database also stores index information data, which records information on the construction methods of several vector indexes; the device includes: The unit is configured to obtain the target attribute value in vector form, which belongs to the target attribute item, based on the query request. The vector index determination unit is configured to determine a target vector index, which is constructed based on the target index information corresponding to each vector under the target attribute item and the target attribute item; the target index information is obtained from the index information data and shows the construction method information of the vector index corresponding to each vector under the target attribute item. The determining unit is configured to determine, based on the target attribute value and the target vector index, a number of nodes or edges that rank highly similar to the target attribute value.
15. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to perform the method of any one of claims 1-13.
16. A computing device comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, it implements the method of any one of claims 1-13.
Citation Information
Patent Citations
Native vector diagram database and data structure construction method
CN118626655A