A Compression Index and Query Method for Attribute Graphs
By constructing the compression structure of the graph structure and attribute compression index, the problem of inefficiency of the existing graph database in large-scale graph data query is solved, and efficient query and storage of the attribute graph database is realized.
Patent Information
- Application Number
- CN202211020089.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-08-24
AI Technical Summary
Existing graph databases are inefficient in querying when processing large-scale graph data, and traditional index establishment methods are difficult to effectively improve query performance.
By constructing the compression structure of the graph structure and attribute compression index, the compressed storage and efficient query of the attribute graph are realized. The specific steps include building an adjacency table for each vertex, calculating the vertex/edge number, obtaining the positioning of the attribute in the attribute compression index, and extracting the attribute value from the index.
It significantly improves the query efficiency of the attribute graph database, reduces the data storage scale, and supports query operations for various types of attribute graphs.
Smart Images

Figure CN115495617B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a compression index and query method for an attribute graph. Background Art
[0002] With the rapid development of fields such as social networks and e-commerce, the application of attribute graph databases is becoming increasingly widespread. It supports efficient complex association relationship analysis, and the efficiency of processing complex and associated network data is much higher than that of traditional relational databases. However, graph data is usually complex and huge, which may cause the graph database to read a large amount of data during query operations, resulting in low query efficiency. Existing graph databases usually improve query efficiency in the following ways: 1. Store the graph structure in the form of an adjacency list. The so-called adjacency list means that each vertex in the attribute graph corresponds to a list. If there are adjacent vertices to this vertex, the adjacent vertices are stored in this list in sequence; by establishing a graph index through the adjacency list, the associated edges and vertices of the vertex can be obtained more quickly. 2. Use a B-tree or the like to establish an attribute index to accelerate the retrieval of vertices and edges on attributes.
[0003] However, the above methods for establishing graph or attribute indexes still have the problem of low query efficiency when the data scale is huge. Summary of the Invention
[0004] In order to solve the above problems existing in the prior art, the present invention provides a compression index and query method for an attribute graph. The technical problems to be solved by the present invention are realized through the following technical solutions:
[0005] An embodiment of the present invention provides a compression index and query method for an attribute graph, which is applied to an attribute graph database system. The attribute graph maintained by the attribute graph database system includes a plurality of vertices, edges, and their corresponding attributes. The method includes:
[0006] Construct an adjacency list for each vertex according to the attribute graph;
[0007] Construct a compressed structure of the graph structure according to the adjacency list of each vertex;
[0008] Calculate the vertex / edge number according to the compressed structure of the graph structure, obtain the location of the attribute in the attribute compression index according to the vertex / edge number and the corresponding bit array, and extract the attribute value from the attribute compression index; the attribute compression index includes a vertex attribute compression index and an edge attribute compression index; the bit array correspondingly records the starting positions of all edge attributes / vertex attributes;
[0009] Query the attribute graph according to the compressed structure of the graph structure and the attribute compression index;
[0010] Among them, the compressed structure of the graph structure includes an encoded sequence, the starting position of the adjacency list of each vertex in the encoded sequence, the total number of differences of the adjacency list of each vertex, and the first vertex number of the adjacency list of each vertex; the encoded sequence is obtained by calculating the differences between adjacent vertex numbers in the adjacency list and encoding the differences.
[0011] In an embodiment of the present invention, the compressed structure of the graph structure includes an out-neighbor compression structure and an in-neighbor compression structure.
[0012] In an embodiment of the present invention, accessing the adjacent vertex numbers corresponding to each vertex in the adjacency list through the compressed structure of the graph structure is expressed as:
[0013]
[0014] Among them, i represents the i-th vertex, Adj i [j] represents the j-th adjacent vertex number accessed by the i-th vertex through the compressed structure of the graph structure, Sam[i] represents the first vertex number of the adjacency list of the i-th vertex, Adj i [j - 1] represents the (j - 1)-th adjacent vertex number accessed by the i-th vertex through the compressed structure of the graph structure, decompress() represents the decoding function, the input parameters are S, pos, and nbits, S represents the encoded sequence, and nbits represents the number of decoding bits. X[i + 1] stores the starting position of the adjacency list of the (i + 1)-th vertex in the encoded sequence, X[i] stores the starting position of the adjacency list of the i-th vertex in the encoded sequence, Ngap[i + 1] stores the total number of differences to the adjacency list of the (i + 1)-th vertex, Ngap[i] stores the total number of differences to the adjacency list of the i-th vertex, and pos = X[i] + (j - 1) * nbits.
[0015] In an embodiment of the present invention, calculating the vertex number according to the compressed structure of the graph structure, obtaining the positioning of the vertex attribute in the vertex attribute compressed index according to the vertex number and the corresponding bit array, and extracting the vertex attribute value from the vertex attribute compressed index includes:
[0016] Establishing a vertex attribute high-order entropy compressed full-text index according to the vertex attribute set; the vertex attribute set includes the vertex attributes of all vertices;
[0017] Calculating the starting position and the ending position of the vertex number in the vertex attribute set according to the vertex number and the corresponding bit array;
[0018] Extracting the corresponding vertex attribute value from the vertex attribute high-order entropy compressed full-text index according to the starting position and the ending position of the vertex attribute.
[0019] In one embodiment of the present invention, edge numbers are calculated according to the compression structure of the graph structure, the positioning of edge attributes in the edge attribute compression index is obtained based on the edge numbers and the corresponding bit arrays, and the extraction of attribute values is performed from the edge attribute compression index, including:
[0020] An edge attribute high-order entropy compression full-text index is established according to the edge attribute set; the edge attribute set includes edge attributes between all adjacent vertices;
[0021] The corresponding edge numbers are obtained according to the vertex numbers;
[0022] The starting position and the ending position of the edge number in the edge attribute set are calculated according to the edge number and the corresponding bit array;
[0023] The corresponding edge attribute values are extracted from the edge attribute high-order entropy compression full-text index according to the starting position and the ending position in the edge attribute set.
[0024] In one embodiment of the present invention, the query of the attribute graph is performed according to the compression structure of the graph structure and the attribute compression index, including:
[0025] Initialize a given vertex number and a preset distance;
[0026] Convert the given vertex number into a given global number;
[0027] Enqueue the given global number into the FIFO queue and initialize the output result set;
[0028] Take out the intermediate vertex number from the FIFO queue, and judge whether the distance between the intermediate vertex number and the given global number is less than the preset distance. If it is less, find the unexplored adjacent vertex numbers corresponding to the given vertex number, enqueue the adjacent vertex numbers into the FIFO queue, and add the adjacent vertex numbers to the output result set. Repeat the above process of taking out from the FIFO queue until the FIFO queue is empty, and return the output result set;
[0029] Among them, the adjacent vertex numbers are obtained by accessing through the compression structure of the graph structure.
[0030] In one embodiment of the present invention, the query of the attribute graph is performed according to the compression structure of the graph structure and the attribute compression index, including:
[0031] Initialize a given vertex number, a preset distance and a preset filtering condition;
[0032] Convert the given vertex number into a given global number;
[0033] Enqueue the given global number into the FIFO queue and initialize the output result set;
[0034] Dequeue the intermediate vertex number from the FIFO queue, and determine whether the distance between the intermediate vertex number and the given global number is less than the preset distance. If it is less, then find the unexplored adjacent vertex numbers corresponding to the given vertex number, and check whether the intermediate vertex number and the adjacent vertex numbers satisfy the preset filtering condition. If they satisfy, enqueue the adjacent vertex numbers into the FIFO queue, and add the adjacent vertex numbers to the output result set. Repeat the above process of dequeuing from the FIFO queue until the FIFO queue is empty, and return the output result set;
[0035] Among them, the adjacent vertex numbers are obtained by accessing through the compressed structure of the graph structure;
[0036] Whether the intermediate vertex number and the adjacent vertex numbers satisfy the preset filtering condition is checked by the edge attribute values extracted by the edge attribute compression index.
[0037] In an embodiment of the present invention, querying the property graph according to the compressed structure of the graph structure and the attribute compression index includes:
[0038] Initialize the first given vertex number, the second given vertex number, and the preset filtering condition;
[0039] Determine whether the first given vertex number and the second given vertex number are the same vertex. If so, return the distance as 0. Otherwise, convert the first given vertex number and the second given vertex number into the first given global number and the second given global number; enqueue the first given global number into the FIFO queue, initialize the distance as 0, dequeue the intermediate vertex number from the FIFO queue, and determine whether the intermediate vertex number and the second given global number are the same vertex. If so, return the distance. Otherwise, find the unexplored adjacent vertex numbers corresponding to the first given vertex number, and check whether the intermediate vertex number and the adjacent vertex numbers satisfy the preset filtering condition. If they satisfy, enqueue the adjacent vertex numbers into the FIFO queue, and increment the distance by 1. Repeat the above process of dequeuing from the FIFO queue until the FIFO queue is empty, and return the distance;
[0040] Among them, the adjacent vertex numbers are obtained by accessing through the compressed structure of the graph structure;
[0041] Whether the intermediate vertex number and the adjacent vertex numbers satisfy the preset filtering condition is checked by the edge attribute values extracted by the edge attribute compression index.
[0042] In an embodiment of the present invention, querying an attributed graph according to the compression structure of the graph structure and the attribute compression index includes:
[0043] Initializing a given vertex number, given vertex attributes, a preset distance, and a preset filtering condition;
[0044] Converting the given vertex number into a given global number;
[0045] Enqueuing the given global number into a first FIFO queue and initializing an output result set;
[0046] Taking out a first intermediate vertex number from the first FIFO queue, determining whether the distance between the first intermediate vertex number and the given global number is less than the preset distance. If it is less, finding an unexplored first adjacent vertex number corresponding to the given vertex number and enqueuing it into the first FIFO queue, checking whether the first intermediate vertex number and the first adjacent vertex number satisfy the preset filtering condition. If they satisfy, finding the vertex attributes of the first adjacent vertex number, comparing whether the vertex attributes of the first adjacent vertex number and the given vertex attributes are of the same attribute. If they are of the same attribute, enqueuing the first adjacent vertex number into a second FIFO queue, taking out a second intermediate vertex number from the second FIFO queue, finding an unexplored second adjacent vertex number of the first adjacent vertex number, obtaining the relevant information of the second intermediate vertex number according to the label of the second adjacent vertex number, and adding the relevant information of the second intermediate vertex number into the output result set. Repeating the above process of taking out from the second FIFO queue until the second FIFO queue is empty, and continuing to repeat the above process of taking out from the first FIFO queue until the first FIFO queue is empty, and returning the output result set;
[0047] Wherein, both the first adjacent vertex number and the second adjacent vertex number are obtained by accessing through the compression structure of the graph structure;
[0048] Whether the first intermediate vertex number and the first adjacent vertex number satisfy the preset filtering condition is checked through the edge attribute value extracted by the edge attribute compression index;
[0049] The relevant information of the second intermediate vertex number is obtained through the vertex attribute value extracted by the vertex attribute compression index.
[0050] Advantages of the present invention:
[0051] The compression index and query method for the attribute graph proposed by the present invention no longer uses the traditional adjacency list storage and query method. Instead, in order to reduce the storage amount of data during storage, a new adjacency list storage and query method is proposed. Specifically: a compressed structure of the graph structure is constructed according to the adjacency list of each vertex, and the compressed structure of the constructed graph structure is used to access the adjacent vertex numbers corresponding to each vertex in the adjacency list, realizing the compressed storage of the attribute graph, so that the query efficiency of the attribute graph database is improved; at the same time, during the query process of the attribute graph database of the present invention, corresponding vertex attribute compressed indexes and edge attribute compressed indexes are also established for vertex attributes and edge attributes, supporting query operations on various types of attribute graphs on the compressed indexes of vertex and edge attributes, further improving the query efficiency of the attribute graph database.
[0052] The following will further elaborate on the present invention in conjunction with the accompanying drawings and embodiments. Brief Description of the Drawings
[0053] Figure 1 is a schematic flow chart of a compression index and query method for an attribute graph provided by an embodiment of the present invention;
[0054] Figure 2 is a schematic diagram of the graph structure of the attribute graph provided by an embodiment of the present invention;
[0055] Figure 3 is a schematic flow chart of extracting edge attributes provided by an embodiment of the present invention;
[0056] Figure 4 is a schematic flow chart of extracting vertex attributes provided by an embodiment of the present invention;
[0057] Figure 5 is a schematic flow chart of querying on the attribute graph according to the compressed structure of the graph structure provided by an embodiment of the present invention;
[0058] Figure 6 is a schematic flow chart of querying on the attribute graph according to the compressed structure of the graph structure and edge attributes provided by an embodiment of the present invention;
[0059] Figure 7 is a schematic flow chart of another querying on the attribute graph according to the compressed structure of the graph structure and edge attributes provided by an embodiment of the present invention;
[0060] Figure 8 is a schematic flow chart of querying on the attribute graph according to the compressed structure of the graph structure, vertex attributes and edge attributes provided by an embodiment of the present invention;
[0061] Figure 9 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed Embodiment
[0062] The present invention will be further described in detail below in conjunction with specific embodiments, but the embodiments of the present invention are not limited thereto.
[0063] To improve the query efficiency of the property graph database, please refer to Figure 1 , an embodiment of the present invention provides a compression index and query method for a property graph, which is applied to a property graph database system. The property graph maintained by the property graph database system includes a plurality of vertices, edges, and their corresponding attributes. The corresponding method includes the following steps:
[0064] S10. Construct an adjacency list for each vertex according to the property graph.
[0065] In the embodiment of the present invention, the adjacency list in S10 can be represented in the form of a traditional adjacency list, which is used here to determine the compressed structure of the graph structure constructed in the subsequent embodiments of the present invention.
[0066] As shown in FIG. 2, the property graph maintained in the property graph database system includes a graph structure composed of a plurality of vertices and edges. Each vertex has its corresponding vertex attributes, including out-neighboring vertices, in-neighboring vertices, and labels. The edges connecting adjacent vertices have their corresponding edge attributes. Among them, the out-neighboring vertices are determined by the out-edges corresponding to the vertex, and the in-neighboring vertices are determined by the in-edges corresponding to the vertex. For example, for the vertex with VID = 0, the other vertices with arrows pointing to the vertex with VID = 0 are its in-neighboring vertices, and the other vertices with arrows pointing from the vertex with VID = 0 are the out-neighboring vertices.
[0067] Taking the vertex with VID = 0 as an example, the vertex attributes of the vertex with VID = 0 include: the external adjacent vertices composed of the vertices with VID = 3, VID = 4, VID = 5, and VID = 7, and the incoming adjacent vertex composed of the vertex with VID = 1. The vertex label of the vertex with VID = 0 is director Tom.23, the vertex label of the vertex with VID = 1 is director Tim.34, the vertex label of the vertex with VID = 3 is film Ben, the vertex label of the vertex with VID = 4 is film Seas, the vertex label of the vertex with VID = 5 is student Lam.MIT, the vertex label of the vertex with VID = 7 is film Bear. The edge attribute of the edge connecting the vertex with VID = 0 and the vertex with VID = 1 is likes and the value is 90, likes and the value is 87. The edge attribute of the edge connecting the vertex with VID = 0 and the vertex with VID = 3 is guide and the date is from 1995 to 2000. The edge attribute of the edge connecting the vertex with VID = 0 and the vertex with VID = 4 is guide and the date is from 1993 to 1995. The edge attribute of the edge connecting the vertex with VID = 0 and the vertex with VID = 5 is likes and the value is 90. The edge attribute of the edge connecting the vertex with VID = 0 and the vertex with VID = 7 is guide and the date is from 2000 to 2010. Other vertices have the same vertex attributes and edge attributes to form a complete attribute graph structure as shown in Figure 2 shown.
[0068] S20. Construct a compressed structure of the graph structure according to the adjacency list of each vertex.
[0069] When directly storing by vertex number VID in the traditional adjacency list, if the amount of VID data is relatively large, there will be a problem of relatively large data storage scale. Due to the large data storage scale of the traditional adjacency list, the query efficiency is low. How to optimize the representation method of the adjacency list to reduce the data storage scale and then improve the query efficiency is a feasible solution. After research by the inventor, a compressed structure of the graph structure can be constructed according to the adjacency list of each vertex. The compressed structure of the graph structure in the embodiment of the present invention includes an encoding sequence, the starting position of the adjacency list of each vertex in the encoding sequence, the total number of differences of the adjacency list of each vertex, and the first vertex number of the adjacency list of each vertex; the encoding sequence is obtained by calculating the differences between adjacent vertex numbers in the adjacency list and encoding the differences.
[0070] As Figure 2 can be seen, for any vertex, there are external adjacent vertices and incoming adjacent vertices. According to the external adjacent vertices and incoming adjacent vertices, the compressed structure of the graph structure corresponding to the vertex can be constructed respectively. That is, the compressed structure of the graph structure in the embodiment of the present invention includes an external adjacent compression structure and an incoming adjacent compression structure.
[0071] For the external adjacent node query, the embodiments of the present invention propose an external adjacent compression structure, including S, X, Ngap, and Sam structures, which are used to access the numbers of adjacent vertices corresponding to each vertex in the adjacency list. S represents the set of encoded sequence obtained by encoding the differences between the numbers of adjacent vertices in the adjacency lists of all vertices, X represents the starting position of the adjacency list of each vertex in the encoded sequence, Ngap represents the total number of differences to the adjacency list of each vertex, and Sam represents the number of the first vertex in the adjacency list of each vertex. Taking Figure 2 the graph structure shown as an example, the adjacency list of each vertex is sorted according to the vertex number ( Figure 2 the vertex numbers in are all the preprocessed global numbers), and the process of constructing the external adjacent compression structure is specifically described. The constructed external adjacent compression structure is the data in the S, X, Ngap, and Sam structures shown in Table 1.
[0072] Table 1 External adjacent compression structure of each vertex
[0073]
[0074] In Table 1, the first row j represents the storage subscript of the external adjacent node element of each vertex in the traditional adjacency list. The storage of the external adjacent nodes of each vertex starts from 0. Adj i represents accessing the adjacency list of the i-th vertex, and gap represents the difference between the numbers of adjacent vertices.
[0075] For example, taking the vertex with VID = 0 as an example, the corresponding external adjacent nodes include the vertices with VID = 3, VID = 4, VID = 5, and VID = 7, which are the access results in Adj in Table 1. i For each vertex, starting from its first adjacent node, then:
[0076] Calculate the differences between the numbers of adjacent vertices in the adjacency list of each vertex, which are the data recorded in gap in Table 1. The vertices recorded in the adjacency list are the vertices with VID = 3, VID = 4, VID = 5, and VID = 7. Starting from the vertex with VID = 3, this point does not need to calculate the gap value. Then calculate the differences gap between the numbers of adjacent vertices in the adjacency list of each vertex as 1 (the difference between 3 and 4), 1 (the difference between 4 and 5), and 2 (the difference between 5 and 7);
[0077] Next, obtain the encoded sequence in the external adjacent compression structure according to the difference encoding. It can be seen from the differences that the maximum value is 2, and only two-bit encoding is required to represent all values. Encode all the differences to obtain the encoded sequence in the external adjacent compression structure. 1 is encoded as 01, 2 is encoded as 10, and the encoded sequence of the differences 1, 1, 2 is recorded as 010110;
[0078] Next, count the starting positions of the adjacency lists of each vertex in the coding sequence. The vertex with VID = 0 is the starting vertex. By default, the starting position X of the adjacency list of the vertex with VID = 0 in the coding sequence is recorded as 0, and the starting position X of the adjacency list of the next vertex with VID = 1 in the coding sequence is recorded as 6. It can be seen that the starting position X of the adjacency list of each vertex is the statistical value of the coding sequences of the adjacency lists of all vertices before this vertex.
[0079] Next, calculate the number of vertex numbers in the adjacency list of each vertex. The vertex with VID = 0 has 4 adjacent vertex numbers in the adjacency list. Then, according to the number of vertex numbers in the adjacency list, count the total difference number Ngap of the adjacency list of each vertex. The vertex with VID = 0 is the starting vertex, and the total difference number Ngap to the adjacency list of this vertex is recorded as 0, and the total difference number Ngap to the adjacency list of the next vertex with VID = 1 is recorded as 4. It can be seen that the total difference number Ngap to the adjacency list of each vertex is the statistical value of the number of adjacent vertices in the adjacency lists of all vertices before this vertex.
[0080] Next, record the first vertex number in the adjacency list of each vertex as Sam. The first vertex number in the adjacency list of the vertex with VID = 0 is the vertex with VID = 3. Then, determine the vertex with VID = 3 as the first vertex number Sam of the vertex with VID = 0 in the external neighbor compression structure.
[0081] Similarly, taking the vertex with VID = 1 as an example, the relevant data of the corresponding external adjacent points are the data in columns j = 0 to 4. It can be seen that finally Adj i The access results include the vertex with VID = 0, the vertex with VID = 0 (with different edge attributes from the first vertex with VID = 0), the vertex with VID = 2, the vertex with VID = 6, the vertex with VID = 6 (with different edge attributes from the first vertex with VID = 6). Taking the vertex with VID = 0 as the starting point, the corresponding differences gap are 0, 2, 4, 0. The maximum difference value is 4, which requires three-bit coding, that is, the coding sequence is 000010100000. The set formed by connecting the coding sequence 010110 of the vertex with VID = 0 and the coding sequence 000010100000 of the vertex with VID = 1 is recorded as the coding sequence S of all vertices. As known above, the starting position X of the adjacency list of the vertex with VID = 1 in the coding sequence is 6, and the corresponding starting position X of the adjacency list of the next vertex with VID = 2 in the coding sequence is 6 + 12 = 18. As known above, the total difference number Ngap to the adjacency list of the vertex with VID = 1 is 4, and the corresponding total difference number Ngap to the adjacency list of the next vertex with VID = 2 is 4 + 5 (5 is the number of external adjacent points of the vertex with VID = 1) = 9. The vertex with VID = 0 is the first vertex number Sam of the adjacency list of the vertex with VID = 1.
[0082] By analogy, the external neighbor compression structure of the embodiment of the present invention as shown in Table 1 is constructed.
[0083] Suppose Figure 2 If the total number of vertices in the attribute graph is n, the total number of edges between adjacent vertices is m, each element of X is represented by log|S| bits, each element of Ngap is represented by log m bits, and each element of Sam is represented by log n bits. Therefore, the encoding length of X is n×log|S|, the encoding length of Ngap is n×log m, and the encoding length of Sam is n×log n. It can be seen that the storage of the adjacency list can be implemented in a concise manner by using the S, X, Ngap, and Sam structures in the embodiment of the present invention.
[0084] For the in-neighbor vertex query, a compression method similar to the external neighbor compression structure is used to implement the compressed storage of the in-neighbor vertices as shown in Table 2. The process of constructing the in-neighbor compression structure will not be elaborated here.
[0085] Table 2 In-neighbor compression structure of each vertex
[0086]
[0087] Tables 1 and 2 clearly analyze how to achieve concise compressed storage of Figure 2 Convert the traditional storage method through the adjacency list into concise storage by constructing the compressed structure of the graph structure.
[0088] Furthermore, for the compressed storage method proposed in the embodiment of the present invention, here, the embodiment of the present invention also designs a method for accessing the adjacent vertex numbers corresponding to each vertex in the adjacency list through the compressed structure of the graph structure, which can be expressed by the formula:
[0089]
[0090] i Among them, i represents the i-th vertex, Adj i [j] represents the j-th adjacent vertex number accessed by the i-th vertex through the compressed structure of the graph structure in the adjacency list, Sam[i] represents the first vertex number of the adjacency list of the i-th vertex, Adj X[i + 1] stores the starting position of the adjacency list of the (i + 1)-th vertex in the coding sequence, X[i] stores the starting position of the adjacency list of the i-th vertex in the coding sequence, Ngap[i + 1] stores the total number of differences to the adjacency list of the (i + 1)-th vertex, Ngap[i] stores the total number of differences to the adjacency list of the i-th vertex, and pos = X[i] + (j - 1) * nbits.
[0091] As can be seen from Table 1 and Table 2, in the embodiment of the present invention, in the storage mode of the compression structure of the graph structure, the data storage scale is greatly reduced. When accessing data, the corresponding vertex can be quickly accessed through formula (1).
[0092] S30. Calculate the vertex / edge number according to the compression structure of the graph structure, obtain the location of the attribute in the attribute compression index according to the vertex / edge number and the corresponding bit array, and extract the attribute value from the attribute compression index; the attribute compression index includes a vertex attribute compression index and an edge attribute compression index; the bit array correspondingly records the starting positions of all edge attributes / vertex attributes.
[0093] The embodiment of the present invention also proposes a method for extracting attribute values by obtaining an attribute compression index based on the compression structure of the graph structure. Specifically, the attribute compression index includes a vertex attribute compression index and an edge attribute compression index, and the attribute values obtained by corresponding attribute extraction are vertex attribute values and edge attribute values.
[0094] For extracting vertex attribute values through the vertex attribute compression index, the embodiment of the present invention provides an optional solution. Figure 2 All vertex attributes are regarded as strings, and a high-order entropy compression full-text index on the vertex attribute set is constructed based on the compression index framework GeCSA proposed by Huo et al., called VIndex. Using VIndex and the extract algorithm, vertex attribute extraction can be achieved. Based on this vertex attribute extraction idea, the embodiment of the present invention combines the above-mentioned compression structure of the graph structure to implement vertex attribute extraction. Specifically:
[0095] Calculate the vertex number according to the compression structure of the graph structure, obtain the location of the vertex attribute in the vertex attribute compression index according to the vertex number and the corresponding bit array, and extract the vertex attribute value from the vertex attribute compression index. Please refer to Figure 3 and includes:
[0096] S301. Establish a high-order entropy compression full-text index of vertex attributes according to the vertex attribute set; the vertex attribute set includes the vertex attributes of all vertices.
[0097] S302. Calculate the starting position and ending position of the vertex number in the vertex attribute set according to the vertex number and the corresponding bit array.
[0098] S303. Extract the corresponding vertex attribute value from the vertex attribute high-order entropy compressed full-text index according to the start position and end position of the vertex attribute.
[0099] For example, Figure 2 taking the property graph G as an example, let T V = T V [0, n V - 1] represent the vertex attribute set of the property graph G defined on the alphabet of size σ V with length n V . Let T V [i, j] represent the substring of T V , that is, the concatenation of the symbols in T V at positions i, i + 1,..., j. Consider all vertex attributes in the property graph G as strings. Based on the compression index framework GeCSA proposed by Huo et al., construct the vertex attribute high-order entropy compressed full-text index on T V , called VIndex. Using VIndex and the extract algorithm, the vertex attribute extraction on the vertex v (∈ V) in the property G can be realized. The vertex attribute extraction query restores the attribute string T V [stpos, stpos + len - 1], where stpos and len are the start position and length of the attribute string of vertex v in T V .
[0100] Table 3 gives Figure 2 the category to which the vertices of the property graph G belong, the vertex attribute set, and the corresponding bit array. In Table 3: The first column represents the vertex number; the second column represents the category to which the given vertex belongs; the third column represents the vertex attribute set (Vertex attributes); the last column represents the bit array VB, which implicitly contains the start position of each vertex attribute string in T V . For 0 ≤ j ≤ n V , if j is the start position of the attribute string in T V , then define the bit array VB[j] = 1, otherwise the bit array VB[j] = 0. The last column in Table 3 shows the positions j that satisfy the bit array VB[j] = 1. For example, the bit array VB
[16] = 1 indicates that the position j = 16 is the start position of the vertex attribute "Alice 30". The start position of this vertex attribute is obtained by calling select1(VB, i), where 1 ≤ i ≤ n. The first sub-column in the third column (Vertex attributes column) of Table 3 is the in-class number vid of each vertex.
[0101] Table 3 The category to which the vertices of the property graph G belong, the vertex attribute set, and the corresponding bit array
[0102] Vid Vertex labels Vertex attributes VB[j]=1 0 director 0 Tom 23 0 1 director 1 Tim 34 8 2 actor 0 Alice 30 16 3 film 0 Ben 26 4 film 1 Seas 31 5 student 0 Lam MIT 37 6 studentactor 0 Mary Harvard 46 7 film 2 Bear 60 8 actor 1 Agela Matrix 66 9 80
[0103] Next, the example in Table 3 is used to illustrate how extractV() performs vertex attribute extraction. Given vertex Vid, the bit array VB is used to locate the start position and end position of the vertex attributes of the given vertex Vid in the vertex attribute high-order entropy compressed full-text index VIndex. Taking Vid = 2 as an example: calling the getPos(2) function, we can get stpos = select1(VB, 3) = 16, endpos = select1(VB, 4) - 1 = 25. Then the length of the attribute string to be retrieved is len = endpos - stpos + 1 = 10. Then call the extractV(16, 10) function, and return the vertex attribute string substr = "0Alice 30", which is the vertex attribute associated with vertex Vid = 2. Among them, the vertex number like Vid = 2 can be accessed through the compressed structure of the graph structure shown in Tables 1 and 2; the specific extractV() algorithm can be implemented using existing solutions and will not be elaborated here.
[0104] It can be seen that the embodiment of the present invention pre-constructs a vertex attribute high-order entropy compressed full-text index according to the vertex attributes of all vertices in the attribute graph as shown in Figure 2 Define a bit array VB, which implicitly contains the start position of each vertex attribute string. The vertex number Vid is quickly obtained through the compressed structure of the graph structure. By calling select1(VB, Vid + 1) and select1(VB, Vid + 2) - 1 respectively, the start position and end position of the vertex attribute string corresponding to the vertex number Vid are calculated, that is, the vertex attribute compression index. Finally, call the extractV() algorithm to obtain the vertex attribute string corresponding to the required vertex number Vid, denoted here as extractV(Vid), and obtain the vertex attribute corresponding to the vertex number Vid through extractV(Vid).
[0105] For extracting edge attribute values through the edge attribute compression index, the embodiment of the present invention provides an optional solution. Figure 2 All edge attributes in are regarded as strings, and based on the compression index framework GeCSA proposed by Huo et al., a high-order entropy compressed full-text index on the edge attribute set is constructed, called EIndex. Using EIndex and the extract algorithm, edge attribute extraction can be realized. Based on this edge attribute extraction idea, the embodiment of the present invention combines the above-mentioned compressed structure of the graph structure to realize edge attribute extraction. Specifically:
[0106] Calculate the edge number according to the compressed structure of the graph structure, obtain the location of the edge attribute in the edge attribute compression index according to the edge number and the corresponding bit array, and extract the attribute value from the edge attribute compression index. Please refer to Figure 4 including:
[0107] S401. Establish a high-order entropy compressed full-text index for edge attributes based on the edge attribute set; the edge attribute set includes the edge attributes between all adjacent vertices;
[0108] S402. Obtain the corresponding edge number according to the vertex number;
[0109] S403. Calculate the starting position and ending position of the edge number in the edge attribute set according to the edge number and the corresponding bit array;
[0110] S404. Extract the corresponding edge attributes from the high-order entropy compressed full-text index of edge attributes according to the starting position and ending position in the edge attribute set.
[0111] For example, taking the Figure 2 attribute graph G as an example, let T E = T E [0, n E - 1] represent the edge attribute set of G defined on the alphabet σ E with length n E . Let T E [i, j] represent the substring of T E , that is, the concatenation of the symbols in T E at positions i, i + 1,..., j. Regard all the edge attributes in the attribute graph G as strings. Based on the compression index framework GeCSA proposed by Huo et al., construct a high-order entropy compressed full-text index for the edge attributes on T E , called EIndex. Using EIndex and the extract algorithm, the extraction of edge attributes on the edge e (∈E) in the graph G can be realized. The edge attribute extraction query restores the attribute string T E [stpos, stpos + len - 1], where stpos and len are the starting position and length of the attribute string associated with the edge e in T E .
[0112] Table 4 shows the Figure 2 edge attribute set and the corresponding bit array of the attribute graph G. In Table 4: the first column represents the associated edge number; the second column represents the edge attribute set; the last column represents the bit array EB, which implicitly contains the starting position of the attribute string on each edge in T E . For 0 ≤ j ≤ n E , if j is the starting position of the attribute string in T E , then define the bit array EB[j] = 1, otherwise the bit array EB[j] = 0. The last column in Table 3 shows the positions j that satisfy the bit array EB[j] = 1. For example, the bit array EB
[15] = 1 indicates that the position j = 15 is the starting position of the edge attribute "guide 1993 - 1995". The starting position of this edge attribute is obtained by calling select1(EB, i), where 1 ≤ i ≤ m.
[0113] Table 4 Edge attributes of the attribute graph G and the corresponding bit arrays
[0114] Eid Edge attributes EB[j]=1 0 guide 1995 - 2000 0 1 guide 1993 - 1995 15 2 likes 90 30 3 guide 2000 - 2010 38 4 likes 90 53 5 likes 87 61 6 likes 85 69 7 likes 85 77 8 likes 87 85 9 likes 100 94 10 serve 1998 - 2001 104 11 likes 78 119 12 likes 100 127 13 136
[0115] Next, use the example in Table 4 to illustrate how extractE() extracts edge attributes. Given the edge Eid, use the bit array EB to locate the start position and end position of the attributes of edge Eid in the edge attribute high-order entropy compressed full-text index EIndex. Taking Eid = 4 as an example, calling the getPos(4) function, we get stpos = select1(EB, 5) = 53, endpos = select1(EB, 6) - 1 = 60. Then the length of the attribute string to be extracted is len = endpos - stpos + 1 = 8. Then call the extractE(53, 8) function, and return the edge attribute string substr = "likes 90", which is the edge attribute associated with edge Eid = 4. Among them, the associated edge number like Eid = 4 is obtained from the vertex number obtained by accessing the compressed structure of the graph structure as shown in Tables 1 and 2. For example, by accessing the vertex with vertex number VID = 0 through the compressed structure of the graph structure, this vertex corresponds to 4 out-neighbor edges and 2 in-neighbor edges. For the out-neighbor edges, their associated edge numbers are corresponding according to the vertex number order, and for the in-neighbor edges, they can be pre-converted into the out-neighbor edges of adjacent vertices, and their associated edge numbers can also be corresponding according to the vertex number order; similarly, by accessing the vertex number Vid = 1 through the compressed structure of the graph structure, determine the associated edge number corresponding to vertex number Vid = 1; and so on, all vertex numbers can be determined to correspond to the associated edge numbers; each out-neighbor edge corresponds to its associated edge number, and further edge attribute extraction is carried out according to these associated edge numbers.
[0116] It can be seen that the embodiment of the present invention pre-constructs a high-order entropy compressed full-text index for edge attributes according to all the edge attributes between adjacent vertices in the attribute graph as Figure 2 shown, defines the bit array EB, which implicitly contains the start position of the edge attributes corresponding to each edge between adjacent vertices. Quickly obtain the vertex number Vid through the compressed structure of the graph structure, find the associated edge number through the vertex number Vid, define Eid as the associated edge number, calculate the start position and end position corresponding to the associated edge number Eid in the edge attribute string, that is, the edge attribute compression index, by calling select1(EB, Eid + 1) and select1(EB, Eid + 2) - 1 respectively. Finally, call the extractE() algorithm to obtain the required edge attribute string, denoted as extractE(Eid) here, and obtain the edge attribute corresponding to the associated edge number Eid through extractE(Eid).
[0117] S40. Query the property graph according to the compressed structure of the graph structure and the attribute compression index.
[0118] Based on the above storage of the compressed structure of the graph structure, an optional solution for querying the property graph according to the compressed structure of the graph structure is proposed in the embodiments of the present invention. Please refer to Figure 5 , including:
[0119] Initialize the given vertex number vid and the preset distance k;
[0120] Convert the given vertex number vid to the given global number Vid;
[0121] Enqueue the given global number Vid into the FIFO queue Q, and initialize the output result set A;
[0122] Take out the intermediate vertex number u from the FIFO queue Q, and judge whether the distance between the intermediate vertex number u and the given global number Vid is less than the preset distance k. If it is less, find the unexplored adjacent vertex number v corresponding to the given vertex number vid, enqueue the adjacent vertex number v into the FIFO queue Q, and add the adjacent vertex number v to the output result set A. Repeat the above process of taking out from the FIFO queue Q until the FIFO queue Q is empty, and return the output result set A;
[0123] Among them, the adjacent vertex number v is obtained by accessing the compressed structure of the graph structure; use u.id to represent the global number Vid of the intermediate vertex number u, then Ngap[u.id + 1] - Ngap[u.id] represents the number of adjacent vertices num of the intermediate vertex number u; and then calculate Adj u.id [j] to obtain the j-th adjacent vertex number v (0 ≤ j ≤ num - 1) of the intermediate vertex number u.
[0124] It can be seen that in the embodiments of the present invention, given the number vid of a person (vertex) and the integer k, all persons whose distance from the given person is at most k can be queried and returned. The query process maintains a FIFO queue Q, which contains vertices that have not explored adjacent point numbers and whose distance from the given global number Vid is less than k. By exploring vertices that meet the preset distance k in the compressed space in the order of the distance between the adjacent vertex number and the given vertex number vid, since the query data storage amount is small and the calculation amount of the query process is small, the query efficiency is improved.
[0125] Further, by Figure 2It can be seen that the edges connecting adjacent vertices in the property graph have different edge attributes. During the query process on the property graph, by using the compressed structure of the graph structure and the edge attributes, that is, by using more information for filtering, the edge attributes of the edges can be fully utilized to further improve the query efficiency. Therefore, the embodiments of the present invention provide an alternative solution for querying the property graph according to the compressed structure of the graph structure and the attribute compression index. Please refer to Figure 6 , including:
[0126] Initialize the given vertex number vid, the preset distance k, and the preset filtering condition;
[0127] Convert the given vertex number vid to the given global number Vid;
[0128] Enqueue the given global number Vid into the FIFO queue Q, and initialize the output result set A;
[0129] Take out the intermediate vertex number u from the FIFO queue Q, and determine whether the distance between the intermediate vertex number u and the given global number Vid is less than the preset distance k. If it is less, then find the unexplored adjacent vertex number v corresponding to the given vertex number vid, and check whether the intermediate vertex number u and the adjacent vertex number v satisfy the preset filtering condition. If they satisfy, then enqueue the adjacent vertex number v into the FIFO queue Q, and add the adjacent vertex number v to the output result set A. Repeat the above process of taking out from the FIFO queue Q until the FIFO queue Q is empty, and return the output result set A;
[0130] Among them, the adjacent vertex number v is obtained by accessing the compressed structure of the graph structure; use u.id to represent the global number Vid of the intermediate vertex number u, then Ngap[u.id + 1] - Ngap[u.id] represents the number of adjacent vertices num of the intermediate vertex number u; and then calculate Adj u.id [j] to obtain the jth adjacent vertex number v (0 ≤ j ≤ num - 1) of the intermediate vertex number u;
[0131] Whether the intermediate vertex number u and the adjacent vertex number v satisfy the preset filtering condition is checked by the edge attribute value extracted by the edge attribute compression index; by using Ngap[k] (u.id ≤ k ≤ u.id + num - 1), the Eid of the kth outer adjacent (out) edge of the intermediate vertex number u can be obtained, and then call extractE(Eid) to obtain the edge attribute on the outer adjacent edge, and this edge attribute represents different filtering conditions, such as Figure 2 "likes", "guide", "serve", etc. in
[0132] It can be seen that in the embodiments of the present invention, given a person's ID vid and an integer k, all persons with preset filtering conditions and a distance of at most k from the given person can be queried and returned. In the embodiments of the present invention, by calling extractE to obtain edge attributes to check the preset filtering conditions between two vertices where there are edges, in addition to maintaining the advantage of small computational complexity in the query process, the query accuracy is further improved.
[0133] Furthermore, the present invention provides another alternative solution for querying an attributed graph based on the compressed structure of the graph structure and the attribute compression index. Please refer to Figure 7 , including:
[0134] Initialize the first given vertex ID vid1, the second given vertex ID vid2, and the preset filtering conditions;
[0135] Judge whether the first given vertex ID vid1 and the second given vertex ID vid2 are the same vertex. If so, return the distance l as 0. Otherwise, convert the first given vertex ID vid1 and the second given vertex ID vid2 into the first given global ID Vid1 and the second given global ID Vid2; enqueue the first given global ID Vid1 into the FIFO queue Q, initialize the distance l as 0, dequeue the intermediate vertex ID u from the FIFO queue Q, judge whether the intermediate vertex ID u and the second given global ID Vid2 are the same vertex. If so, return the distance l. Otherwise, find the unexplored adjacent vertex ID v corresponding to the first given vertex ID vid1, check whether the intermediate vertex ID u and the adjacent vertex ID v satisfy the preset filtering conditions. If satisfied, enqueue the adjacent vertex ID v into the FIFO queue Q and increment the distance l by 1. Repeat the above process of dequeuing from the FIFO queue Q until the FIFO queue Q is empty, and return the distance l;
[0136] Among them, the adjacent vertex ID v is obtained by accessing the compressed structure of the graph structure; use u.id to represent the global ID Vid of the intermediate vertex ID u, then Ngap[u.id + 1] - Ngap[u.id] represents the number of adjacent vertices num of the intermediate vertex ID u; and then calculate Adj u.id [j] to obtain the jth adjacent vertex ID v (0 ≤ j ≤ num - 1) of the intermediate vertex ID u.
[0137] Whether the intermediate vertex ID u and the adjacent vertex ID v satisfy the preset filtering conditions is checked by the edge attribute values extracted by the edge attribute compression index; by using Ngap[k] (u.id ≤ k ≤ u.id + num - 1), the Eid of the kth outer adjacent (out) edge of the intermediate vertex ID u can be obtained, and then call extractE(Eid) to obtain the edge attributes on the outer adjacent edge.
[0138] It can be seen that the embodiments of the present invention can solve the query of relatively complex property graphs, such as the query of the property graph Q13 problem. Given PersonX and PersonY, find the shortest path between them in the subgraph induced by the Knows relationship. The implementation process includes:
[0139] (a1) Determine whether the two given vertex numbers vid1 and the given vertex number vid2 are the same vertex. If so, return the distance l as 0; otherwise, continue.
[0140] (a2) Convert the given vertex number vid1 and the given vertex number vid2 into the global numbers Vid1 and Vid2.
[0141] (a3) Enqueue the global number Vid1 into the FIFO queue Q, initialize the distance l as 0, and execute the following steps until the FIFO queue Q is empty:
[0142] (b1) Dequeue the intermediate vertex number u from the FIFO queue Q.
[0143] (b2) If the intermediate vertex number u and the global number Vid2 are the same vertex, return the distance l.
[0144] (b3) Find the unexplored adjacent vertex number v of the intermediate vertex number u, check whether the edge attribute of the edge connecting the intermediate vertex number u and the adjacent vertex number v has a friendship relationship (here the preset filtering condition is designed as a friendly relationship). If it has a friendship relationship, enqueue the adjacent vertex number v into the FIFO queue Q, and at the same time increment the distance l by 1.
[0145] (a4) If there is no path between the given vertex number vid1 and the given vertex number vid2, return -1.
[0146] Further, from Figure 2 It can be seen that each vertex in the property graph also includes label information. During the query process on the property graph, the compressed structure of the graph structure, vertex attributes, and edge attributes are utilized, that is, more information is used for filtering, improving the query efficiency. Therefore, the embodiments of the present invention provide an alternative solution for querying the property graph according to the compressed structure of the graph structure and the attribute compression index. Please refer to Figure 8 , including:
[0147] Initialize the given vertex number vid, the given vertex attributes, the preset distance k, and the preset filtering condition.
[0148] Convert the given vertex number vid into the given global number Vid.
[0149] Enqueue the given global number Vid into the first FIFO queue Q, and initialize the output result set A.
[0150] Take the first intermediate vertex number u from the first FIFO queue Q, and determine whether the distance between the first intermediate vertex number u and the given global number Vid is less than the preset distance k. If it is less, then find the unexplored first adjacent vertex number v corresponding to the given vertex number vid and enqueue it into the first FIFO queue Q. Check whether the first intermediate vertex number u and the first adjacent vertex number v satisfy the preset filtering condition. If they satisfy, find the vertex attribute of the first adjacent vertex number v, and compare whether the vertex attribute of the first adjacent vertex number v and the given vertex attribute are of the same attribute. If they belong to the same attribute, enqueue the first adjacent vertex number v into the second FIFO queue P. Take the second intermediate vertex number p from the second FIFO queue P, find the unexplored second adjacent vertex number v1 of the first adjacent vertex number v, obtain the relevant information of the second intermediate vertex number p according to the label of the second adjacent vertex number v1, and add the relevant information of the second intermediate vertex number p to the output result set A. Repeat the above process of taking from the second FIFO queue P until the second FIFO queue P is empty, and continue to repeat the above process of taking from the first FIFO queue Q until the first FIFO queue Q is empty, and return the output result set A;
[0151] Among them, the first adjacent vertex number v and the second adjacent vertex number v1 are obtained by accessing the second adjacency list; use u.id to represent the global number Vid of the first intermediate vertex number u, then Ngap[u.id + 1] - Ngap[u.id] represents the number of adjacent vertices num of the first intermediate vertex number u; and then calculate Adj u.id [j] to obtain the jth first adjacent vertex number v (0 ≤ j ≤ num - 1) of the first intermediate vertex number u; similarly, the adjacent vertex number v1 can be obtained by calculating Adj u1.id [j];
[0152] Whether the first intermediate vertex number u and the first adjacent vertex number v satisfy the preset filtering condition is checked by the edge attribute value extracted by the edge attribute compression index; use Ngap[k] (u.id ≤ k ≤ u.id + num - 1) to obtain the Eid of the kth outer adjacent (out) edge of the first intermediate vertex number u, and then call extractE(Eid) to obtain the edge attribute value on the outer adjacent edge;
[0153] The relevant information of the second intermediate vertex number p is obtained through the vertex attribute value extracted by the vertex attribute compression index; according to the label of the second intermediate vertex number p, call extractV(v1.id) to obtain the vertex attribute of the second adjacent vertex number v1, and obtain the relevant information of the second intermediate vertex number p according to the corresponding vertex attribute, or call extractE(v1.id) to obtain the edge attribute of the second adjacent vertex number v1, and then call extractV(v1.id) to obtain the vertex attribute of the second adjacent vertex number v1, and obtain the relevant information of the second intermediate vertex number p according to the corresponding vertex attribute; the label of the second intermediate vertex number p can be obtained during the process of converting the given vertex number vid to the given global number Vid.
[0154] It can be seen that the embodiments of the present invention can solve the queries of relatively complex property graphs, such as the property graph Q1 problem query, given a name (first name) and its number vid, return at most k = 20 friends with the same name, and output them in ascending order of the distance from the given person (at most 3); for people within the same distance, sort them by the last name, and the returned results should include a list of workplaces and study locations. Specifically:
[0155] First, perform preprocessing to convert the given vertex number vid to the global number Vid.
[0156] Enqueue the global number Vid into the FIFO queue Q, and execute the following steps until the FIFO queue Q is empty:
[0157] (c1), Take out the intermediate vertex number u from the FIFO queue Q;
[0158] (c2), If the distance from the intermediate vertex number u to the global number Vid is less than k, then find the unexplored adjacent vertex number v of the intermediate vertex number u and enqueue it into the FIFO queue Q;
[0159] (c3), Check the outer adjacent edge attributes of the intermediate vertex number u and the adjacent vertex number v to check whether there is a friendship relationship (here the preset filtering condition is designed as a friendly relationship);
[0160] (c4), If there is a friendly relationship, then obtain the name attribute on the adjacent vertex number v by calling extractV(v.id), compare it with the name of the given vertex number vid, if they are the same name, then enqueue the adjacent vertex number v into the FIFO queue P, and execute the following steps until the FIFO queue P is empty:
[0161] (d1), Take out the intermediate vertex number p from the FIFO queue P;
[0162] (d2), find an unexplored adjacent vertex number v1 of the intermediate vertex number p.
[0163] (d3) For each adjacent vertex number v1 of the intermediate vertex number p, determine its vertex attribute. Here, the vertex attribute is a label. If the label is "place", call extractV(v.id) to obtain the vertex attribute index on this vertex, and correspondingly determine that the vertex attribute is the location. If the label is "organization", call extractE to obtain the edge attribute index on the outgoing edge from the intermediate vertex number p to the adjacent vertex number v1, and correspondingly determine the edge attribute. If the edge attribute value on this edge includes "workAt" or "studyAt", then continue to call extractV(v.id) to obtain information related to the intermediate vertex number p, such as the work place or study place and its date.
[0164] (d4) Obtain the vertex attribute on this vertex by calling extractV(v.id). Here, the vertex attribute is the name attribute. Compare it with the name of the given vertex number vid. If they are the same, enqueue the adjacent vertex number v1 into the FIFO queue P.
[0165] It should be noted that in the query process as shown in Figure 5 、 Figure 6 、 Figure 7 、 Figure 8 , there may be the same numbers among vertices of different classes, and preprocessing is required to map the given vertex number vid to the global number Vid. The embodiment of the present invention provides an optional solution for mapping the given vertex number vid to the global number Vid: use a mapping function to obtain the starting position of the given vertex number vid in the global system, and then use the minimum perfect hash function MPHash to map the given vertex number vid to the intra-class number of the class it belongs to. The sum of this starting position and the intra-class number gives the global number Vid associated with the given vertex number vid. At the same time, during the mapping process, the global number Vid carries the label information of the given vertex number vid. For example, there are two types of vertices, person and city; there are 10 vertices in the person category, and the number vid is from 1 to 10; there are also 10 vertices in the city category, and the number vid is also from 1 to 10; the numbered vertices vid in the two categories have duplicate phenomena, which will affect the query. Therefore, it is necessary to map the number vid to the global number Vid using hashing so that there will definitely be no same number Vid among different classes.
[0166] It should be noted that in the embodiments of the present invention, an external neighbor compression structure and an incoming neighbor compression structure can be constructed. The query processes are different, and the situations of calling the external neighbor compression structure and the incoming neighbor compression structure are different. During the query process, the external neighbor compression structure and / or the incoming neighbor compression structure can be considered as query conditions according to actual needs. Which query needs to call the external neighbor compression structure and which query needs to call the incoming neighbor compression structure will not be elaborated here.
[0167] In summary, the compression index and query method for the property graph proposed in the embodiments of the present invention no longer adopt the traditional adjacency list storage and query method. Instead, in order to reduce the storage amount of data during the storage process, a new adjacency list storage and query method is proposed. Specifically: a compression structure of the graph structure is constructed according to the adjacency list of each vertex, and the corresponding adjacent vertex numbers of each vertex in the adjacency list are accessed by the constructed compression structure of the graph structure, realizing the compressed storage of the property graph and improving the query efficiency of the property graph database; at the same time, during the query process of the property graph database in the embodiments of the present invention, corresponding vertex attribute compression indexes and edge attribute compression indexes are also established for vertex attributes and edge attributes, supporting query operations on various types of property graphs on the vertex and edge attribute compression indexes, further improving the query efficiency of the property graph database.
[0168] Based on the same inventive concept of the above method, please refer to Figure 9 , the embodiments of the present invention provide an electronic device, including a processor 901, a communication interface 902, a memory 903, and a communication bus 904. Among them, the processor 901, the communication interface 902, and the memory 903 complete mutual communication through the communication bus 904;
[0169] The memory 903 is used to store a computer program;
[0170] The processor 901 is used to implement the steps of the above compression index and query method for the property graph when executing the program stored in the memory 903.
[0171] The embodiments of the present invention provide a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the above compression index and query method for the property graph are implemented.
[0172] For the embodiments of the apparatus / electronic device / storage medium, since they are basically similar to the method embodiments, the description is relatively simple. For related parts, please refer to the partial description of the method embodiments.
[0173] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0174] Although the present invention has been described in conjunction with various embodiments, however, in the process of implementing the claimed invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by referring to the specification and its accompanying drawings. In the specification, the term "comprising" does not exclude other components or steps, and "a" or "one" does not exclude a plurality. A single processor or other unit may implement several functions recited in the embodiments of the present invention. Certain measures are recited in different embodiments, but this does not mean that these measures cannot be combined to produce good results.
[0175] The above content is a further detailed description of the present invention in conjunction with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A compression index and query method for an attribute graph, characterized in that, Applied to an attribute graph database system, the attribute graph maintained by the attribute graph database system includes a plurality of vertices, edges, and their corresponding attributes. The method includes: Constructing an adjacency list for each vertex according to the attribute graph; Constructing a compressed structure of the graph structure according to the adjacency list of each vertex; Calculating vertex / edge numbers according to the compressed structure of the graph structure, obtaining the location of the attribute in the attribute compression index according to the vertex / edge number and the corresponding bit array, and extracting the attribute value from the attribute compression index; the attribute compression index includes a vertex attribute compression index and an edge attribute compression index; the bit array correspondingly records the starting positions of all edge attributes / vertex attributes; Querying the attribute graph according to the compressed structure of the graph structure and the attribute compression index; Wherein, the compressed structure of the graph structure includes a coding sequence, the starting position of the adjacency list of each vertex in the coding sequence, the total difference of the adjacency list of each vertex, and the first vertex number of the adjacency list of each vertex; the coding sequence is obtained by calculating the difference between adjacent vertex numbers in the adjacency list and encoding the difference.
2. The compression index and query method for an attribute graph according to claim 1, characterized in that, The compressed structure of the graph structure includes an out-neighbor compression structure and an in-neighbor compression structure.
3. The compression index and query method for an attribute graph according to claim 1, characterized in that, Accessing the corresponding adjacent vertex number of each vertex in the adjacency list through the compressed structure of the graph structure is expressed as: where i represents the i-th vertex, and Adj i [j] represents the j-th adjacent vertex number in the adjacency list accessed by the i-th vertex through the compressed structure of the graph structure, Sam[i] represents the number of the first vertex in the adjacency list of the i-th vertex, and Adj i [j - 1] represents the (j - 1)-th adjacent vertex number in the adjacency list accessed by the i-th vertex through the compressed structure of the graph structure, decompress() represents the decoding function, with input parameters S, pos, and nbits, S represents the encoded sequence, and nbits represents the number of decoding bits, X[i + 1] stores the starting position of the adjacency list of the (i + 1)-th vertex in the encoded sequence, X[i] stores the starting position of the adjacency list of the i-th vertex in the encoded sequence, Ngap[i + 1] stores the total number of differences to the adjacency list of the (i + 1)-th vertex, Ngap[i] stores the total number of differences to the adjacency list of the i-th vertex, and pos = X[i] + (j - 1) * nbits.
4. The compression index and query method for an attribute graph according to claim 1, characterized in that, Calculating the vertex number according to the compressed structure of the graph structure, obtaining the location of the vertex attribute in the vertex attribute compression index according to the vertex number and the corresponding bit array, and extracting the vertex attribute value from the vertex attribute compression index, including: Establishing a vertex attribute high-order entropy compressed full-text index according to the vertex attribute set; the vertex attribute set includes the vertex attributes of all vertices; Calculating the starting position and ending position of the vertex number in the vertex attribute set according to the vertex number and the corresponding bit array; Extracting the corresponding vertex attribute value from the vertex attribute high-order entropy compressed full-text index according to the starting position and the ending position of the vertex attribute.
5. The compression index and query method for an attribute graph according to claim 1, characterized in that, Calculating the edge number according to the compressed structure of the graph structure, obtaining the location of the edge attribute in the edge attribute compression index according to the edge number and the corresponding bit array, and extracting the attribute value from the edge attribute compression index, including: Establishing an edge attribute high-order entropy compressed full-text index according to the edge attribute set; the edge attribute set includes the edge attributes between all adjacent vertices; Obtaining the corresponding edge number according to the vertex number; Calculating the starting position and ending position of the edge number in the edge attribute set according to the edge number and the corresponding bit array; Extracting the corresponding edge attribute value from the edge attribute high-order entropy compressed full-text index according to the starting position and the ending position in the edge attribute set.
6. The compression index and query method for an attribute graph according to claim 1, characterized in that, Querying the attribute graph according to the compressed structure of the graph structure and the attribute compression index, including: Initializing a given vertex number and a preset distance; Converting the given vertex number into a given global number; Enqueuing the given global number into a FIFO queue and initializing an output result set; Take out the intermediate vertex number from the FIFO queue, and determine whether the distance between the intermediate vertex number and the given global number is less than the preset distance. If it is less, find the unexplored adjacent vertex numbers corresponding to the given vertex number, enqueue the adjacent vertex numbers into the FIFO queue, and add the adjacent vertex numbers to the output result set. Repeat the above process of taking out from the FIFO queue until the FIFO queue is empty, and return the output result set; Among them, the adjacent vertex numbers are obtained by accessing the compressed structure of the graph structure.
7. The compression index and query method for an attribute graph according to claim 1, characterized in that, Query the property graph according to the compressed structure of the graph structure and the attribute compression index, including: Initialize the given vertex number, preset distance, and preset filtering condition; Convert the given vertex number to a given global number; Enqueue the given global number into the FIFO queue and initialize the output result set; Take out the intermediate vertex number from the FIFO queue, and determine whether the distance between the intermediate vertex number and the given global number is less than the preset distance. If it is less, find the unexplored adjacent vertex numbers corresponding to the given vertex number, and check whether the intermediate vertex number and the adjacent vertex numbers meet the preset filtering condition. If they meet, enqueue the adjacent vertex numbers into the FIFO queue, and add the adjacent vertex numbers to the output result set. Repeat the above process of taking out from the FIFO queue until the FIFO queue is empty, and return the output result set; Among them, the adjacent vertex numbers are obtained by accessing the compressed structure of the graph structure; Whether the intermediate vertex number and the adjacent vertex numbers meet the preset filtering condition is checked by the edge attribute values extracted by the edge attribute compression index.
8. The compression index and query method for the attribute graph according to claim 1, characterized in that, Query the property graph according to the compressed structure of the graph structure and the attribute compression index, including: Initialize the first given vertex number, the second given vertex number, and the preset filtering condition; Determine whether the first given vertex number and the second given vertex number are the same vertex. If so, return a distance of 0. Otherwise, convert the first given vertex number and the second given vertex number to the first given global number and the second given global number; enqueue the first given global number into the FIFO queue, initialize the distance to 0, take out the intermediate vertex number from the FIFO queue, and determine whether the intermediate vertex number and the second given global number are the same vertex. If so, return the distance. Otherwise, find the unexplored adjacent vertex numbers corresponding to the first given vertex number, and check whether the intermediate vertex number and the adjacent vertex numbers meet the preset filtering condition. If they meet, enqueue the adjacent vertex numbers into the FIFO queue and increment the distance by 1. Repeat the above process of taking out from the FIFO queue until the FIFO queue is empty, and return the distance; Among them, the adjacent vertex numbers are obtained by accessing the compressed structure of the graph structure; Whether the intermediate vertex number and the adjacent vertex number satisfy the preset filtering condition is checked by the edge attribute value extracted by the edge attribute compression index.
9. The compression index and query method for the attribute graph according to claim 1, characterized in that, Querying the property graph according to the compressed structure of the graph structure and the attribute compression index includes: Initializing a given vertex number, given vertex attributes, a preset distance, and a preset filtering condition; Converting the given vertex number to a given global number; Enqueuing the given global number into a first FIFO queue and initializing an output result set; Taking out a first intermediate vertex number from the first FIFO queue, determining whether the distance between the first intermediate vertex number and the given global number is less than the preset distance. If it is less, finding the unexplored first adjacent vertex number corresponding to the given vertex number, enqueuing it into the first FIFO queue, checking whether the first intermediate vertex number and the first adjacent vertex number satisfy the preset filtering condition. If they satisfy, finding the vertex attributes of the first adjacent vertex number, comparing whether the vertex attributes of the first adjacent vertex number and the given vertex attributes are of the same attribute. If they are of the same attribute, enqueuing the first adjacent vertex number into a second FIFO queue, taking out a second intermediate vertex number from the second FIFO queue, finding the unexplored second adjacent vertex number of the first adjacent vertex number, obtaining the relevant information of the second intermediate vertex number according to the label of the second adjacent vertex number, and adding the relevant information of the second intermediate vertex number into the output result set. Repeating the above process of taking out from the second FIFO queue until the second FIFO queue is empty, and continuing to repeat the above process of taking out from the first FIFO queue until the first FIFO queue is empty, and returning the output result set; Wherein, the first adjacent vertex number and the second adjacent vertex number are both obtained by accessing the compressed structure of the graph structure; Whether the first intermediate vertex number and the first adjacent vertex number satisfy the preset filtering condition is checked by the edge attribute value extracted by the edge attribute compression index; The relevant information of the second intermediate vertex number is obtained by the vertex attribute value extracted by the vertex attribute compression index.
Citation Information
Patent Citations
Attribute graph processing method and device, electronic equipment and readable storage medium
CN113822315A
Code attribute graph compression method and device for source code vulnerability detection
CN113987522A