A method and system for optimizing the implementation of multi-layer adjacency queries based on a graph storage model
By introducing a cache layer into the graph storage model, recording the number of data repetitions, the efficiency of multi-layer adjacency query is optimized, the problems of low efficiency and large memory usage in the existing technology are solved, and efficient result set query is achieved.
Patent Information
- Application Number
- CN202211671712.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-12-26
AI Technical Summary
The traversal method of existing graph storage models is inefficient, occupies a large computer memory and is prone to memory overflow errors.
Introduce a cache layer during multi-layer adjacency query, record the number of data repetitions through the cache mechanism, and optimize the result set query.
Improve query efficiency, reduce memory usage, and avoid memory overflow errors.
Smart Images

Figure CN115858873B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of optimizing adjacency queries for graph storage models, and particularly relates to a method and system for optimizing multi-layer adjacency queries based on a graph storage model. Background Art
[0002] Graph storage models store data based on vertices and edges, and are a set of points and edges. Among them, "points" represent entities, and "edges" represent the relationships between entities. Graph storage models can be divided into directed graphs and undirected graphs. A directed graph is marked with an arrow to indicate the direction, as shown in Figure 1 ; conversely, it is an undirected graph, as shown in Figure 2 . In a directed graph, a certain point as the starting point of an edge is called the out-degree, and as the ending point of an edge is called the in-degree. For example, in Figure 1 , the out-degree of vertex v1 is 2 and the in-degree is 1.
[0003] Currently, there are mainly two ways to traverse a graph: depth-first search and breadth-first search. Depth-first search starts from a certain vertex v in the graph, visits this vertex, and then recursively traverses the graph from the unvisited adjacent vertices of vertex v until all vertices in the graph that are connected to v by a path are visited; if there are still vertices in the graph that have not been visited at this time, then select another unvisited vertex in the graph as the starting point and repeat the above process until all vertices in the graph are visited. Breadth-first search starts from a certain vertex v in the graph. After visiting v, it successively visits the unvisited adjacent vertices of v, and then starts from these adjacent vertices and successively visits their adjacent vertices, and makes the "adjacent vertices of the vertex that is visited first" precede the "adjacent vertices of the vertex that is visited later" until all adjacent vertices of the vertices that have been visited in the graph are visited; if there are still vertices in the graph that have not been visited at this time, then select another unvisited vertex in the graph as the starting point and repeat the above process until all vertices in the graph are visited. In other words, the process of breadth-first search traversing a graph starts from v, from near to far, and successively visits the vertices that are connected to v by a path and the path lengths are 1, 2, 3...
[0004] Depth-first search starts from a vertex and keeps recursively visiting. Although there is not much additional space overhead, the efficiency is low; while breadth-first search needs to cache the results retrieved for each layer. When the data volume is very large, it occupies a large amount of computer memory and is prone to reporting memory overflow errors. Summary of the Invention
[0005] Objective of the Invention: To solve the problems existing in the existing graph traversal methods, such as low efficiency, large computer memory occupation, and easy occurrence of memory overflow errors, etc., the present invention proposes an optimized implementation method and system for multi-layer adjacency query based on a graph storage model. A cache layer is introduced during multi-layer adjacency query, and the result set is optimized without duplicate removal. The optimization effect is more obvious when the number of adjacent layers queried is larger.
[0006] Technical Solution: An optimized implementation method for multi-layer adjacency query based on a graph storage model, comprising the following steps:
[0007] Step 1: Determine the total number of adjacent query layers n according to the adjacent query request;
[0008] Step 2: Starting from the starting vertex, through querying the graph storage model, store the queried adjacent vertex information into the first-layer result cache L1. In the result cache, the adjacent vertex information is cached with (vertex, count) as an element; wherein, the count represents the number of times the vertex is repeated; until the first-layer result cache L1 is filled or all the adjacent vertices of the starting vertex are stored into the first-layer result cache L1; then go to Step 3;
[0009] Step 3: Determine whether the current result cache is the (n - 1)-th layer result cache L n-1 , if not, then go to Step 4; if so, then go to Step 5;
[0010] Step 4: Denote the current result cache as L i , i = 1, 2, 3... (n - 2), take out the first element from the current result cache L i , find the adjacent vertices of the first element. If the found adjacent vertices have not been stored into the result cache L i+1 , then the count of the found adjacent vertices follows the count of the parent node and is stored into the result cache L i+1 , if the found adjacent vertices are already in the result cache L i+1 , then the count of the found adjacent vertices is accumulated on the basis of following the count of the parent node and is stored into the result cache L i+1 ; take out the second element from the current result cache L i , find the adjacent vertices of the second element. If the found adjacent vertices have not been stored into the result cache L i+1 , then the count of the found adjacent vertices follows the count of the parent node and is stored into the result cache L i+1 , if the found adjacent vertices are already in the result cache L i+1 , then the count of the found adjacent vertices is accumulated on the basis of following the count of the parent node and is stored into the result cache L i+1 ; and so on, until the result cache L i+1 is filled or the current result cache Li All adjacent points of all elements in have been stored in the result cache L i+1 In the caching process, go to step 3; during the caching process, if one or more adjacent points of one or more elements cannot be cached to the next level result cache because the result cache is full, the current position is recorded through the cursor;
[0011] Step 5: Cache L from the n-1th layer result n-1 Take out the first element and find the adjacent points of the first element. If the adjacent points found have not been stored in the result cache L n In the case of , the count of the adjacent points found will follow the count of the parent node and be stored in the result cache L n If the adjacent point found is already in the result cache L n In the case of , the count of the adjacent nodes found is accumulated based on the count of the parent node and stored in the result cache L n Middle; from the n-1th layer result cache L n-1 Take out the second element and find the adjacent point of the second element. If the adjacent point found has not been stored in the result cache L n In the case of , the count of the adjacent points found will follow the count of the parent node and be stored in the result cache L n If the adjacent point found is already in the result cache L n In the case of , the count of the adjacent nodes found is accumulated based on the count of the parent node and stored in the result cache L n and so on, until the result cache L is filled for the first time n Or the n-1th level result cache L n-1 All adjacent points of all elements in have been stored in the result cache L n In the cache, the current result is L n The result in is returned to the user; during the caching process, if one or more adjacent points of one or more elements cannot be cached to the next level result cache because the result cache is full, the current position is recorded through the cursor; go to step 6;
[0012] Step 6: Clear the result cache L n , according to the result cache L of the n-1th layer n-1 At the corresponding cursor position, find the corresponding adjacent points of one or more elements that have not been traversed in order, update the count of the adjacent points, and store them in the result cache L n Go to step 7;
[0013] Step 7: Determine the result cache L at this time n Is it full? If it is full, cache the current result L n The result in is returned to the user, and then the result cache is cleared. n ; If not filled, go to step 8;
[0014] Step 8: Determine whether the elements in the current result cache have been traversed. If so, clear the current result cache, start searching upward from the upper-layer result cache to find the cursor position. According to the cursor position, find the corresponding adjacent nodes of one or more untraversed elements in sequence, update the count of the adjacent nodes, and store them in the lower-layer result cache. And so on until the result cache L n is filled and returned to the user, or when the first-layer result cache L1 is empty, the result of the current result cache L n is returned to the user.
[0015] Further, in step 4, the count of the found adjacent nodes is accumulated on the basis of following the count of the parent node, including:
[0016] Assume that the count of the kth element is a, the adjacent nodes of the kth element include vertex s, the count of the qth element is b, the adjacent nodes of the qth element include vertex s, and assume that vertex s does not exist in the result cache to be cached; then the count of vertex s cached in the result cache is a + b.
[0017] Further, returning the result in the result cache L n to the user includes:
[0018] Returning the vertices in the result cache L n to the user.
[0019] The present invention also discloses a multi-layer adjacency query optimization implementation system based on a graph storage model. The system includes a network interface, a memory, and a processor; wherein,
[0020] The network interface is used for receiving and sending signals during the process of receiving and sending information with other external network elements;
[0021] The memory is used for storing computer program instructions that can run on the processor;
[0022] The processor is used for executing the steps of a multi-layer adjacency query optimization implementation method based on a graph storage model when running the computer program instructions.
[0023] The present invention also discloses a computer storage medium. The computer storage medium stores a program of a multi-layer adjacency query optimization implementation method based on a graph storage model. When the program of the multi-layer adjacency query optimization implementation method based on a graph storage model is executed by at least one processor, the steps of a multi-layer adjacency query optimization implementation method based on a graph storage model are implemented.
[0024] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0025] In the traditional sense, the adjacency query adopts the depth-first search or breadth-first search method. For complex graph storage models, a large amount of duplicate data is included in the result set. The present invention adopts a caching mechanism to record the number of data duplicates, greatly improving the query efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of a directed graph;
[0027] Figure 2 It is a schematic diagram of an undirected graph;
[0028] Figure 3 It is a schematic diagram of a multi-layer adjacency query;
[0029] Figure 4 It is a schematic diagram of a social network;
[0030] Figure 5 It is a schematic diagram of the traditional non-duplicate removal method;
[0031] Figure 6 It is a schematic diagram of adopting the method of the present invention; <(
[0032] Figure 7 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0033] The technical solution of the present invention will be further described below in conjunction with the drawings and embodiments.
[0034] Embodiment 1:
[0035] This embodiment discloses an optimized implementation method for multi-layer adjacency query based on a graph storage model, realizing multi-layer adjacency query based on a graph storage model and not removing duplicates from the result set. Now, taking Figure 3 as an example, the implementation method of this embodiment will be described. As Figure 7 shown, it mainly includes the following steps:
[0036] Step 1: When performing an n-layer adjacency query on a certain vertex, store the queried adjacent vertex data in the L1 result cache; the user can configure or adjust the cache size according to the hardware performance of the computer. The adjacent vertex data is cached in the form of key-value, where the key is the vertex and the value is the number of times the vertex is repeated, which is called the count hereinafter. When storing the queried adjacent vertex data in the L1 result cache, if the adjacent vertex already exists, increment the count by 1 until the first layer of the L1 result cache is full, or all the adjacent vertices of the starting point have been queried and stored in the L1 result cache.
[0037] Step 2: Determine whether the current layer is an intermediate layer. Here, the intermediate layer can be recognized as a layer that is not the penultimate layer. If so, go to Step 3; otherwise, go to Step 4.
[0038] Step 3: Collect the intermediate layer cache: Let the current layer be i, where i = 1, 2, 3... (n - 2), that is, before the penultimate layer. Retrieve the first vertex from the L i result cache, find its adjacent vertices, expand the adjacent vertices according to the count of the parent node (that is, multiply by the count of the parent node because repeated traversal is required), and put them into the L i+1 result cache (if it does not exist, directly put it into the L i+1 result cache; if it exists, add the count to the count in the L i+1 result cache); until the L i+1 result cache is filled, otherwise retrieve the next vertex from the L i result cache and repeat the above operations until the L i+1 result cache is full, or the L i result cache is empty. At this time, enter Step 2 again. If it is the intermediate layer, continue to collect; if it is the penultimate layer, enter Step 4.
[0039] Step 4: This layer is the penultimate layer. Retrieve a point from the L n-1 result cache, traverse its adjacent vertices (retain the count), and put them into the L n result set cache. If the L n is full, return it to the user. If the L n is not full, then retrieve the next point from the L n-1 result cache and repeat the above operations until the L n-1 result cache is empty, then return to the previous layer. Repeat similar steps until the result cache of layer L1 is empty.
[0040] Step 5: Continue to traverse the adjacent vertices of the starting point and repeat Steps 1 to 4 until all the adjacent vertices of the starting point have been traversed.
[0041] In Step 4, if there is no cache in the last layer, since the adjacent query results of the penultimate layer are already the final results, they can be directly returned to the user. If there is a cache in the last layer, then cache the adjacent query results of the penultimate layer L n-1 into the result cache of the last layer L n . When the L n result cache is filled or the points in the L n-1 result cache have been queried completely, return the results in the L n result cache to the user.
[0042] Return the results in the result cache to the user according to the following processing method:
[0043] If the points in the result set returned to the user are expanded on the server side and there is no count information, the receiving end does not need to perform special processing. If a certain result is repeated 100 times, the server will actually save this result 100 times and send it to the receiving end. At this time, both the memory consumption and the network communication consumption are relatively large.
[0044] If the points in the result set returned to the user retain the count information, the receiving end needs to perform special processing, that is, expand it on the receiving end. This processing can save a large amount of communication cost. If a certain point is repeated 100 times, when the server sends data to the receiving end, it only sends one piece of this data but will carry the repetition count. At this time, the special processing of the receiving end is to expand this piece of data according to the repetition count when using it.
[0045] Now, a method for optimizing the implementation of multi - layer adjacency query disclosed in this embodiment is adopted. In Figure 4 In the interpersonal network shown, "find the friends of the friends of the friends of Xiaoming, and the result set is not de - duplicated". In this figure, vertices represent "people" and edges represent the relationship of "knowing". For example, "find the friends of the friends of the friends of Xiaoming, and the result set is not de - duplicated", that is, a four - layer adjacency query and the result set is not de - duplicated. The specific operations include:
[0046] Assume that the cache size of each layer in this embodiment is 4. Refer to Figure 6 to understand this embodiment. Figure 6 The 4 circles in the rectangle in
[0047] indicate that 4 data can be stored in the cache with a size of 4. In the circle, the front represents the name of the person, and the back number represents the repetition count. When there are two data in the circle, the data with a strikethrough above represents the state when it is first (or in the middle) added to the cache, and the following is the final updated state. The dotted line in the figure indicates backtracking. When multiple caches are drawn on each layer, it represents the second batch, third batch of data added to the cache of this layer, rather than understanding that there are multiple caches on the same layer.
[0048] S101: Find Zhang San and Li Si, friends of Xiaoming. After adding them to the L1 result cache, L1 = {(Zhang San, 1), (Li Si, 1)}. At this time, all friends of Xiaoming have been stored in the L1 result cache.
[0049] S102: There are a total of 4 - layer adjacency queries. This layer is the first layer (i.e., not the penultimate layer), so it is the middle layer, and execute S103.
[0050] S103: Take the first element Zhang San in the L1 result cache, and then find his friends Xiaoming, Li Si, and Wang Wu. Since the count of Zhang San is 1, after multiplying by the count of the parent node, the count of his friends is still 1. Therefore, after adding them to the L2 result cache, L2 = {(Xiaoming, 1), (Li Si, 1), (Wang Wu, 1)}.
[0051] Take the second element Li Si in L1 and find his first friend Xiao Ming. Since Xiao Ming already exists in L2, after adding, L2 = {(Xiao Ming, 2), (Li Si, 1), (Wang Wu, 1)}; then find the second friend Zhang San of Li Si. Similarly, after adding, L2 = {(Xiao Ming, 2), (Li Si, 1), (Wang Wu, 1), (Zhang San, 1)}; then find the third friend Zhao Liu of Li Si. Since the cache size is 4 and L2 is full at this time, it cannot be added. Record the current cursor position of the L2 result cache as "Xiao Ming / Li Si / Zhao Liu".
[0052] At this time, it is judged that the L2 result cache is also the middle layer, so continue to collect the middle layer cache.
[0053] Take the first element Xiao Ming in L2 and find his friends Zhang San and Li Si. Since Xiao Ming's count is 2, after multiplying by 2 and adding to the L3 result cache, L3 = {(Zhang San, 2), (Li Si, 2)}; then take the second element Li Si in the L2 result cache and find his friends Xiao Ming, Zhang San, and Zhao Liu. And Li Si's count is 1. After adding, L3 = {(Zhang San, 3), (Li Si, 2), (Xiao Ming, 1), (Zhao Liu, 1)}; then take the third element Wang Wu in L2 and find his friends Zhang San and Zhao Liu and add them to L3. At this time, L3 = {(Zhang San, 4), (Li Si, 2), (Xiao Ming, 1), (Zhao Liu, 2)}; finally, take the fourth element Zhang San in the L2 result cache and find his friends Xiao Ming, Li Si, and Wang Wu. Since the first two friends already exist, after adding, L3 = {(Zhang San, 4), (Li Si, 3), (Xiao Ming, 2), (Zhao Liu, 2)}. Because the L3 cache size is 4 and it is full, record the current cursor position as "Xiao Ming / Li Si / Zhang San / Wang Wu".
[0054] At this time, it is judged that the L3 result cache is the penultimate layer, so enter S104.
[0055] S104: Take the first element Zhang San in the L3 result cache and find his friends Xiao Ming, Li Si, and Wang Wu. Since Zhang San's count is 4, after multiplying by the count of the parent node, the L4 result cache is L4 = {(Xiao Ming, 4), (Li Si, 4), (Wang Wu, 4)}; then take the second element Li Si in L3 and find his first friend Xiao Ming. And since his count is 3, after adding to the cache, L4 = {(Xiao Ming, 7), (Li Si, 4), (Wang Wu, 4)}, then find his second friend Zhang San and add to the cache, getting L4 = {(Xiao Ming, 7), (Li Si, 4), (Wang Wu, 4), (Zhang San, 3)}. Then find his third friend Zhao Liu, but at this time the L4 cache is full. Record the current position of the cursor, and then return the first batch of results {(Xiao Ming, 7), (Li Si, 4), (Wang Wu, 4), (Zhang San, 3)} to the user.
[0056] Then clear the L4 result cache, and then add the sixth friend of Li Si, Zhao Liu, to it. Since Li Si's count is 3, so L4 = {(Zhao Liu, 3)}; Then take the third element of L3, Xiao Ming, find his friends Zhang San and Li Si, and since Xiao Ming's count is 2, after adding to L4, L4 = {(Zhao Liu, 3), (Zhang San, 2), (Li Si, 2)}; Then take the fourth element of L3, Zhao Liu, and similarly add Zhao Liu's friends to the L4 result cache, then L4 = {(Zhao Liu, 3), (Zhang San, 2), (Li Si, 4), (Wang Wu, 2)}; At this time, the traversal of L3 is completed, clear L3, and return to the upper layer.
[0057] Repeat similar steps until L1 is empty. <I
[0058] Therefore, add the cursor position of L3 recorded above, "Xiao Ming / Li Si / Zhang San / Wang Wu", to it. Since the repetition count of Zhang San in L2 is 1, the appearance count of his friends after multiplying by the parent node count is still 1. Therefore, L3 = {(Wang Wu, 1)}; At this time, the traversal of L2 is completed; Clear L2, and add the "Xiao Ming / Li Si / Zhao Liu" pointed to by the cursor of L2 recorded above to L2. At this time, L2 = {(Zhao Liu, 1)}; Add Wang Wu and Li Si, the friends of Zhao Liu, to L3. At this time, L3 = {(Wang Wu, 2), (Li Si, 1)}.
[0059] Take the first element Wang Wu in L3, find his friends Zhang San and Zhao Liu. Since Wang Wu's count is 2, the appearance count of his friends is also 2. After adding, L4 = {(Zhao Liu, 5), (Zhang San, 4), (Li Si, 4), (Wang Wu, 2)}; Then take the second element Li Si in L3, find his first friend Xiao Ming. At this time, the L4 cache is full, so return {(Zhao Liu, 5), (Zhang San, 4), (Li Si, 4), (Wang Wu, 2)} as the second batch of results to the user and clear L4; Continue to add Xiao Ming, the friend of the second element Li Si in L3, to L4. At this time, L4 = {(Xiao Ming, 1)}; Similarly, add the other two friends of Li Si, Zhang San and Zhao Liu, to L4. At this time, L4 = {(Xiao Ming, 1), (Zhang San, 1), (Zhao Liu, 1)}. At this time, the traversal of L3 is completed, and when returning to L2 and L1, it is found that both have been traversed.
[0060] S105: At this time, the starting point Xiao Ming has no new adjacent nodes, so return the third batch of results in L4, {(Xiao Ming, 1), (Zhang San, 1), (Zhao Liu, 1)}, to the user. Thus, the program result is that the user obtains all the query results.
[0061] To verify the accuracy of the method in this embodiment, the traditional method is now used for adjacent query, including: starting from the vertex Xiao Ming for searching, finding his friend Zhang San, then finding Zhang San's friend Li Si, then finding Li Si's friend Zhao Liu, and finally finding Zhao Liu's friend Wang Wu. This is just one four - layer adjacent query route. All four - layer adjacent query routes are asFigure 5 As shown, in the final result, Xiaoming appears 8 times, Zhang San appears 8 times, Li Si appears 8 times, Wang Wu appears 6 times, and Zhao Liu appears 6 times, for a total of 36 pieces of data.
[0062] The results obtained by using the method of this embodiment are the same as those obtained by using the traditional method. Xiaoming appears 8 times, Zhang San appears 8 times, Li Si appears 8 times, Wang Wu appears 6 times, and Zhao Liu appears 6 times, for a total of 36 pieces of data. The results are correct.
[0063] Embodiment 2:
[0064] This embodiment discloses a multi-layer adjacency query optimization implementation system, which is a system that performs adjacency queries using the method disclosed in the above embodiment.
[0065] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0066] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0067] The above-described embodiments only represent several implementation manners of this application. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of the patent of this application should be subject to the appended claims.
Claims
1. An optimized implementation method for multi-layer adjacency query based on a graph storage model, characterized in that: Including the following steps: Step 1: Determine the total number of layers n of the adjacency query according to the adjacency query request; Step 2: Starting from the starting vertex, through the query graph storage model, store the queried adjacency point information into the first-layer result cache L1. In the result cache, the adjacency point information is cached with <vertex, count> as an element; where the count represents the number of times the vertex is repeated; until the first-layer result cache L1 is full or all the adjacency points of the starting vertex are stored in the first-layer result cache L1; then go to Step 3; Step 3: Determine whether the current result cache is the result cache L of the (n - 1)-th layer. n-1 If not, go to Step 4; if so, go to Step 5. Step 4: Denote the current result cache as L i , where i = 1, 2, 3... n - 2, take out the first element from the current result cache L i . Search for the adjacent nodes of the first element. If the found adjacent nodes have not been stored in the result cache L i+1 , then the count of the found adjacent nodes follows that of the parent node and is stored in the result cache L i+1 . If the found adjacent nodes are already in the result cache L i+1 , then the count of the found adjacent nodes is accumulated on the basis of following the count of the parent node and is stored in the result cache L i+1 ; Take out the second element from the current result cache L i . Search for the adjacent nodes of the second element. If the found adjacent nodes have not been stored in the result cache L i+1 , then the count of the found adjacent nodes follows that of the parent node and is stored in the result cache L i+1 . If the found adjacent nodes are already in the result cache L i+1 , then the count of the found adjacent nodes is accumulated on the basis of following the count of the parent node and is stored in the result cache L i+1 ; And so on until the result cache L i+1 is filled or all the adjacent nodes of all elements in the current result cache L i have been stored in the result cache L i+1 , then go to Step 3; During the caching process, if one or more adjacent nodes of one or more elements cannot be cached to the next-level result cache due to the result cache being full, record the current position through a cursor; Step 5: Take out the first element from the result cache L of the (n - 1)-th layer n-1 and find the adjacent nodes of the first element. If the found adjacent nodes have not been stored in the result cache L n , then the count of the found adjacent nodes follows the count of the parent node and is stored in the result cache L n . If the found adjacent nodes are already in the result cache L n , then the count of the found adjacent nodes is accumulated on the basis of following the count of the parent node and is stored in the result cache L n ; Take out the second element from the result cache L of the (n - 1)-th layer n-1 and find the adjacent nodes of the second element. If the found adjacent nodes have not been stored in the result cache L n , then the count of the found adjacent nodes follows the count of the parent node and is stored in the result cache L n . If the found adjacent nodes are already in the result cache L n , then the count of the found adjacent nodes is accumulated on the basis of following the count of the parent node and is stored in the result cache L n ; And so on until the result cache L is filled for the first time n or all the adjacent nodes of all the elements in the result cache L of the (n - 1)-th layer n-1 have been completely stored in the result cache L n , return the result in the current result cache L n to the user; During the caching process, if one or more adjacent nodes of one or more elements cannot be cached to the next-layer result cache due to the result cache being full, record the current position through a cursor; go to Step 6; Step 6: Clear the result cache L n , according to the cursor position corresponding to the result cache L of the (n - 1)-th layer n-1 , find the corresponding adjacent points of one or more un-traversed elements in sequence, update the count of the adjacent points, and store them in the result cache L n ; go to Step 7; Step 7: Determine whether the result cache L n is full at this time. If it is full, return the result in the current result cache L n to the user, and then clear the result cache L n ; if it is not full, go to Step 8; Step 8: Determine whether the elements in the current result cache have been traversed completely. If so, clear the current result cache, start searching upward from the upper-level result cache to find the cursor position. According to the cursor position, find the corresponding adjacent points of one or more untraversed elements in sequence, update the count of the adjacent points, and store them in the next-level result cache. And so on until the result cache L n is filled and returned to the user or when the first-level result cache L1 is empty, return the result of the current result cache L n to the user.
2. The method for optimizing the implementation of multi-layer adjacency query based on a graph storage model according to claim 1, wherein: In Step 4, the count of the found adjacency points is accumulated on the basis of following the count of the parent node, including: Assume that the count of the kth element is a, the adjacency points of the kth element include vertex s, the count of the qth element is b, the adjacency points of the qth element include vertex s, and assume that vertex s does not exist in the result cache to be cached; then the count of vertex s cached in the result cache is a + b.
3. The method for optimizing the implementation of multi-layer adjacency query based on a graph storage model according to claim 1, characterized in that: Return the result cached in L n to the user, including: Return the vertices in result cache L n to the user.
4. A multi - layer adjacency query optimization implementation system based on a graph storage model, characterized in that, The system includes a network interface, a memory, and a processor; where The network interface is used for receiving and sending signals during the process of receiving and sending information with other external network elements; The memory is used for storing computer program instructions that can run on the processor; The processor is used for executing the steps of a method for optimizing the implementation of a multi-layer adjacency query based on a graph storage model according to any one of claims 1 to 3 when running the computer program instructions.
5. A computer storage medium, characterized in that, The computer storage medium stores a program of a method for optimizing the implementation of a multi-layer adjacency query based on a graph storage model. When the program of the method for optimizing the implementation of a multi-layer adjacency query based on a graph storage model is executed by at least one processor, the steps of a method for optimizing the implementation of a multi-layer adjacency query based on a graph storage model according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
A caching method for a query intermediate result set of a distributed database system
CN109947796A
Graph data path retrieval processing method and device, server and storage medium
CN110941741A