A Fast Multi-Layer Line Expansion Method for Graph Data Based on BFS + DFS
By adopting the combination method of BFS+DFS in the graph database, combining forward and reverse search, and introducing a backtracking mechanism, the problem of large resource occupation and performance degradation in multi-layer line expansion operations is solved, and efficient relationship object network search is achieved.
Patent Information
- Application Number
- CN202211094038.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-09-08
AI Technical Summary
Traditional relational databases cannot effectively process complex related network data. Graph databases have problems such as large resource usage and performance degradation in multi-layer line expansion operations, especially during deep query.
A fast multi-layer line expansion method based on BFS+DFS is adopted. By presetting the object set and associated attributes of the target type, combining forward and reverse search, a backtracking mechanism is introduced, the search path is optimized, and efficiency is improved.
It effectively improves the search efficiency of the relational object network, reduces resource usage, avoids OutOfMemory problems, and improves system performance.
Smart Images

Figure CN116304195B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for rapidly expanding multi - layer lines of graph data based on BFS + DFS, belonging to the technical field of graph databases. Background Art
[0002] With the rapid development of the Internet of Things (IoT) technology, a vast amount of data that requires correlation analysis has emerged in various industries. However, traditional relational databases cannot meet the storage, query, and analysis of complex correlation network data, and graph databases exhibit unprecedented performance advantages.
[0003] Regarding the unique multi - layer line expansion problem of graph data, since the data increases exponentially as the query level deepens, resource occupancy will increase significantly, and performance will drop significantly. For example, when querying the 5 - layer relationships of 10 nodes, with an average of 100 relationships per node, the number of relationships retrieved in the first layer is 10 * 100 = 1000, in the second layer is 10 * 100^2 = 100000, in the third layer is 10 * 100^3 = 10000000, in the fourth layer is 10 * 100^4 = 1000000000, … Without considering duplicate relationship data, the relationship data at the fourth - layer query has reached the billion - level. If there are super nodes, the data volume will increase exponentially.
[0004] Multi - layer line expansion belongs to search - type operations. Generally, there are two methods: DFS (Depth - First Search) and BFS (Breadth - First Search).
[0005] DFS is to brute - force search all paths. It uses the backtracking method, saves the current position, searches deeply, and when all searches are completed, it backtracks to search the next position until all the deepest positions are searched. Many of the paths found may be useless. However, since it only searches deeply in one direction at a time, it will not cause resource tension problems, but it may take a long time.
[0006] BFS is to find all possible paths at each layer, then select one direction and continue to find all possible paths. Next, select another direction from the previous layer and continue to find all possible paths until all directions at each layer are traversed. Since more and more paths of the first N layers need to be cached as the layer deepens, it is easy to cause resource problems, but the search time is less.
[0007] It can be seen from this that DFS trades time for space, and BFS trades space for time. The traversal efficiency of DFS is too slow, and BFS is prone to causing server - side resource tension and even OutOfMemory problems. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a fast multi-layer line expansion method for graph data based on BFS+DFS, integrate the connection between the two methods, introduce a new search strategy, and can effectively improve the search efficiency of the relationship object network.
[0009] The present invention adopts the following technical solutions to solve the above technical problems: The present invention designs a fast multi-layer line expansion method for graph data based on BFS+DFS. Based on a preset target type object set, according to the relationship search request under a preset association attribute, it realizes the search for the relationship object network under the preset association attribute starting from the target object in the preset target type object set; includes the following steps:
[0010] Step A. Based on the equal division of the preset target type object set, obtain each preset target type object subset, then initialize i = 1 and the result path set as an empty set, and enter Step B;
[0011] Step B. According to the preset maximum relationship layer number K greater than 1, set each target object in the i-th preset target type object subset as the 0th relationship layer, initialize the 1st relationship layer to the K-1th relationship layer corresponding to the i-th preset target type object subset, and initialize n = 1, then enter Step C;
[0012] Step C. For the nth target object in the i-th preset target type object subset, perform the relationship search corresponding to the preset maximum relationship layer number K. If the relationship search corresponding to the preset maximum relationship layer number K for the nth target object is successfully completed, obtain the relationship network corresponding to the preset maximum relationship layer number K for the nth target object, and enter Step D; if the relationship search corresponding to the preset maximum relationship layer number K for the nth target object is not successfully completed, enter Step F;
[0013] Step D. For the relationship network corresponding to the preset maximum relationship layer number K of the latest obtained nth target object, search for each relationship object path corresponding to the nth target object from the K-1th relationship layer to the 0th relationship layer direction, add it to the result path set, and judge whether the number of relationship object paths in the result path set meets the path number requirement in the relationship search request. If so, complete the search for the relationship object network under the preset association attribute starting from the target object; otherwise, enter Step E;
[0014] Step E. For the relationship network corresponding to the preset maximum relationship layer number K of the nth target object, perform a new relationship search corresponding to the preset maximum relationship layer number K in the direction from the K-1th relationship layer to the 0th relationship layer. If the new relationship search corresponding to the preset maximum relationship layer number K for the nth target object is successfully completed, obtain the relationship network corresponding to the preset maximum relationship layer number K for the nth target object, and return to Step D; if the new relationship search corresponding to the preset maximum relationship layer number K for the nth target object is not successfully completed, enter Step F;
[0015] Step F. Determine whether n is equal to the number of target objects in the i-th preset target type object subset. If yes, go to Step G; otherwise, update n by adding 1 and return to Step C.
[0016] Step G. Determine whether the value of i is equal to the number of preset target type object subsets. If yes, end the search for the relationship object network starting from the target object in the preset target type object set under the preset associated attribute; otherwise, update the value of i by adding 1, initialize n = 1, and then return to Step C.
[0017] As a preferred technical solution of the present invention: Step C is executed according to the following Step C1 to Step C8;
[0018] Step C1. Initialize k = 1, and select the n-th target object in the i-th preset target type object subset as the object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object, and go to Step C2;
[0019] Step C2. Search and determine whether there are associated relationship objects for each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset associated attribute. If yes, obtain the relationship objects associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset associated attribute, and go to Step C3; otherwise, go to Step C6;
[0020] Step C3. Determine whether k is equal to K - 1. If yes, go to Step C5; otherwise, go to Step C4;
[0021] Step C4. Determine whether the number of relationship objects associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset associated attribute is greater than the preset analysis upper limit L. If yes, randomly select L relationship objects from the associated relationship objects as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, update k by adding 1, and then return to Step C2; otherwise, directly use the associated relationship objects as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, update k by adding 1, and then return to Step C2;
[0022] Step C5. Use the relationship objects associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset associated attribute as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, that is, successfully complete the relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K, obtain the relationship network of the n-th target object corresponding to the preset maximum number of relationship layers K, and go to Step D;
[0023] Step C6. Determine whether k - 1 is equal to 0. If so, it means that the relationship search for the preset maximum relationship layer K corresponding to the nth target object has not been successfully completed, and then proceed to step F; otherwise, proceed to step C7;
[0024] Step C7. Determine whether there are any relationship objects among the relationship objects associated with each object to be analyzed in the (k - 2)th relationship layer corresponding to the nth target object with respect to the preset association attribute that are not the objects to be analyzed in the (k - 1)th relationship layer corresponding to the nth target object. If so, only use these relationship objects as the relationship objects associated with each object to be analyzed in the (k - 2)th relationship layer corresponding to the nth target object with respect to the preset association attribute, then update the value of k by subtracting 1, and then return to step C4; otherwise, proceed to step C8;
[0025] Step C8. Update the value of k by subtracting 1, and determine whether k - 2 is less than 0. If so, it means that the relationship search for the preset maximum relationship layer K corresponding to the nth target object has not been successfully completed, and then proceed to step F; otherwise, return to step C7.
[0026] As a preferred technical solution of the present invention: The said step E is executed according to the following steps E1 to E9;
[0027] Step E1. Initialize g = K - 1, and proceed to step E2;
[0028] Step E2. Determine whether g - 2 is less than 0. If so, it means that the new relationship search for the preset maximum relationship layer K corresponding to the nth target object has not been successfully completed, and then proceed to step F; otherwise, proceed to step E3;
[0029] Step E3. Determine whether there are any relationship objects among the relationship objects associated with each object to be analyzed in the (g - 2)th relationship layer corresponding to the nth target object with respect to the preset association attribute that are not the objects to be analyzed in the (g - 1)th relationship layer corresponding to the nth target object. If so, only use these relationship objects as the relationship objects associated with each object to be analyzed in the (g - 2)th relationship layer corresponding to the nth target object with respect to the preset association attribute, then update the value of g by subtracting 1, and then proceed to step E5; otherwise, proceed to step E4;
[0030] Step E4. Update the value of g by subtracting 1, and return to step E2;
[0031] Step E5. Determine whether the number of relationship objects associated with each object to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute is greater than the preset analysis upper limit L. If so, randomly select L relationship objects from the associated relationship objects as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, increment the value of g by 1, and then proceed to Step E6; otherwise, directly use the associated relationship objects as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, increment the value of g by 1, and then proceed to Step E6;
[0032] Step E6. Search and determine whether there are any associated relationship objects for each object to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute. If so, obtain the relationship objects associated with each object to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute, and proceed to Step E7; otherwise, proceed to Step E9;
[0033] Step E7. Determine whether g is equal to K - 1. If so, proceed to Step E8; otherwise, return to Step E5;
[0034] Step E8. Use the relationship objects associated with each object to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, that is, successfully complete the new relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K, obtain the relationship network of the n-th target object corresponding to the preset maximum number of relationship layers K, and return to Step D;
[0035] Step E9. Determine whether g - 1 is equal to 0. If so, it means that the new relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K is not successfully completed, and then proceed to Step F; otherwise, return to Step E3.
[0036] As a preferred technical solution of the present invention: in the process of obtaining the relationship objects associated with each object to be analyzed in the relationship layer with respect to the preset association attribute, if the number of objects to be analyzed in this relationship layer is less than or equal to the preset maximum task data MAX_TASK_NUM that can be started, directly search for the relationship objects associated with each object to be analyzed in this relationship layer with respect to the preset association attribute for each object to be analyzed in this relationship layer;
[0037] If the number of objects to be analyzed in this relationship layer is greater than the preset maximum task data MAX_TASK_NUM that can be started, each time search for the relationship objects associated with each object to be analyzed in this relationship layer with respect to the preset association attribute for no more than MAX_TASK_NUM objects to be analyzed in this relationship layer. Through each search, obtain the relationship objects associated with each object to be analyzed in this relationship layer with respect to the preset association attribute.
[0038] As a preferred technical solution of the present invention: according to a preset maximum relationship layer number K greater than 1, initialize the path pickers corresponding to each relationship layer from the 1st relationship layer to the (K - 1)th relationship layer, and each path picker of each relationship layer respectively searches for the relationship objects associated with each object to be analyzed in the previous relationship layer according to the preset association attributes.
[0039] As a preferred technical solution of the present invention: the preset target type object set is a literature set, and the corresponding preset association attribute is author association; or the preset target type object set is a target population set, and the corresponding preset association attribute is call association.
[0040] Compared with the prior art, the fast multi - layer line expansion method for graph data based on BFS + DFS of the present invention has the following technical effects by adopting the above technical solutions:
[0041] The fast multi - layer line expansion method for graph data based on BFS + DFS designed by the present invention is based on a preset target type object set. First, it is designed to perform a search of the relationship network of the preset maximum relationship layer number K corresponding to each target object in a forward search manner; then, it performs a search of the relationship object network starting from each target object and regarding the preset association attributes in a reverse search manner; and in the forward search and the reverse search, in order to meet the requirements of the corresponding maximum relationship layer number K, a backtracking mechanism is introduced. By designing such a logical strategy and adding the specific search methods of BFS and DFS, finally, a search of the relationship object network starting from the target objects in the preset target type object set and regarding the preset association attributes is realized, overcoming the deficiencies of the prior art and effectively improving the search efficiency of the relationship object network search. Brief Description of the Drawings
[0042] Figure 1 is a schematic diagram of the fast multi - layer line expansion method for graph data based on BFS + DFS designed by the present invention;
[0043] Figure 2 is an application schematic diagram of step C4 and step E5 designed by the present invention;
[0044] Figure 3 is an application schematic diagram of the embodiment designed by the present invention. Detailed Description of the Invention
[0045] The following further details the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.
[0046] The present invention designs a fast multi-layer line expansion method for graph data based on BFS+DFS. Based on a preset target type object set, according to a relationship search request under a preset associated attribute, it realizes the search for a relationship object network starting from a target object in the preset target type object set and regarding the relationship under the preset associated attribute. In practical applications, depth-first search is generally adopted. When searching for a single-layer path, breadth-first search is used. That is, when discovering vertices in the nth layer, breadth-first search is performed on the vertices in the (n-1)th layer until the searched path exceeds the threshold (i.e., the layer path threshold, which refers to the maximum amount of paths allowed to be discovered in a single complete search of a single-layer path), then the search stops and depth-first search continues. When no reachable vertices can be found in the nth layer or the deepest layer has been reached, it backtracks to the (n-1)th layer and repeats the above process until all paths are searched out.
[0047] In practical applications, the specific implementation is designed as follows in steps A to G.
[0048] Step A. Based on the equal division of the preset target type object set, obtain each preset target type object subset. Then initialize i = 1 and the result path set as an empty set, and enter step B.
[0049] Step B. According to the preset maximum relationship layer number K greater than 1, initialize the 1st to (K-1)th relationship layers corresponding to the i-th preset target type object subset with each target object in the i-th preset target type object subset corresponding to the 0th relationship layer, and initialize n = 1, then enter step C.
[0050] Step C. For the nth target object in the i-th preset target type object subset, perform a relationship search corresponding to the preset maximum relationship layer number K. If the relationship search corresponding to the nth target object for the preset maximum relationship layer number K is successfully completed, obtain the relationship network corresponding to the nth target object for the preset maximum relationship layer number K, and enter step D; if the relationship search corresponding to the nth target object for the preset maximum relationship layer number K is not successfully completed, enter step F.
[0051] In practical applications, the above step C is specifically designed and implemented as follows in steps C1 to C8.
[0052] Step C1. Initialize k = 1, and select the nth target object in the i-th preset target type object subset as the object to be analyzed corresponding to the (k-1)th relationship layer of the nth target object, and enter step C2.
[0053] Step C2. Search and determine whether there are any related relationship objects for each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute. If so, obtain the relationship objects associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute, and proceed to Step C3; otherwise, proceed to Step C6.
[0054] Step C3. Determine whether k is equal to K - 1. If so, proceed to Step C5; otherwise, proceed to Step C4.
[0055] Step C4. Determine whether the number of relationship objects associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute is greater than the preset analysis upper limit L. If so, randomly select L relationship objects from the associated relationship objects as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, update the value of k by adding 1, and then return to Step C2; otherwise, directly use the associated relationship objects as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, update the value of k by adding 1, and then return to Step C2.
[0056] Step C5. Use the relationship objects associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, that is, successfully complete the relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K, obtain the relationship network of the n-th target object corresponding to the preset maximum number of relationship layers K, and proceed to Step D.
[0057] Step C6. Determine whether k - 1 is equal to 0. If so, it means that the relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K is not successfully completed, and then proceed to Step F; otherwise, proceed to Step C7.
[0058] Step C7. Determine whether there are any relationship objects among the relationship objects associated with each object to be analyzed in the (k - 2)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute that are not used as the objects to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object. If so, only use these relationship objects as the relationship objects associated with each object to be analyzed in the (k - 2)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute, then update the value of k by subtracting 1, and then return to Step C4; otherwise, proceed to Step C8.
[0059] Step C8. Update the value of k by subtracting 1, and determine whether k - 2 is less than 0. If so, it means that the relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K is not successfully completed, and then proceed to Step F; otherwise, return to Step C7.
[0060] Step D. For the relationship network of the nth target object corresponding to the preset maximum relationship level K obtained most recently, search for the relationship object paths corresponding to the nth target object from the (K - 1)th relationship level to the 0th relationship level in the direction towards the 0th relationship level, add them to the result path set, and determine whether the number of relationship object paths in the result path set meets the path number requirement in the relationship search request. If yes, complete the search for the relationship object network starting from the target object and regarding the preset association attributes; otherwise, proceed to Step E.
[0061] Step E. For the relationship network of the nth target object corresponding to the preset maximum relationship level K, perform a new relationship search corresponding to the preset maximum relationship level K from the (K - 1)th relationship level to the 0th relationship level. If the new relationship search for the nth target object corresponding to the preset maximum relationship level K is successfully completed, obtain the relationship network of the nth target object corresponding to the preset maximum relationship level K, and return to Step D; if the new relationship search for the nth target object corresponding to the preset maximum relationship level K is not successfully completed, then proceed to Step F.
[0062] In practical applications, regarding the above Step E, the specific design and execution are as follows: Steps E1 to E9.
[0063] Step E1. Initialize g = K - 1, and proceed to Step E2.
[0064] Step E2. Determine whether g - 2 is less than 0. If yes, it means that the new relationship search for the nth target object corresponding to the preset maximum relationship level K is not successfully completed, and then proceed to Step F; otherwise, proceed to Step E3.
[0065] Step E3. Determine whether there are any relationship objects among the relationship objects associated with each object to be analyzed in the (g - 2)th relationship level corresponding to the nth target object regarding the preset association attributes that are not the relationship objects of the objects to be analyzed in the (g - 1)th relationship level corresponding to the nth target object. If yes, only use these relationship objects as the relationship objects associated with each object to be analyzed in the (g - 2)th relationship level corresponding to the nth target object regarding the preset association attributes, then update the value of g by subtracting 1, and then proceed to Step E5; otherwise, proceed to Step E4.
[0066] Step E4. Update the value of g by subtracting 1, and return to Step E2.
[0067] Step E5. Determine whether the number of relationship objects associated with the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute is greater than the preset analysis upper limit L. If so, randomly select L relationship objects from the associated relationship objects as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, increment the value of g by 1, and then proceed to Step E6; otherwise, directly use the associated relationship objects as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, increment the value of g by 1, and then proceed to Step E6.
[0068] Step E6. Search and determine whether there are associated relationship objects for the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute. If so, obtain the relationship objects associated with the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute, and proceed to Step E7; otherwise, proceed to Step E9.
[0069] Step E7. Determine whether g is equal to K - 1. If so, proceed to Step E8; otherwise, return to Step E5.
[0070] Step E8. Use the relationship objects associated with the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, that is, successfully complete the new relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K, obtain the relationship network of the n-th target object corresponding to the preset maximum number of relationship layers K, and return to Step D.
[0071] Step E9. Determine whether g - 1 is equal to 0. If so, it means that the new relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K is not successful, and then proceed to Step F; otherwise, return to Step E3.
[0072] Step F. Determine whether n is equal to the number of target objects in the i-th preset target type object subset. If so, proceed to Step G; otherwise, increment the value of n by 1 and return to Step C.
[0073] Step G. Determine whether the value of i is equal to the number of preset target type object subsets. If so, end the search for the relationship object network starting from the target objects in the preset target type object set with respect to the preset association attribute; otherwise, increment the value of i by 1, initialize n = 1, and then return to Step C.
[0074] In the actual implementation of the above design method, during the process of obtaining the relationship objects associated with each object to be analyzed in the relationship layer with respect to the preset association attribute, due to the excessive data volume and too deep query level during the multi-layer extended line query path, it will cause the OutOfMemory problem on the Server side. Combining the graph data characteristics, a combined traversal method of BFS+DFS is applied for concurrent optimization. Specifically, if the number of objects to be analyzed in this relationship layer is less than or equal to the preset maximum task data MAX_TASK_NUM that can be started, directly for each object to be analyzed in this relationship layer, find the relationship objects associated with each object to be analyzed with respect to the preset association attribute; if the number of objects to be analyzed in this relationship layer is greater than the preset maximum task data MAX_TASK_NUM that can be started, each time for no more than MAX_TASK_NUM objects to be analyzed in this relationship layer, find the relationship objects associated with each object to be analyzed with respect to the preset association attribute. Through each search, obtain the relationship objects associated with each object to be analyzed in this relationship layer with respect to the preset association attribute.
[0075] Apply the above-designed fast multi-layer extended line method for graph data based on BFS+DFS to the actual situation. According to the preset maximum number of relationship layers K greater than 1, initialize the path pickers corresponding to each relationship layer in the 1st relationship layer to the (K-1)th relationship layer. Each path picker of each relationship layer searches for the relationship objects associated with each object to be analyzed in the previous relationship layer according to the preset association attribute. The initial situation of the path picker selector is as follows.
[0076]
[0077] The relationship between path pickers is expressed in a form similar to a linked list; the path picker uses the last node of the path returned by the parent path picker as the starting node for query expansion. Only a fixed amount of data is cached at each layer. If the data volume reaches the limit number, the data is returned. If the specified layer has been searched, the relationships that have formed the response result path in the cache of each layer can be cleared, and new relationships can be cached for in-depth extended line.
[0078] Specifically, the path picker selector is divided into the 0th layer picker ZeroLevelSelector and the standard picker StandardLevelSelector. Because except for the selector of the 0th layer, there are both child and parent in other layers, and the 0th layer is the parent, that is, the starting node layer, which is the parent path for the first layer traversal. The path for the first layer traversal is the sub-path of the starting node. The specific linked list structure is as Figure 1 shown.
[0079] In the above steps C4 and E5, the judgment on whether the number of relationship objects associated with each object to be analyzed in the corresponding relationship layer of the target object with respect to the preset association attribute is greater than the preset analysis upper limit L is specifically as follows Figure 2 As shown; assume that the preset analysis upper limit L for each layer is 100 path results. Then, starting from the target object in Level-0, only one target object is traversed. If the relationship object path has reached 100, the traversal of the second target object will not continue; start the Level-1 traversal to expand the line from the termination node of the existing relationship object path to obtain 100 two-layer paths in Level-2; traverse deeper in turn until reaching the preset maximum relationship layer K, then start traversing the second node from the starting node, similar to the traversal of the first node.
[0080] In the actual implementation and application of the entire solution, the preset target type object set is a literature set, and the corresponding preset association attribute is author association; or the preset target type object set is a target population set, and the corresponding preset association attribute is call association; that is, the above design execution method is used to implement the search for the relationship object network under the preset association attribute starting from the target object in the preset target type object set of each type.
[0081] Regarding the query of tasks in the entire solution design, since the entire traversal process involves the start of multiple query threads, and in addition to traversal queries, there are various other algorithms and ordinary query tasks on the entire graph database server, and the time consumption of each task varies greatly. Therefore, according to the possible time consumption of the query tasks, the response threads of the graph server are decomposed and managed. For algorithmic analysis tasks, a long-task thread pool is applied, and for ordinary one-time query requests, a short-task thread pool is applied. If the number of tasks being executed in the short-task thread pool is equal to the core thread number and the number of tasks waiting to be executed is too large (exceeding half of the maximum thread number), an emergency thread pool can be rented. If the emergency thread pool is idle (the number of threads being executed is less than 75% of the maximum thread pool), renting is allowed, otherwise it is not allowed. With this refined task decomposition management method, the execution response of various different queries is improved, thereby improving the query efficiency and avoiding the impact on the response efficiency of other queries due to a certain query task being overly time-consuming and resource-consuming. The specific analysis is as follows.
[0082] 1. Since G concurrent tasks will be started in the relationship traversal iterator object in the Selector of each layer to iteratively obtain the RelationIterator data, a total of N*G tasks need to be started, and the tasks of each layer are not completed by running alone, but need to run according to the order of path picking in the Selector.
[0083] 2. In addition to traversal tasks, the Server side also has other query tasks and algorithm analysis tasks. General query tasks generally take less time and occupy fewer resources. Algorithm analysis tasks generally take longer and occupy more resources. The traverser task directly occupies N*G threads, but the time consumption within each task is short. It's just that the entire search process is a combined search of BFS+DFS, and there are dependencies between tasks at each level. A single task at a certain level cannot be directly run to completion.
[0084] 3. When the Server starts, initialize the resource management module; the resource management module contains three resource pools: long-task, short-task, and emergency-task resource pools, as Figure 3 shown.
[0085] 4. The long-task resource pool can use a cache queue to cache tasks. Since long tasks take a long time, the number of core threads is set relatively small, and the maximum number of threads can be set to 4 times the number of core threads; the short-task resource pool can also use a cache queue to cache tasks, but the number of core threads is slightly more. Since short tasks themselves take a short time, the maximum number of threads can be set to 2.5 times the number of core threads; the emergency resource pool can be leased on the basis that both the long-task and short-task thread pools are full, and both the number of core threads and the maximum number of threads are set relatively small.
[0086] 5. When a task request is sent to the resource management module, the task dispatcher determines whether the task is submitted and which resource pool it is submitted to; the specific resource pool is responsible for the execution and recycling of tasks and for responding to requests;
[0087] 6. The task monitor is responsible for:
[0088] Monitoring the number of threads being executed in each thread pool;
[0089] Monitoring the number of queues waiting to be executed in each thread pool;
[0090] Triggering the early release of some long tasks as needed.
[0091] The above technical solution designs a fast multi-layer graph data line expansion method based on BFS+DFS. Based on a preset target type object set, it is designed to first perform a search of the relationship network of the preset maximum relationship layer K corresponding to each target object in a forward search manner; then perform a search of the relationship object network starting from each target object and regarding the preset association attributes in a reverse search manner; and in the forward search and reverse search, in order to meet the requirements of the corresponding maximum relationship layer K, a backtracking mechanism is introduced. Such a design logic strategy is combined with the specific search methods of BFS and DFS, and finally realizes the search of the relationship object network starting from the target objects in the preset target type object set and regarding the preset association attributes, overcoming the deficiencies of the existing technology and effectively improving the search efficiency of the relationship object network search.
[0092] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments, and various changes can be made without departing from the spirit of the present invention within the scope of knowledge possessed by those of ordinary skill in the art.
Claims
1. A fast multi-layer line expansion method for graph data based on BFS + DFS, characterized in that: Based on a preset set of target type objects, according to a relationship search request under a preset associated attribute, implement a relationship object network search starting from a target object in the preset set of target type objects and regarding the relationship under the preset associated attribute; including the following steps: Step A. Based on the equal division of the preset set of target type objects, obtain each subset of the preset target type objects, then initialize i = 1 and the result path set as an empty set, and enter Step B; Step B. According to the preset maximum relationship level K greater than 1, with each target object in the i-th subset of the preset target type objects corresponding to the 0-th relationship level, initialize the 1-st relationship level to the K-1 relationship level corresponding to the i-th subset of the preset target type objects, and initialize n = 1, then enter Step C; Step C. For the n-th target object in the i-th subset of the preset target type objects, perform a relationship search corresponding to the preset maximum relationship level K. If the relationship search corresponding to the preset maximum relationship level K for the n-th target object is successfully completed, obtain the relationship network corresponding to the preset maximum relationship level K for the n-th target object, and enter Step D; If the relationship search corresponding to the preset maximum relationship level K for the n-th target object is not successfully completed, enter Step F; Step D. For the relationship network corresponding to the preset maximum relationship level K for the n-th target object obtained latest, search for each relationship object path corresponding to the n-th target object from the K-1 relationship level to the 0-th relationship level direction, add it to the result path set, and determine whether the number of relationship object paths in the result path set meets the path number requirement in the relationship search request. If so, complete the relationship object network search starting from the target object and regarding the relationship under the preset associated attribute; otherwise enter Step E; Step E. For the relationship network corresponding to the preset maximum relationship level K for the n-th target object, perform a new relationship search corresponding to the preset maximum relationship level K in the direction from the K-1 relationship level to the 0-th relationship level. If the new relationship search corresponding to the preset maximum relationship level K for the n-th target object is successfully completed, obtain the relationship network corresponding to the preset maximum relationship level K for the n-th target object, and return to Step D; If the new relationship search corresponding to the preset maximum relationship level K for the n-th target object is not successfully completed, enter Step F; Step F. Determine whether n is equal to the number of target objects in the i-th subset of the preset target type objects. If so, enter Step G; Otherwise, update n by adding 1 and return to Step C; Step G. Determine whether the value of i is equal to the number of subsets of the preset target type objects. If so, end the relationship object network search starting from the target object in the preset set of target type objects and regarding the relationship under the preset associated attribute; otherwise update the value of i by adding 1, initialize n = 1, and then return to Step C.
2. The method for quickly expanding lines in multiple layers of graph data based on BFS+DFS according to claim 1, wherein: The said Step C is executed according to the following Step C1 to Step C8; Step C1. Initialize k = 1, and select the n-th target object in the i-th subset of the preset target type objects as the object to be analyzed corresponding to the k-1 relationship level of the n-th target object, and enter Step C2; Step C2. Search and determine whether there are any related relationship objects for each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute. If so, obtain each relationship object associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute, and proceed to Step C3; Otherwise, proceed to Step C6; Step C3. Determine whether k is equal to K - 1. If so, proceed to Step C5; Otherwise, proceed to Step C4; Step C4. Determine whether the number of relationship objects associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute is greater than the preset analysis upper limit L. If so, randomly select L relationship objects from the associated relationship objects as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, update the value of k by adding 1, and then return to Step C2; Otherwise, directly use the associated relationship objects as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, update the value of k by adding 1, and then return to Step C2; Step C5. Use each relationship object associated with each object to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute as the objects to be analyzed in the k-th relationship layer corresponding to the n-th target object, that is, successfully complete the relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K, obtain the relationship network of the n-th target object corresponding to the preset maximum number of relationship layers K, and proceed to Step D; Step C6. Determine whether k - 1 is equal to 0. If so, it means that the relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K has not been successfully completed, and then proceed to Step F; otherwise, proceed to Step C7; Step C7. Determine whether there are any relationship objects among the relationship objects associated with each object to be analyzed in the (k - 2)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute that have not been used as the objects to be analyzed in the (k - 1)-th relationship layer corresponding to the n-th target object. If so, only use these relationship objects as the relationship objects associated with each object to be analyzed in the (k - 2)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute, then update the value of k by subtracting 1, and then return to Step C4; otherwise, proceed to Step C8; Step C8. Update the value of k by subtracting 1, and determine whether k - 2 is less than 0. If so, it means that the relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K has not been successfully completed, and then proceed to Step F; otherwise, return to Step C7.
3. A fast multi-layer line expansion method for graph data based on BFS+DFS according to claim 1, characterized in that: The said Step E is executed according to the following Steps E1 to E9; Step E1. Initialize g = K - 1, and proceed to Step E2; Step E2. Determine whether g - 2 is less than 0. If so, it means that the new relationship search for the n-th target object corresponding to the preset maximum number of relationship layers K has not been successfully completed, and then proceed to Step F; otherwise, proceed to Step E3; Step E3. Determine whether there are any relationship objects among the objects to be analyzed in the (g - 2)-th relationship layer corresponding to the n-th target object that are associated with the preset association attribute and are not the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object. If so, only use these relationship objects as the objects to be analyzed in the (g - 2)-th relationship layer corresponding to the n-th target object that are associated with the preset association attribute, then update by decrementing the value of g by 1, and then proceed to Step E5; otherwise, proceed to Step E4; Step E4. Update by decrementing the value of g by 1, and return to Step E2; Step E5. Determine whether the number of relationship objects associated with the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute is greater than the preset analysis upper limit L. If so, randomly select L relationship objects from the associated relationship objects as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, then update by incrementing the value of g by 1, and then proceed to Step E6; otherwise, directly use the associated relationship objects as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, then update by incrementing the value of g by 1, and then proceed to Step E6; Step E6. Search and determine whether there are any associated relationship objects for the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute. If so, obtain the relationship objects associated with the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute, and proceed to Step E7; otherwise, proceed to Step E9; Step E7. Determine whether g is equal to K - 1. If so, proceed to Step E8; otherwise, return to Step E5; Step E8. Use the relationship objects associated with the objects to be analyzed in the (g - 1)-th relationship layer corresponding to the n-th target object with respect to the preset association attribute as the objects to be analyzed in the g-th relationship layer corresponding to the n-th target object, that is, successfully complete the new relationship search for the n-th target object corresponding to the preset maximum relationship layer number K, obtain the relationship network of the n-th target object corresponding to the preset maximum relationship layer number K, and return to Step D; Step E9. Determine whether g - 1 is equal to 0. If so, it means that the new relationship search for the n-th target object corresponding to the preset maximum relationship layer number K has not been successfully completed, and then proceed to Step F; otherwise, return to Step E3.
4. A fast multi-layer line expansion method for graph data based on BFS+DFS according to any one of claims 1 to 3, characterized in that: In the process of obtaining the relationship objects associated with the objects to be analyzed in each relationship layer with respect to the preset association attribute, if the number of objects to be analyzed in this relationship layer is less than or equal to the preset maximum task data that can be started MAX_TASK_NUM, then directly search for the relationship objects associated with each object to be analyzed in this relationship layer with respect to the preset association attribute; If the number of objects to be analyzed in the relationship layer is greater than the preset maximum task data MAX_TASK_NUM that can be started, then each time for no more than MAX_TASK_NUM objects to be analyzed in the relationship layer, find the relationship objects associated with each object to be analyzed with respect to the preset association attribute. Through each search, obtain the relationship objects associated with each object to be analyzed in the relationship layer with respect to the preset association attribute.
5. A method for fast multi-layer line expansion of graph data based on BFS+DFS according to any one of claims 1 to 3, characterized in that: According to the preset maximum number of relationship layers K greater than 1, initialize the path pickers corresponding to each relationship layer in the first relationship layer to the K - 1 relationship layer. The path pickers of each relationship layer respectively find the relationship objects associated with the objects to be analyzed in the previous relationship layer according to the preset association attribute.
6. A fast multi - layer line expansion method for graph data based on BFS + DFS according to any one of claims 1 to 3, characterized in that: The preset target type object set is a literature set, and the corresponding preset association attribute is author association; Or the preset target type object set is a target population set, and the corresponding preset association attribute is call association.
Citation Information
Patent Citations
A method and a system for tracing multi-level association paths of two enterprises
CN106126614A
Data processing method and device for user relationship strength, computer equipment and storage medium
CN111274495A