Graph data processing method and device, computer equipment, storage medium and computer program product
By dividing the target nodes in the graph data into sub-blocks and storing them in a binary indexed tree, the construction process of the index structure is simplified, the problem of high memory overhead in existing technologies is solved, and efficient weighted sampling is achieved.
Patent Information
- Application Number
- CN202410567358.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-11-11
AI Technical Summary
Existing graph data processing methods are complex and memory-intensive when building index information, resulting in excessive resource consumption.
By identifying target nodes in the graph data, sub-blocks are divided according to the identifiers of neighboring nodes, and the sampling weights of neighboring nodes are stored in the sub-blocks in the form of a binary indexed tree to construct an index structure for weighted sampling.
It simplifies the sub-block partitioning process, reduces memory overhead, and improves the efficiency and speed of neighbor node sampling.
Smart Images

Figure CN120929642A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a graph data processing method, apparatus, computer equipment, storage medium, and computer program product. Background Technology
[0002] With the development of artificial intelligence technology, the application scenarios of graph data are becoming increasingly widespread. For example, in applications such as social networking, e-commerce, and the Internet of Things, vast and complex relationship networks are formed, and the number of nodes in graph data is also growing exponentially. Taking the training phase of a graph neural network framework as an example, during training, it is necessary to sample the nodes of the graph data, obtain the features of the graph nodes, and then perform graph convolution processing.
[0003] However, in order to achieve weighted sampling of graph data, it is necessary to build index information based on the weight data of nodes in the graph data. However, the existing index information construction process is complex and has a large memory overhead. Summary of the Invention
[0004] Therefore, it is necessary to provide a graph data processing method, apparatus, computer device, computer-readable storage medium, and computer program product that can reduce memory overhead in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a graph data processing method. The method includes:
[0006] Identify target nodes with neighboring nodes from graph data;
[0007] For each target node, sub-blocks are divided according to the identifiers of each neighboring node to determine the sub-block to which each neighboring node belongs; wherein, the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block.
[0008] The sampling weights of each neighboring node belonging to the same sub-block are stored in the sub-block in the form of a binary indexed tree;
[0009] Based on each of the sub-blocks, an index structure is constructed for the target node, and the index structure is used to perform weighted sampling of the neighboring nodes of the target node.
[0010] Secondly, this application also provides a graph data processing apparatus. The apparatus includes:
[0011] The target node determination module is used to identify target nodes with neighboring nodes from graph data;
[0012] The sub-block partitioning module is used to partition each target node into sub-blocks according to the identifiers of each neighboring node, and determine the sub-block to which each neighboring node belongs; wherein the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block.
[0013] The weight storage module is used to store the sampling weights of each neighboring node belonging to the same sub-block in the form of a tree array to the sub-block;
[0014] An index building module is used to build an index structure for the target node based on each of the sub-blocks, and the index structure is used to perform weighted sampling of the neighboring nodes of the target node.
[0015] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0016] Identify target nodes with neighboring nodes from graph data;
[0017] For each target node, sub-blocks are divided according to the identifiers of each neighboring node to determine the sub-block to which each neighboring node belongs; wherein, the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block.
[0018] The sampling weights of each neighboring node belonging to the same sub-block are stored in the sub-block in the form of a binary indexed tree;
[0019] Based on each of the sub-blocks, an index structure is constructed for the target node, and the index structure is used to perform weighted sampling of the neighboring nodes of the target node.
[0020] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0021] Identify target nodes with neighboring nodes from graph data;
[0022] For each target node, sub-blocks are divided according to the identifiers of each neighboring node to determine the sub-block to which each neighboring node belongs; wherein, the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block.
[0023] The sampling weights of each neighboring node belonging to the same sub-block are stored in the sub-block in the form of a binary indexed tree;
[0024] Based on each of the sub-blocks, an index structure is constructed for the target node, and the index structure is used to perform weighted sampling of the neighboring nodes of the target node.
[0025] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0026] Identify target nodes with neighboring nodes from graph data;
[0027] The sub-blocks are divided according to the identifiers of each neighboring node to determine the sub-block to which each neighboring node belongs; wherein, the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block.
[0028] The sampling weights of each neighboring node belonging to the same sub-block are stored in the sub-block in the form of a binary indexed tree;
[0029] Based on each of the sub-blocks, an index structure is constructed for the target node, and the index structure is used to perform weighted sampling of the neighboring nodes of the target node.
[0030] The aforementioned graph data processing method, apparatus, computer equipment, storage medium, and computer program product identify target nodes in the graph data that have neighboring nodes. During the process of constructing an index structure for the target nodes, sub-blocks are divided according to the identifiers of the neighboring nodes. The sub-blocks only need to consider the size relationship of the identifiers between the sub-blocks, without needing to consider the order of the identifiers within the sub-blocks. This simplifies the sub-block division process and reduces the data processing resources required for sub-block division. The sampling weights of the neighboring nodes are stored in the sub-blocks in the form of a binary indexed tree. In the sub-blocks, the binary indexed tree is combined with the unordered identifiers to construct an index structure for weighted sampling of the neighboring nodes of the target node. This reduces the memory overhead required for unordered storage of the sampling weights of the neighboring nodes and facilitates the rapid implementation of weighted sampling of neighboring nodes. Attached Figure Description
[0031] Figure 1 This is an application environment diagram of the data processing method in one embodiment;
[0032] Figure 2 This is a flowchart illustrating a data processing method in one embodiment;
[0033] Figure 3 This is a schematic diagram of the topology of the graph data in one embodiment;
[0034] Figure 4 This is a schematic diagram of the sampling weight accumulation interval of a binary indexed tree in one embodiment.
[0035] Figure 5 This is a schematic diagram illustrating the principle of Hall effect partitioning in one embodiment;
[0036] Figure 6 This is a schematic diagram of the process for updating the sampling weights of neighboring nodes in one embodiment;
[0037] Figure 7 This is a flowchart illustrating the process of deleting the sampling weights of neighboring nodes in one embodiment;
[0038] Figure 8 This is a schematic diagram illustrating the process of inserting sampling weights for neighboring nodes in one embodiment;
[0039] Figure 9 This is a schematic diagram illustrating the results of different update processes applied to the sampling weights in one embodiment.
[0040] Figure 10 This is a schematic diagram of the sampling process for neighboring nodes in one embodiment;
[0041] Figure 11 This is a schematic diagram of the identifier compression process of the graph data processing method in one embodiment;
[0042] Figure 12 This is a schematic diagram of the structure of a data processing system in one embodiment;
[0043] Figure 13 This is a schematic diagram of the index tree structure after the sub-blocks are split in one embodiment;
[0044] Figure 14 This is a schematic diagram illustrating the calculation method of a binary indexed tree (BIT) in one embodiment;
[0045] Figure 15 This is a schematic diagram of the principle of FSTable in one embodiment;
[0046] Figure 16 This is a schematic diagram comparing the dynamic graph construction speed of one embodiment with that of the prior art;
[0047] Figure 17 This is a schematic diagram comparing memory usage in one embodiment with that in the prior art;
[0048] Figure 18 This is a schematic diagram comparing the batch update performance of one embodiment with that of the prior art;
[0049] Figure 19 This is a schematic diagram comparing the sampling performance of one embodiment with that of the prior art;
[0050] Figure 20This is a schematic diagram comparing the sampling performance of an embodiment with that of the prior art at different hop counts;
[0051] Figure 21 This is a structural block diagram of the data processing device in one embodiment;
[0052] Figure 22 This is an internal structural diagram of a computer device in one embodiment;
[0053] Figure 23 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0055] The graph data processing method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server. Terminal 102 and server 104 can work together to execute graph data processing methods, or they can be used independently to execute graph data processing methods. Taking the collaborative graph data processing method of terminal 102 and server 104 as an example, terminal 102 sends graph data to server 104. Server 104 identifies the target node with neighboring nodes from the graph data, divides it into sub-blocks according to the identifier of each neighboring node, and determines the sub-block to which each neighboring node belongs. Among them, the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block. Then, server 104 stores the sampling weights of each neighboring node belonging to the same sub-block in the form of a binary indexed tree in the sub-block. Based on each sub-block, an index structure for the target node is constructed. The index structure is used to perform weighted sampling of the neighboring nodes of the target node.
[0056] The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0057] The multiple index structures obtained by executing the graph data processing method provided in the embodiments of this application can be applied to weighted sampling of graph data. More specifically, in the field of artificial intelligence, weighted sampling of graph data and dynamic updating of graph data can also be used to train models such as graph neural networks.
[0058] Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce new intelligent machines that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess perception, reasoning, and decision-making capabilities.
[0059] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0060] The technical solution in this application can be specifically applied to the graph data processing process in machine learning. Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and many other disciplines. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and pre-trained learning. Pre-trained models are the latest development in deep learning, integrating the above techniques.
[0061] In one embodiment, such as Figure 2 As shown, a graph data processing method is provided. Taking the application of this method to a computer device as an example, the computer device can be a terminal, a server, or a system composed of a terminal and a server, etc. The graph data processing method specifically includes the following steps:
[0062] Step 202: Identify the target node with neighboring nodes from the graph data.
[0063] Graph data is a type of data represented in the form of a graph, which includes nodes and edges. Nodes represent entities, and edges represent relationships between different entities. Graph data can capture many-to-many relationships, hierarchical relationships, and attribute relationships between different entities. This makes graph data more intuitive and effective in representing and storing complex relational information. Graph data can be applied to scenarios such as social network analysis, knowledge graphs, and recommendation systems.
[0064] For example, in the field of friend recommendation, nodes in graph data can be user identifiers, node features can be user attribute features, and edges in graph data can indicate whether the users corresponding to two connected nodes are friends. In the field of information recommendation, nodes in graph data can represent promotional information or user identifiers, node features can be content features of promotional information or user attribute features, and edges in graph data can represent whether a user clicked on the promotional information or not. In other embodiments, graph data can also be knowledge graphs. It is understood that in practical application scenarios, the information actually represented by nodes and edges in graph data constructed for different scenarios will also be different. Edges in graph data can also include weight data, which is used to characterize the strength of the connection between the two nodes connected by the edge. In some embodiments, nodes in graph data include source nodes and destination nodes. It is understood that for directed graphs, the edges between different nodes are directional; for the edge, the node in the direction the arrow points is the destination node, and the node in the opposite direction of the arrow is the source node. For undirected graphs, the edges between different nodes are undirectional, so the two nodes connected by the edge can be each other's source and destination nodes.
[0065] In some embodiments, the computer device acquires graph data in the form of a topological structure, i.e., a topological graph. The computer device then converts this topological graph into an edge representation. For example, please refer to... Figure 3 , Figure 3 This is a graph topology where nodes 1-7 represent seven distinct nodes. Lines connecting nodes indicate an edge between the first and last nodes. Each edge is weighted to represent the sampling weight, reflecting the strength of the relationship between the two connected nodes. Figure 3 In the graph data, the topological graph has 7 nodes and 5 edges. Nodes can be accessed via... To represent, edges can be expressed as follows: To express.
[0066] The target node refers to the source node in the topology of the graph data. The destination node, which is connected to the source node through an edge, is the neighbor node of the target node. For example, in Figure 3 In the example, the target nodes with neighboring nodes are node 1 and node 3. The neighboring nodes of node 1 include node 2, node 3 and node 5, and the neighboring nodes of node 3 include node 4 and node 7.
[0067] Step 204: For each target node, divide it into sub-blocks according to the identifiers of each neighboring node, and determine the sub-block to which each neighboring node belongs.
[0068] Each node in the graph data has a corresponding identifier to uniquely represent the node in the graph data, for example... Figure 3In the diagram, 1 to 7 represent the identifiers of nodes 1 to 7. A sub-block is the result of dividing the target node into blocks based on its neighboring nodes. Each sub-block corresponds to at least one set of neighboring nodes of the target node. The identifiers within the same sub-block are not ordered, and the largest identifier in the preceding sub-block is less than or equal to the smallest identifier in the following sub-block.
[0069] For example, if node 1 has 6 neighboring nodes, the first 3 can be grouped into one sub-block based on their identifier numbers, and the last 3 into another. For instance, if the identifier numbers of node 1's 6 neighboring nodes are 3, 5, 7, 6, 2, and 4, and the identifier numbers of the first 3 are 2, 3, and 4, and the last 2 are 5, 6, and 7, then the neighboring nodes with identifier numbers 2, 3, and 4 can be grouped into sub-block 1, and the neighboring nodes with identifier numbers 5, 6, and 7 into sub-block 2. It should be noted that the identifier numbers of the 3 neighboring nodes in sub-block 1 can be arranged in either 2, 3, and 4, or 3, 2, and 4. Similarly, the identifier numbers of the 3 neighboring nodes in sub-block 2 can be arranged in either 5, 6, and 7, or 5, 7, and 6. The specific arrangement can be determined based on the sub-block partitioning method, as long as the maximum identifier number in the preceding sub-block is less than or equal to the minimum identifier number in the following sub-block.
[0070] In some embodiments, the method of dividing sub-blocks according to the identifiers of neighboring nodes can be implemented by a fast selection algorithm, such as by size comparison, including single-path cyclic fast selection, double-path cyclic fast selection, or by Hall effect partitioning.
[0071] Step 206: Store the sampling weights of each neighboring node belonging to the same sub-block in the sub-block in the form of a binary indexed tree.
[0072] The sampling weight is a weighted data representing the strength of the relationship between the target node and its neighboring nodes. The stronger the relationship, the larger the sampling weight value, and the greater the probability of the target node being sampled during random sampling. By adding sampling weights for storage, it is easy to implement weighted sampling of neighboring nodes for the target node.
[0073] A binary indexed tree, also known as a Fenwick tree, is a data structure with a time complexity of log(n) for both querying and modifying elements. It's primarily used for quickly finding the sum of all elements between any two points, making it a very practical data structure. It uses node i to record the array index... This section contains information about all numbers in the given interval, where k is the number of trailing zeros in the binary representation of i. The goal is to perform lookup and update of the array data in O(log n) time.
[0074] like Figure 4 As shown, sub-block 1 includes 6 neighboring nodes, whose sampling weights are represented by W0-W5, where 0-5 represent the position of the weight data in the array, and their corresponding binary representations are 0000, 0001, 0010, 0011, 0100, 0101, respectively. Correspondingly, according to... In the interval shown, the stored value of the position record corresponding to W0 is W0, the stored value of the position record corresponding to W1 is W0+W1, the stored value of the position record corresponding to W2 is W2, the stored value of the position record corresponding to W3 is W0+W1+W2+W3, the stored value of the position record corresponding to W4 is W4, and the stored value of the position record corresponding to W5 is W5+W6.
[0075] Step 208: Based on each of the sub-blocks, construct an index structure for the target node, the index structure being used to perform weighted sampling of the target node's neighboring nodes.
[0076] The index structure is a data structure used to locate the neighboring nodes of the target node. Specifically, the index structure can be an index table or an index tree, among other data structures. Since each sub-block stores the weight data of neighboring nodes in the form of a binary indexed tree, the index tree facilitates weighted sampling of the target node's neighboring nodes.
[0077] The graph data processing method described above identifies target nodes with neighboring nodes in the graph data. During the index structure construction for the target nodes, sub-blocks are divided according to the identifiers of the neighboring nodes. The sub-blocks only need to consider the size relationship of the identifiers between sub-blocks, without needing to consider the order of the identifiers within the sub-blocks. This simplifies the sub-block division process and reduces the data processing resources required for sub-block division. The sampling weights of the neighboring nodes are stored in the sub-blocks in the form of a binary indexed tree. In the sub-blocks, the binary indexed tree is combined with the unordered identifiers to construct an index structure for weighted sampling of the neighboring nodes of the target node. This reduces the memory overhead required for unordered storage of the sampling weights of the neighboring nodes and facilitates the rapid implementation of weighted sampling of neighboring nodes.
[0078] In one embodiment, the step of dividing the data into sub-blocks according to the identifiers of each neighboring node, and determining the sub-block to which each neighboring node belongs, includes:
[0079] Hall partitioning is performed based on the identifiers of each neighboring node of the target node to locate the target identifier that meets the median condition.
[0080] The target identifier is used as the split point to divide the blocks into sub-blocks until the number of identifiers contained in each sub-block meets the sub-block capacity threshold condition.
[0081] Hall partitioning refers to updating the position of a reference element through elements in an interactive sequence, such that after the update, the values of elements preceding the reference element are all less than or equal to the reference element's value, and the values of elements following the reference element are all greater than or equal to the reference element's value. For example... Figure 5 The diagram shown illustrates the principle of Hall effect partitioning, using a random element as the pivot and selecting the leftmost element as the reference point. Figure 5 In this algorithm, pointer j is responsible for finding elements smaller than the pivot point from right to left, and pointer i is responsible for finding elements larger than the pivot point from left to right. Once found, they are swapped until i and j intersect. Finally, the pivot point is swapped with either i or j.
[0082] The median condition refers to the range of middle positions determined based on the number of elements in an array. If, after swapping the pivot element, its position falls within this middle range, then the pivot element's corresponding identifier is the target identifier that satisfies the median condition. For example, in an array with 13 elements, using the leftmost element as the pivot element and performing Hall partitioning, the element to be swapped with is located at the center of the array. After the swap, the element at the center becomes the pivot element. All elements before the center have values less than or equal to the pivot element, and all elements after the center have values greater than or equal to the pivot element. Based on the definition of the median, the pivot element located at the center of the array is the median of the array.
[0083] In one embodiment, for the ID sequence composed of the identifiers of neighboring nodes, a reference point k is first found. This reference point k can be the first value of the ID sequence, the middle value of the first, last, and mIDpoints, or a random selection. Then, the ID sequence is processed using the Hoare Partition Algorithm to ensure that two constraints are met:
[0084] 1) For any j, when j < k, ;
[0085] 2) For any j, when j > k, ,in yes The i-th ID value in the middle.
[0086] If the above two conditions are met, the split can be performed based on the position after the exchange of the reference point k.
[0087] The specific process is as follows: Assuming that the split is strictly based on the median of the ID sequence, a baseline point v is selected in each round, let's say v = 1. Then, the Hall partitioning algorithm is executed to ensure that all IDs to the left of v are smaller than v, and all IDs to the right of v are larger than v. Then, based on the number of identifiers contained in the ID sequence, this process is repeated using a divide-and-conquer method until the pivot point is exactly at the center of the ID sequence, i.e., the median of the ID sequence. At this center, the sequence is split. If the number of identifiers in the split sub-blocks does not meet the sub-block capacity threshold, the sub-block splitting process continues until the number of identifiers in the split sub-blocks all meet the sub-block capacity threshold.
[0088] In this embodiment, by using Hall partitioning, sub-blocks are split by setting a sub-block capacity threshold condition, which ensures that the number of identifiers contained in each sub-block is less than or equal to the sub-block capacity threshold, thus ensuring balanced splitting of sub-blocks. A balanced number of sub-blocks is conducive to the construction of the index structure, with a time complexity of O(N), so that the subsequent sampling process can be executed quickly and the sampling efficiency can be improved.
[0089] In one embodiment, the median condition can be either the median itself or a median interval range. Taking a median interval range as an example, the median condition determination method includes: obtaining relaxation parameters that match the target node; and determining the median interval range corresponding to the relaxation parameters as the median condition that the sub-block partitioning needs to satisfy.
[0090] The relaxation parameter is used to expand the median interval range corresponding to the median condition. For example, when the relaxation parameter is 1, the median interval range includes the center position of the sequence, the position before the center, and the position after the center, a total of three positions. As another example, when the relaxation parameter is 4, the median interval range includes the center position of the sequence, the four positions before the center, and the four positions after the center, a total of nine positions. Different relaxation parameters can be configured for different graph data. In some specific embodiments, different relaxation parameters can also be configured for different types of target nodes for the same graph data. The specific configuration can be based on actual needs.
[0091] In one specific embodiment, an approximate reference position pivot is selected based on the reference point k. And satisfy the following inequalities:
[0092] Where α is a user-defined relaxation parameter. It is assumed that a soft split is performed according to the median, but in practice, iteration can stop as long as the distance between the pivot position and the actual center position is within α, and the split can be directly performed based on... Split.
[0093] Specifically, assuming the ID sequence length is N, then splitting according to the median is actually...
[0094] The fast selection algorithm, in In the algorithm, as long as pivot is Within a certain range, division can begin. Clearly, the larger the α value, the faster the division, but the more uneven the division, and vice versa.
[0095] In this embodiment, sub-block splitting is performed by obtaining relaxation parameters that match the target node, which can effectively improve the splitting efficiency compared to the splitting method based on median splitting.
[0096] In some embodiments, the sub-block includes multiple key-value pairs; each key-value pair corresponds one-to-one with a neighboring node; the key-value pair uses the identifier of the neighboring node as the index key and the storage value corresponding to the sampling weight of the neighboring node as the index value; wherein the storage values contained in the same sub-block constitute a binary indexed tree that sequentially records the sampling weights of each neighboring node.
[0097] In this context, a key-value pair refers to a data storage method that associates two pieces of data using an index key and an index value. The index key allows for quick location of the corresponding index value, enabling rapid data retrieval. In this embodiment, within a sub-block, each neighboring node corresponds to a key-value pair. This key-value pair uses the neighboring node's identifier as the index key and the stored value corresponding to the neighboring node's sampling weight as the index value.
[0098] It should be noted that the sampling weight is not directly used as the index value in the key-value pair. Instead, the storage value corresponding to the sampling weight is used as the index value. The storage value is the cumulative value of the sampling weight of an interval. The range of this interval is determined based on the position of the neighboring node in the sub-block, so that the storage values contained in the same sub-block form a tree array that sequentially records the sampling weight of each neighboring node.
[0099] In this embodiment, the relationship between the identifier and the stored value is established by using key-value pairs, which makes it easy to quickly obtain the position of the neighboring node in the sub-block and the sampling weight of the neighboring node.
[0100] In some embodiments, the index structure may be an index table, an index tree, etc. Taking an index tree as an example, the method of constructing an index structure for the target node based on each of the sub-blocks, i.e., constructing the index tree, includes: generating hierarchical index information based on the key-value pairs stored in each of the sub-blocks; and constructing an index tree for the target node based on each of the sub-blocks and each of the hierarchical index information.
[0101] The index tree is a hierarchical index structure, including leaf nodes and non-leaf nodes. A leaf node is a node in the index tree that has a parent node but no children. A non-leaf node is a node in the index tree that has children. The root node of the index tree is a non-leaf node that has only children and no parent. Hierarchical index information consists of the information of the non-leaf nodes at each level of the index tree. In this embodiment, each sub-block is a leaf node of the index tree. Based on the leaf nodes, hierarchical index information is generated sequentially as the non-leaf nodes at each level. Then, based on the leaf and non-leaf nodes, an index tree for the target node is constructed.
[0102] In this embodiment, by constructing an index tree with a hierarchical structure as the index structure, the sampling process can be indexed based on the hierarchical structure of the index tree, which can effectively simplify the index complexity of the sampling process and improve the sampling efficiency.
[0103] In one embodiment, the hierarchical index information is characterized by key-value pairs contained in each hierarchical index sub-block. Generating hierarchical index information based on the key-value pairs stored in each sub-block includes:
[0104] For each sub-block, the identifier with the smallest value is selected from the key-value pairs stored in the sub-block, and the sum of the sampling weights of each neighboring node assigned to the sub-block is determined; using the identifier and the sum of the weights as key-value pairs, a new level of hierarchical index sub-blocks is generated sequentially until the new level of hierarchical index sub-blocks become the top-level index sub-blocks.
[0105] In this system, the sub-blocks obtained by dividing the sub-blocks using identifiers are called leaf node sub-blocks. Sub-blocks generated based on these leaf node sub-blocks are called hierarchical index sub-blocks. Each new level of hierarchical index sub-block includes information from the current level's hierarchical index sub-blocks. A key-value pair in a new level's hierarchical index sub-block corresponds to a hierarchical index sub-block in the current level. The top-level index sub-block refers to the highest-level hierarchical index sub-block and serves as the root node of the index tree.
[0106] In one embodiment, the number of identifiers in the same sub-block satisfies the sub-block capacity threshold condition. The step of generating new-level hierarchical index sub-blocks sequentially using the identifier and the sum of the weights as key-value pairs, until the new-level hierarchical index sub-blocks are the top-level index sub-blocks, includes: using the identifier and the sum of the weights as key-value pairs, and according to the sub-block capacity threshold condition, sequentially generating new-level hierarchical index sub-blocks until the number of new-level hierarchical index sub-blocks is 1. By setting the sub-block capacity threshold condition, it is possible to ensure a balance in the number of elements contained in each sub-block and each level of generated hierarchical index sub-blocks, thereby improving the overall sampling efficiency through balanced data distribution.
[0107] In one specific embodiment, the target node has 64 neighboring nodes, and the sub-block capacity threshold for each sub-block is 4. Therefore, the 64 neighboring nodes are divided into 16 sub-blocks, meaning the index tree has 16 leaf nodes. For these 16 sub-blocks, a first-level hierarchical index is generated. This first level includes four hierarchical index entries, each containing four key-value pairs, each pointing to a sub-block. The index key of the key-value pair is the smallest identifier in the sub-block, and the index value is the sum of the sampling weights of all neighboring nodes in the sub-block. For these four hierarchical index entries, a new level of hierarchical index is generated. Since this new level of hierarchical index has only one entry, it becomes the top-level index.
[0108] In this embodiment, by selecting the smallest identifier from the sub-blocks and determining the sum of the sampling weights of each neighboring node assigned to the sub-block to construct a new key-value pair, and sequentially generating a new level of hierarchical index sub-blocks, the sum of the sampling weights of all sub-blocks included in the top-level index information can be used to determine the range of values for random sampling parameters according to the sum of weights, thereby improving the comprehensiveness of the sampling range and ensuring the randomness of the sampling results.
[0109] The above describes the process of building the index tree for the target node. After building the index tree, it is also necessary to maintain the information of the target node's neighbor nodes, specifically including updating the sampling weights of neighbor nodes, adding neighbor nodes, and deleting neighbor nodes. The following describes the processing methods for these three cases: updating the sampling weights of neighbor nodes, adding neighbor nodes, and deleting neighbor nodes. Figure 9 The figures shown are schematic diagrams of the processing results in the above three cases.
[0110] The process involves a computer device acquiring a node update request, extracting the source node identifier, destination node identifier, and sampling weights carried in the request, and searching for a target node with the same identifier as the source node. If the target node exists, its index tree is retrieved, and a search is conducted based on the index tree to determine if a neighboring node with the same identifier as the destination node exists. If such a neighboring node exists, it is determined whether the node update request carries a deletion flag. If a deletion flag is present, the node update request is identified as a request to delete the sampling weights of a neighboring node of the target node. If not, the node update request is identified as a request to update the sampling weights of a neighboring node of the target node. If no neighboring node exists, the node update request is identified as a request to insert a neighboring node of the target node.
[0111] It is understood that the first neighbor node, the second neighbor node, and the third neighbor node in the following embodiments are merely to illustrate several ways of maintaining neighbor node information under different circumstances. There is no necessary relationship between the three. The first neighbor node, the second neighbor node, and the third neighbor node can be the same neighbor node or different neighbor nodes.
[0112] In one embodiment, the update of the sampling weights of neighboring nodes specifically includes the following process:
[0113] Step 602: In response to the sampling weight update request for the first neighbor node in the target node, determine the updated sampling weight of the first neighbor node and the first identifier representing the first neighbor node.
[0114] Step 604: Based on the first identifier, find the first sub-block where the first neighbor node is located through the index tree, and find the first key-value pair of the first neighbor node from the first sub-block.
[0115] Step 606: Based on the position of the first stored value of the first key-value pair in the binary indexed tree, determine the associated stored value that is related to the first stored value.
[0116] Step 608: Update the first stored value and the associated stored value according to the updated sampling weight.
[0117] The sampling weight update request includes a first identifier representing the first neighbor node and its updated sampling weight. Since the identifiers among sub-blocks are unordered, but there is a relationship between their sizes, the computer device, using an index tree, can first determine the first sub-block containing the first neighbor node based on its identifier size. Then, based on the first identifier, it finds the first key-value pair of the first neighbor node within the sub-block. The index value of the first key-value pair is the first stored value corresponding to the sampling weight of the first neighbor node. Based on the position of the first stored value in the binary indexed tree, the sampling weight accumulation interval, including the associated stored value at that position, can be determined. Finally, the first stored value and the associated stored value are updated according to the updated sampling weight.
[0118] Further, updating the first stored value and the associated stored value according to the updated sampling weight includes: determining the weight difference between the updated sampling weight and the target weight data based on the target weight data matched by the first stored value; and updating the first stored value and the associated stored value according to the weight difference.
[0119] In one specific embodiment, the first sub-block's binary indexed tree (FSTable) has a total of [number] sub-blocks. Given an element, the index of the key-value pair corresponding to its first neighbor node is i. Now, we need to update the sampling weight of the first neighbor node at index i to w. First, we use FSTable in the worst-case scenario... The time complexity is to calculate the original sampling weight value w. i Then, the original sampling weight value w is subtracted from the new sampling weight value w. i Get the increment value In each iteration of the non-leaf nodes in the index tree, first find the parent node at position i from the previous level of non-leaf nodes, and then add the position of the parent node. Then iterates through the parent node's parent node, and so on. This process continues until the exponent exceeds the number of elements in the FSTable. Clearly, the time complexity of updating the sampling weights of existing neighbor nodes is O(n log n). .
[0120] By using the index tree provided in the above embodiments, in the process of updating the sampling weight of an existing node, only a part of the data in the binary indexed tree is needed to achieve weighted sampling, which greatly reduces the array information dimension caused by the change of one of the sampling weights and can effectively improve the time complexity of sampling weight update processing.
[0121] In one embodiment, regarding the neighbor node deletion process, such as Figure 7 As shown, the specific process includes the following:
[0122] Step 702: In response to the deletion request for the second neighbor node in the target node, determine the second identifier of the second neighbor node.
[0123] Step 704: Based on the second identifier, find the second sub-block where the second neighbor node is located through the index tree, and find the second key-value pair of the second neighbor node from the second sub-block.
[0124] Step 706: Swap the contents of the second key-value pair with the last key-value pair of the second character block.
[0125] Step 708: Update the sampling weight based on the second key-value pair after content swapping, and delete the last key-value pair after content swapping.
[0126] For ease of naming, the sub-block containing the second neighbor node will be referred to as the second sub-block, and the key-value pair of the second neighbor node will be referred to as the second key-value pair.
[0127] The deletion of the second neighbor node requires deleting its second key-value pair. To minimize the impact of other key-value pairs in the second sub-block during deletion, the second key-value pair is swapped with the last key-value pair in the second sub-block before deletion. This ensures the second key-value pair to be deleted is at the end of the sub-block. Deleting the last key-value pair does not affect the positions of other key-value pairs in the sub-block, nor does it affect the accumulated sampling weight range. Simultaneously, the problem of deleting the key-value pair at the position of the second key-value pair is transformed into a problem of updating sampling weights. The processing method for updating sampling weights is the same as described above and will not be repeated.
[0128] In a specific embodiment, deleting key-value pairs in a sub-block hinges on ensuring that the deletion does not affect the positions of other key-value pairs within the sub-block, i.e., it does not change the indices representing their positions within the sub-block. Specifically, this is achieved by swapping the sampling weights of the key-value pair to be deleted with those of the last key-value pair, transforming the deletion operation into a sampling weight update operation. In the actual processing, the sampling weights of the neighboring nodes represented by the last key-value pair are first restored using a binary indexed tree (BIT). Then use Update the weight at position i, and finally delete the last element. Since the deletion operation is ultimately transformed into a sampling weight update operation, the time complexity of the deletion process is O(log n). The entire process does not reconstruct the FSTable, simplifying the complexity of deleting sampling weights.
[0129] In one embodiment, regarding the neighbor node insertion process, such as Figure 8 As shown, the specific steps include:
[0130] Step 802: In response to the insertion request for the third neighbor node in the target node, obtain the target sampling weight of the third neighbor node, and search for the target sub-block that meets the insertion conditions from the index tree corresponding to the target node.
[0131] Step 804: If the target sub-block satisfies the key-value pair addition condition, determine the sampling weight accumulation interval corresponding to the third neighbor node according to the tree array formed by the stored values in the target sub-block.
[0132] Step 806: The sampling weight of the third neighbor node is summed with the sampling weights corresponding to the sampling weight accumulation interval to obtain the storage value corresponding to the sampling weight of the third neighbor node.
[0133] Step 808: The third identifier and the corresponding storage value of the third neighbor node are used as key-value pairs and inserted at the end of the target sub-block.
[0134] The insertion request for the third neighbor node includes the third neighbor node's third identifier and its sampling weight. First, based on the value of the third identifier, the target sub-block that meets the insertion criteria is searched from the leaf nodes of the index tree. The sub-blocks, which are the leaf nodes of the index tree, are arranged according to their identifier values. The method for determining the target sub-block is as follows: if there exists a sub-block whose minimum identifier value is less than or equal to the third identifier value, and the minimum identifier value of the next sub-block is greater than the third identifier value, then that sub-block is a target sub-block that meets the insertion criteria.
[0135] The key-value pair addition condition is that the number of key-value pairs contained in the sub-block is less than the sub-block capacity threshold. If the number of key-value pairs contained in the target sub-block is less than the sub-block capacity threshold, it means that the third neighbor node can be inserted at the end of the target sub-block, thus determining the insertion position of the third neighbor node as the end of the target sub-block.
[0136] For the key-value pair that the third neighbor node needs to be inserted at the end of the target sub-block, its index key is the third identifier. The index value is determined based on the binary indexed tree (BIT) formed by the stored values in the target sub-block. First, the sampling weight accumulation interval corresponding to the third neighbor node needs to be determined based on the number of existing key-value pairs in the target sub-block. The sampling weight accumulation interval is implemented as described in the above embodiment and will not be repeated here. After determining the sampling weight accumulation interval, the sampling weight of the third neighbor node is accumulated with each sampling weight corresponding to the sampling weight accumulation interval to obtain the stored value corresponding to the sampling weight of the third neighbor node, i.e., the index value of the key-value pair. Thus, the key-value pair that the third neighbor node needs to be inserted at the end of the target sub-block is determined. Inserting this key-value pair at the end of the target sub-block completes the insertion process for the third neighbor node.
[0137] In a specific application, given an index tree and a neighbor node to be inserted, with identifier v, return a new index tree after inserting neighbor node v. First, find the target sub-block to insert the neighbor node, and then find a path P using a depth-first traversal algorithm. (where H is the height of the index tree). In the search path, the IDs contained in each non-leaf node are arranged in ascending order. Therefore, during the search, we find the position j of the smallest ID greater than or equal to v, and then continue to search for the j-th child, that is, the j-th non-leaf node of the next level, and so on, until a leaf node is found. This leaf node is the target sub-block.
[0138] At a leaf node, if the neighbor node v is in the leaf node If the key-value pair exists, update the sampling weight according to the updated sampling weight of the neighbor node v; otherwise, perform an append operation on the key-value pair of the neighbor node v, inserting the key-value pair of the neighbor node v into the leaf node. At the end of the page, after updating the leaves, it was found... If the number of elements exceeds the threshold `max`, a split operation is triggered. The process of determining the split node is implemented as described in the above embodiments. Finally, the path P from the leaf node is updated. to the root node The key-value pairs of all nodes are used to complete the insertion process for the neighbor node v.
[0139] Since the append operation only adds key-value pairs at the end and does not change the arrangement of other key-value pairs in the FSTable within the sub-block, it does not cause the binary indexed tree (BIT) to be restructured. Specifically, the current key-value pair's position in the target sub-block is initialized to index i, the current stored value s is w, and k is a set of 1s followed by consecutive zeros. In each iteration, the child of i is found, i.e., index x is calculated to see if it satisfies the following formula; if it does, it is considered a child.
[0140]
[0141] in This is a bitwise AND operation, where k is the number of consecutive trailing zeros in the binary representation of x.
[0142] The calculation of F[i] can be represented by the following formula:
[0143]
[0144] in, The collection of all children.
[0145] If F[x] is a child of F[i], then s is added to F[x]. In each iteration, k is incremented by 1, and the termination condition is... Finally, it appends s to the position F[i], and theoretical proof shows that the time complexity is O(logN).
[0146] Next, we will introduce the process of sampling neighbor nodes based on the index tree, such as... Figure 10 As shown, the neighbor node sampling process includes:
[0147] Step 1002: In response to the sampling request for the neighbor nodes of the target node, obtain the sum of the sampling weights of each neighbor node of the target node.
[0148] Step 1004: Using the sum of the weights as the upper limit of the sampling parameter values, randomly obtain the sampling parameters.
[0149] Step 1006: From the leaf nodes of the index tree matched by the target node, find the sampling sub-block that matches the sampling parameters, and determine the node sampling parameters for the sampling sub-block.
[0150] Step 1008: Based on the key-value pairs contained in the sampling sub-block, sample neighboring nodes that match the node sampling parameters.
[0151] In the index tree, since the index values of the key-value pairs contained in the non-leaf nodes of each level are the sum of the index values of the previous level, the sum of the index values of the key-value pairs contained in the root node of the index tree is the sum of the sampling weights of the target node's neighboring nodes. Therefore, the sum of the sampling weights is obtained from the root node of the index tree. With the sum of the weights as the upper limit of the sampling parameter value, the determined range of sampling parameter values can fully cover all neighboring nodes. By randomly obtaining the sampling parameters from the range of sampling parameter values, random sampling across the entire range can be achieved, improving the balanced distribution of the sampling results.
[0152] The sampling sub-block that matches the sampling parameters is the sub-block corresponding to the leaf node obtained by the last location of the layer-by-layer indexing. During the layer-by-layer indexing process, the sampling parameters are updated once for each indexing until the node sampling parameters for the sampling sub-block are obtained. Finally, according to the node sampling parameters, the neighboring nodes that match the node sampling parameters are sampled from the key-value pairs contained in the sampling sub-block.
[0153] For example, by combining the ITS method for non-leaf nodes and the FTS method for leaf nodes, a neighbor node is sampled from a given target node. First, within the range... Generate a random number R, where This is the sum of the sampling weights of all neighboring nodes of s. Then, based on the random number R, a layer-by-layer search is performed in the index tree Samtree of the target node. For each non-leaf node in Samtree, the ITS method is used to search for key-value pairs that satisfy the index value among the key-value pairs contained in the non-leaf node. Find the smallest key-value pair i, then jump to the i-th child, i.e., the i-th non-leaf node of the next level, and continue searching. Specifically, C[-1] = 0. When the search reaches a leaf node, the FTSmethod is used, and... The input is used to obtain the final index p in the leaf node. The sampled neighbor node is the neighbor node represented by the identifier corresponding to the p-th position in the leaf node.
[0154] In this embodiment, the sampling parameters are indexed and updated layer by layer through the index tree until the node sampling parameters for the sampling sub-block are obtained. According to the node sampling parameters, neighboring nodes that match the node sampling parameters are sampled from the sampling sub-block. The index tree based on the above structure can quickly sample neighboring nodes, effectively reduce the memory overhead required for sampling, and simplify the complexity of the sampling process, thereby improving sampling efficiency.
[0155] In one embodiment, the identifier is represented by a binary number; the method of storing the identifier in the sub-block specifically includes: for each sub-block at each level in the index tree, based on the common prefix of each identifier contained in the sub-block, splitting each identifier into a prefix and a suffix; storing each identifier contained in the same sub-block according to a combination of one prefix and multiple suffixes.
[0156] To further reduce memory consumption, a novel compression algorithm is proposed based on the above-described index tree structure. This algorithm primarily compresses the ID list composed of identifiers in both leaf and non-leaf nodes of the index tree. Existing graph compression algorithms, such as ZipG, require preprocessing of data, leading to significant performance overhead as the entire graph is decompressed and recompressed every time the graph data is updated. To address this issue, this application proposes a compression algorithm that uses a binary prefix compression method to compress the ID list, which is referred to as the "Dynamic Prefix Compression Algorithm," storing the ID list as a string.
[0157] For example, leaf nodes or non-leaf nodes in an index tree contain Given a set of IDs, assuming that the first z bits of all IDs are the same, then CP-IDs are represented as follows:
[0158] z│prefix│suf(v0)│suf(v1)│…│suf(v L )│
[0159] Where prefix is the ID prefix of the first z bytes, and suf(v) is the 8-z byte suffix of point v.
[0160] like Figure 11 As shown, an example is given (in this example, the first 7 bytes are the prefix). In some embodiments, the number of prefix bytes can be selected from {0, 4, 6, 7} to further improve the efficiency of suffix lookup.
[0161] This application also provides an application scenario for GNN training dynamic graph storage. This application scenario utilizes the aforementioned graph data processing method and can be used in recommendation scenarios such as video accounts, live streaming, and e-commerce, as well as in risk control fields such as black market detection. Specifically, the graph data processing method is applied in this scenario as follows:
[0162] The system architecture for GNN training dynamic graph storage is as follows: Figure 12 As shown, the system architecture consists of two layers. The first layer, from top to bottom, is the TF operator layer, which mainly supports the operation interface of the upper-layer model. The second layer is the dynamic graph storage layer, which is used to store the dynamic graph topology. Among them, Samtree is an index tree, which is mainly used to maintain graph topology information; Fenwick Tree is used to reduce the update time complexity of the probability table from O(N) to O(logN); key-value pairs (kv) are used to store the mapping between nodes and attributes.
[0163] Part 1: Introduction to Samtree (Index Tree)
[0164] Before describing the Samtree structure, we first describe graph topology storage and show how to use Samtree to describe the storage graph topology structure.
[0165] like Figure 13 As shown in the figure There are 7 points and 5 edges, that is and Specifically, using an index tree, Samtree. Let G(V,E,W) describe all the neighbors of point u. It has two source points, 1 and 3. Each source point maintains its corresponding Samtree to keep track of all its neighbors. In this example, the maximum tree node capacity of the Samtree is set to 2; that is, a split will occur when the tree node size exceeds 2.
[0166] by For example, there are only 2 neighbors in total, which is within the maximum capacity, so a single leaf node is sufficient for storage. The leaf node contains two parts: an ID list storing neighbor node IDs and a sampling probability table (FSTable). The FSTable is used for sampling and is maintained using a Fenwick Tree to improve the efficiency of probability table updates.
[0167] To enable weighted sampling, an additional sampling probability table needs to be maintained. A Fenwick Tree is used to address the issue of dynamically updating the weights of the leaf nodes in a Samtree. A Fenwick Tree is a data structure that can efficiently compute prefix sums. For example... Figure 14 As shown, each F[i] in the Fenwick Tree is the sum of weights in the interval j~i. When using the Fenwick Tree in the scenario described above, when a certain neighbor's weight... When a change occurs, only a small portion of the values in F will be modified; that is, the F values within the range i will be updated, while others will not. Theoretically, the update time complexity is proven to be O(logN).
[0168] An FSTable was constructed using Fenwick Tree, leaf node IDs were deordered, and insertion / deletion was converted into swap and append operations to solve the index inconsistency problem. Then, a new sampling algorithm, the FTS method, was proposed to solve the sampling problem.
[0169] The principle of FSTable is: given a source node s, and its neighboring nodes forming an index tree (Samtree) , There is a leaf node ( The weight array (of elements) is FSTable is a table of length 10 ... An array, where the i-th element is:
[0170]
[0171]
[0172] in, Let j be the sampling weight of the neighboring node j of the target node s. k is the number of consecutive trailing zeros in the binary representation of i.
[0173] For example, when i=10, since the binary representation of 10 ends with only one consecutive 0, From the above formula, we can see that FSTable has a property that each element is the sum of the interval F[g(i) + 1] to F[i].
[0174] In addition to the properties mentioned above, FSTable also has the property of a tree. The i-th value F[i] in FSTable is equal to the prefix sum of all its children. Therefore, if all the children of i are known, F[i] can be calculated directly. If there is an index x, and x satisfies the following formula, then F[x] is a child of F[i]:
[0175]
[0176] in, Let F[i] be a bitwise AND operation, where k is the number of consecutive trailing zeros in the binary representation of x. Therefore, the calculation of F[i] can be transformed into the following formula:
[0177]
[0178] Where, Xi Let F[i] be the set of all children of F[i].
[0179] The following example illustrates the calculation of FSTable in detail. Figure 15 As shown, there are 3 points in a leaf node of a samtree, and the weight array is A={0.3,0.4,0.1}.
[0180] When i = 0, since g(0) = -1, g(0) + 1 = 0, Therefore, the actual value of F[0] is w0, which means F[0] = 0.3.
[0181] Similarly, when i = 1, we can obtain the following: In addition, F[0] is a child of F[1] because index 0 satisfies the above child determination conditions.
[0182] When i = 2, since g(2) = 1, g(2) + 1 = 2. And F[2] has no children.
[0183] Again For example, there are a total of 3 neighbors. Since the maximum capacity limit is exceeded, a split is required. In this example, after the split, there are 2 leaf nodes, and 1 non-leaf node is formed based on the 2 leaf nodes. The information contained in the leaf nodes is the same as above.
[0184] The information contained within non-leaf nodes is explained in detail below:
[0185] 1) An ID list is used, where each ID corresponds to a leaf node. The i-th ID value is its i-th child, which is the smallest ID among the corresponding i-th leaf nodes. For example, if there are 8 leaf nodes, and non-leaf node 1 corresponds to leaf nodes 1, 2, 3, and 4, then leaf nodes 1, 2, 3, and 4 are called children of non-leaf node 1. Similarly, non-leaf node 2 corresponds to leaf nodes 5, 6, 7, and 8, then leaf nodes 5, 6, 7, and 8 are called children of non-leaf node 2. The second ID in the ID list of non-leaf node 1 is the smallest ID among leaf nodes 2, and the second ID in the ID list of non-leaf node 2 is the smallest ID among leaf nodes 6.
[0186] 2) CSTable, the first value of CSTable is the sum of the sample weights of the first child (0.5), and the second value is the sum of the sample weights of the first child and the second child (0.5+0.2=0.7).
[0187] Part Two: Tree Node Splitting
[0188] The leaf node ID list of a Samtree is unordered. Maintaining this unordered leaf node order is primarily for the purpose of using a Fenwick Tree to manage the dynamic probability table, and secondly, to improve the performance of insertion, update, and deletion operations (append replaces insert). If the IDs were ordered, splitting a tree node would be as simple as splitting at the midpoint. However, since the Samtree ID list is unordered, the α-split algorithm is proposed to handle the splitting problem. The α-split algorithm is based on the fast selection algorithm, and it reduces the average time complexity of splitting to O(N) without requiring strictly ordered leaf nodes.
[0189] In the leaf node First, a pivot point k is found. The pivot point k can be the first value in the ID sequence, a value selected from the first, last, and mIDpoints, or a random value. Then, the ID sequence is processed using the Hall partitioning algorithm to ensure that two constraints are satisfied:
[0190] 1) For any j, when j < k,
[0191] 2) For any j, when j > k,
[0192] in, yes The i-th ID value is given. At this point, splitting can be performed based on position k.
[0193] The specific process is as follows: Assuming that the split is strictly based on the median of the ID sequence, a pivot v is selected in each round. Then, perform Hall partitioning to ensure that all IDs to the left of v are smaller than v, and all IDs to the right of v are larger than v. Then, use a divide-and-conquer approach to continuously perform this process until the pivot is exactly the median.
[0194] Using the fast selection algorithm to solve the splitting problem, although the average time complexity is O(N), it still incurs significant copying overhead. To improve splitting efficiency, an approximate pivot is chosen. And satisfy the following inequalities:
[0195]
[0196] Where α is a user-defined relaxation parameter.
[0197] Assuming a soft split based on the median, but in reality, the iteration can stop as long as the distance between the pivot and the actual median is within α, and the split can be directly based on... The above algorithm is called the α-split algorithm.
[0198] Specifically, assuming leaf nodes If the ID sequence length is N, then splitting according to the median is actually... The fast selection algorithm, in the α-split algorithm, requires that the pivot is... Within a certain range, division can begin. Clearly, the larger the α value, the faster the division, but the more uneven the division, and vice versa.
[0199] Part 3: Graph Topology Update
[0200] The key to graph topology updates is actually the Samtree update. In a specific application, given an index tree and a neighbor node to be inserted, with the identifier v, return a new index tree after inserting neighbor node v. In the pseudocode, the target sub-block for inserting the neighbor node is first searched, which finds a path P using a depth-first traversal algorithm. (where H is the height of the index tree). In the search path, the IDs contained in each non-leaf node are arranged in ascending order. Therefore, during the search, we find the position j of the smallest ID greater than or equal to v, and then continue to search for the j-th child, that is, the j-th non-leaf node of the next level, and so on, until a leaf node is found. This leaf node is the target sub-block.
[0201] The pseudocode describes how to operate on a leaf node if its neighbor node v is a leaf node. If the key-value pair exists, update the sampling weight according to the updated sampling weight of the neighbor node v; otherwise, perform an append operation on the key-value pair of the neighbor node v, inserting the key-value pair of the neighbor node v into the leaf node. At the end of the page, after updating the leaves, it was found... If the number of elements exceeds the threshold `max`, a split operation is triggered. The process of determining the split node is implemented as described in the above embodiments. Finally, the path P from the leaf node is updated. to the root node The key-value pairs of all nodes are used to complete the insertion process for the neighbor node v. Figure 15 To insert a new edge Examples.
[0202] The following explains how to update the weight of an existing point in the FSTable (In-place Update), insert a weight into a specified position (Insertion), and delete the weight at position i (Deletion).
[0203] For example, the first sub-block's binary indexed tree FSTableF, the binary indexed tree has a total of Given an element, the index of the key-value pair corresponding to its first neighbor node is i. Now, we need to update the sampling weight of the first neighbor node at index i to w. First, we use FSTable in the worst-case scenario... Time complexity for calculating the original sampling weight values Then, the original sampling weight value is subtracted from the new sampling weight value w. Get the increment value In each iteration of the non-leaf nodes in the index tree, first find the parent node at position i from the previous level of non-leaf nodes, and then add the position of the parent node. Then iterates through the parent node's parent node, and so on. This process continues until the exponent exceeds the number of elements in the FSTable. Clearly, the time complexity of updating the sampling weights of existing neighbor nodes is O(n log n). .
[0204] Insert a weight at a specified position (Insertion). Given an index tree and a neighbor node to be inserted, with identifier v, return a new index tree after inserting neighbor node v. First, it searches for the target sub-block to which the neighbor node should be inserted, finding a path P using a depth-first traversal algorithm. (where H is the height of the index tree). In the search path, the IDs contained in each non-leaf node are arranged in ascending order. Therefore, during the search, we find the position j of the smallest ID greater than or equal to v, and then continue to search for the j-th child, that is, the j-th non-leaf node of the next level, and so on, until a leaf node is found. This leaf node is the target sub-block.
[0205] If the neighbor node v is in a leaf node If the key-value pair exists, update the sampling weight according to the updated sampling weight of the neighbor node v; otherwise, perform an append operation on the key-value pair of the neighbor node v, inserting the key-value pair of the neighbor node v into the leaf node. At the end of the page, after updating the leaves, it was found... If the number of elements exceeds the threshold `max`, a split operation is triggered. The process of determining the split node is implemented as described in the above embodiments. Finally, the path P from the leaf node is updated. to the root node The append operation performs key-value pair insertion for all nodes, including the neighbor node v. Since the append operation only adds key-value pairs at the end and does not change the arrangement of other key-value pairs in the FSTable within the sub-blocks, it does not cause the binary indexed tree (BIT) to be restructured.
[0206] The key to deleting the weight at position i is ensuring that the deletion does not affect the positions of other IDs and weights within the leaf nodes, i.e., their indices. This is achieved by swapping the sampled weight to be deleted with the last sampled weight, transforming the deletion operation into a sampled weight update operation. During processing, the sampled weights of the neighboring nodes represented by the last key-value pair are first restored using the FSTable. Then use Update the weight at position i, and finally delete the last element. Since the deletion operation is ultimately transformed into a sampling weight update operation, the time complexity of the deletion process is O(log n). The entire process does not reconstruct the FSTable, simplifying the complexity of deleting sampling weights.
[0207] Part Four: Weighted Sampling
[0208] The FTS (Fenwick Tree-based Sampling Method) method states that the element at position i in the FSTable is equal to the sum of the weights of a certain interval. This is based on the "tree-like property" of the FSTable: for an integer k > 0, the element at position i in the FSTable is equal to the sum of the weights of that interval. The elements are equal to the interval. All weights sum.
[0209] like Figure 14 As shown, there are 6 elements, and when k = 1, When k = 2, .
[0210] Based on the above properties, a range-narrow sampling method is used. The principle of range-narrowing is to continuously shorten the sampling range in the FSTable until it is finally shortened to a certain index, which is the sampled index. First, a range is generated... Let R be a random number between S and S, where S is the sum of all weights. S can be obtained by calling the getAllSum() function with a time complexity of O(n log n). Then find the one that satisfies... The smallest Then, using a range-narrow strategy in [0, The actual index of the range search is the sampled index.
[0211] Specifically, in each iteration, the midpoint index mID is first found. If mID is greater than... If it exceeds the bounds, then the upper bound `right` is set to `mID`, and the next loop iteration continues; otherwise, it compares... And R, the index is finally found through the above range-narrow.
[0212] Combining the ITS method for non-leaf nodes and the FTS method for leaf nodes, a neighbor node is sampled from a given target node. First, within the range... Generate a random number R, where This is the sum of the sampling weights of all neighboring nodes of s. Then, based on the random number R, a layer-by-layer search is performed in the index tree Samtree of the target node. For each non-leaf node in Samtree, the ITS method is used to search for key-value pairs that satisfy the index value among the key-value pairs contained in the non-leaf node. Find the smallest key-value pair i, then jump to the i-th child, i.e., the i-th non-leaf node of the next level, and continue searching. Specifically, C[-1] = 0. When the search reaches a leaf node, the FTS method is used, and... The input is used to obtain the final index p in the leaf node. The sampled neighbor node is the neighbor node represented by the identifier corresponding to the p-th position in the leaf node.
[0213] To further reduce memory consumption, a novel compression algorithm is proposed based on the above-described index tree structure. This algorithm primarily compresses the ID list composed of identifiers in both leaf and non-leaf nodes of the index tree. Existing graph compression algorithms, such as ZipG, require preprocessing of data, leading to significant performance overhead as the entire graph is decompressed and recompressed every time the graph data is updated. To address this issue, this application proposes a compression algorithm that uses a binary prefix compression method to compress the ID list, which is referred to as the "Dynamic Prefix Compression Algorithm," storing the ID list as a string.
[0214] For example, leaf nodes or non-leaf nodes in an index tree contain Given a set of IDs, assuming that the first z bits of all IDs are the same, then CP-IDs are represented as follows:
[0215] z│prefix│suf(v0)│suf(v1)│…│suf(v L )│
[0216] Here, prefix is the ID prefix of the first z bytes, and suf(v) is the 8-z byte suffix of point v. In some embodiments, the number of prefix bytes can be selected from {0, 4, 6, 7} to further improve the efficiency of suffix lookup.
[0217] To further verify the effectiveness of this approach, it was compared with two other existing graph data processing methods (PlatoGL and Aligraph) using three different datasets. The datasets are described in detail below. Figure 16 As shown.
[0218] The graph data processing method in this application, compared to other existing graph data processing methods, such as... Figure 17 The comparison of dynamic graph construction speed shown indicates that dynamic graph loading speed is up to 6.3 times faster than existing graph data processing methods. For example... Figure 18 The memory usage comparison shown indicates that memory usage can be reduced by up to 79.8% during operation.
[0219] In addition, update experiments were conducted under different batches, such as... Figure 19 As shown in the figure, the batch update performance comparison shows that, among the given batch sizes, the update speedup can reach up to 5.4 times compared to existing graph data processing methods.
[0220] like Figure 20 The figures show the experimental results for three datasets. (a) to (c) represent the performance of sampling neighbors (1 hop). The experimental figures show that the performance is up to 3.2 times better than the existing graph data processing method. (d) to (f) represent the performance of sampling subgraphs (2 hops). The experimental figures show that the performance is up to 13.7 times better than the existing graph data processing method.
[0221] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0222] Based on the same inventive concept, this application also provides a graph data processing apparatus for implementing the graph data processing method described above. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations in one or more graph data processing apparatus embodiments provided below can be found in the limitations of the graph data processing method described above, and will not be repeated here.
[0223] In one embodiment, such as Figure 21As shown, a graph data processing device is provided, including: a target node determination module 2102, a sub-block partitioning module 2104, a weight storage module 2106, and an index construction module 2108, wherein:
[0224] The target node determination module 2102 is used to identify target nodes with neighboring nodes from graph data.
[0225] The sub-block partitioning module 2104 is used to partition sub-blocks according to the identifiers of each neighboring node, and to determine the sub-block to which each neighboring node belongs; wherein the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block.
[0226] The weight storage module 2106 is used to store the sampling weights of each neighbor node belonging to the same sub-block in the form of a tree array to the sub-block.
[0227] The index building module 2108 is used to build an index structure for the target node based on each of the sub-blocks, and the index structure is used to perform weighted sampling on the neighboring nodes of the target node.
[0228] In one embodiment, the sub-block partitioning module 2104 is used to perform Hall partitioning based on the identifiers of each neighboring node of the target node, locate the target identifier that meets the median condition, and use the target identifier as the split point to partition the sub-blocks until the number of identifiers contained in each sub-block meets the sub-block capacity threshold condition.
[0229] In one embodiment, the sub-block partitioning module 2104 is used to obtain relaxation parameters that match the target node; and to determine the median interval range corresponding to the relaxation parameters as the median condition that the sub-block partitioning needs to satisfy.
[0230] In one embodiment, the sub-block includes multiple key-value pairs; each key-value pair corresponds one-to-one with a neighboring node; the key-value pair uses the identifier of the neighboring node as the index key and the storage value corresponding to the sampling weight of the neighboring node as the index value; wherein the storage values contained in the same sub-block constitute a binary indexed tree that sequentially records the sampling weights of each neighboring node.
[0231] In one embodiment, the index structure includes an index tree; an index building module 2108 is used to generate hierarchical index information based on the key-value pairs stored in each of the sub-blocks; and to build an index tree for the target node based on each of the sub-blocks and each of the hierarchical index information.
[0232] In one embodiment, the hierarchical index information is characterized by key-value pairs contained in each hierarchical index sub-block; the index construction module 2108 is used to, for each sub-block, select the identifier with the smallest value from the key-value pairs stored in the sub-block, and determine the sum of the sampling weights of each neighboring node assigned to the sub-block; using the identifier and the sum of the weights as key-value pairs, a new level of hierarchical index sub-block is generated sequentially until the new level of hierarchical index sub-block is the top-level index sub-block.
[0233] In one embodiment, the number of identifiers in the same sub-block satisfies the sub-block capacity threshold condition; the index construction module 2108 is used to generate a new level of hierarchical index sub-blocks sequentially according to the sub-block capacity threshold condition, using the identifier and the sum of the weights as key-value pairs, until the number of new level hierarchical index sub-blocks is 1.
[0234] In one embodiment, the graph data processing apparatus further includes an update processing module, configured to, in response to a sampling weight update request for a first neighbor node in the target node, determine an update sampling weight for the first neighbor node and a first identifier representing the first neighbor node; based on the first identifier, search for the first sub-block where the first neighbor node is located through an index tree, and search for the first key-value pair of the first neighbor node in the first sub-block; based on the position of the first stored value of the first key-value pair in the binary indexed tree, determine an associated stored value that is associated with the first stored value; and update the first stored value and the associated stored value according to the update sampling weight.
[0235] In one embodiment, the update processing module is configured to determine the weight difference between the updated sampling weight and the target weight data based on the target weight data matched by the first stored value; and update the first stored value and the associated stored value according to the weight difference.
[0236] In one embodiment, the graph data processing apparatus further includes an update processing module, configured to, in response to a deletion request for a second neighbor node in the target node, determine a second identifier of the second neighbor node; based on the second identifier, search for the second sub-block where the second neighbor node is located through an index tree, and search for the second key-value pair of the second neighbor node in the second sub-block; swap the content of the second key-value pair with the last key-value pair of the second sub-block; update the sampling weight based on the second key-value pair after the content swap, and delete the last key-value pair after the content swap.
[0237] In one embodiment, the graph data processing apparatus further includes an update processing module, configured to, in response to an insertion request for a third neighbor node in a target node, obtain the target sampling weight of the third neighbor node and search for a target sub-block that meets the insertion conditions from the index tree corresponding to the target node; if the target sub-block meets the key-value pair addition conditions, determine the sampling weight accumulation interval corresponding to the third neighbor node according to the tree array formed by the stored values in the target sub-block; accumulate the sampling weight of the third neighbor node with the sampling weights corresponding to the sampling weight accumulation interval to obtain the storage value corresponding to the sampling weight of the third neighbor node; and insert the third identifier and the storage value corresponding to the third neighbor node as a key-value pair at the end of the target sub-block.
[0238] In one embodiment, the graph data processing apparatus further includes a sampling processing module, configured to, in response to a sampling request for neighboring nodes of a target node, obtain the sum of the sampling weights of each neighboring node of the target node; randomly obtain sampling parameters with the sum of weights as the upper limit of the sampling parameter values; search for sampling sub-blocks that match the sampling parameters from the leaf nodes of the index tree matched by the target node, and determine the node sampling parameters for the sampling sub-blocks; and sample neighboring nodes that match the node sampling parameters based on the key-value pairs contained in the sampling sub-blocks.
[0239] In one embodiment, the identifier is represented by a binary number; the graph data processing device further includes an identifier compression module, which is used to split each identifier into a prefix and a suffix for each sub-block at each level in the index tree based on the same prefix of each identifier contained in the sub-block; and to store each identifier contained in the same sub-block in a combination of one prefix and multiple suffixes.
[0240] The aforementioned graph data processing device identifies target nodes in the graph data that have neighboring nodes. During the process of constructing an index structure for the target nodes, it divides the data into sub-blocks according to the identifiers of the neighboring nodes. The sub-blocks only need to consider the size relationship of the identifiers between the sub-blocks, without needing to consider the order of the identifiers within the sub-blocks. This simplifies the sub-block division process and reduces the data processing resources required for sub-block division. The sampling weights of the neighboring nodes are stored in the sub-blocks in the form of a binary indexed tree. In the sub-blocks, the binary indexed tree is combined with the unordered identifiers to construct an index structure for weighted sampling of the neighboring nodes of the target node. This reduces the memory overhead required for unordered storage of the sampling weights of the neighboring nodes and facilitates the rapid implementation of weighted sampling of the neighboring nodes.
[0241] Each module in the aforementioned data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0242] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 22 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores graph data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a graph data processing method.
[0243] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 23As shown, the computer device includes a processor, memory, input / output interface, communication interface, display unit, and input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interface. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interface is used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a graph data processing method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0244] Those skilled in the art will understand that Figure 22 or Figure 23 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0245] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0246] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0247] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0248] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0249] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0250] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0251] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A graph data processing method, characterized in that, The method includes: Identify target nodes with neighboring nodes from graph data; For each target node, sub-blocks are divided according to the identifiers of each neighboring node to determine the sub-block to which each neighboring node belongs; wherein, the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block. The sampling weights of each neighboring node belonging to the same sub-block are stored in the sub-block in the form of a binary indexed tree; Based on each of the sub-blocks, an index structure is constructed for the target node, and the index structure is used to perform weighted sampling of the neighboring nodes of the target node.
2. The method according to claim 1, characterized in that, The step of dividing the land into sub-blocks according to the identifiers of each neighboring node, and determining the sub-block to which each neighboring node belongs, includes: Hall partitioning is performed based on the identifiers of each neighboring node of the target node to locate the target identifier that meets the median condition. The target identifier is used as the split point to divide the blocks into sub-blocks until the number of identifiers contained in each sub-block meets the sub-block capacity threshold condition.
3. The method according to claim 2, characterized in that, The method further includes: Obtain the relaxation parameters that match the target node; The median interval range corresponding to the relaxation parameter is determined as the median condition that the sub-block division needs to satisfy.
4. The method according to claim 1, characterized in that, The sub-block includes multiple key-value pairs; each key-value pair corresponds one-to-one with a neighboring node. The key-value pair uses the identifier of the neighboring node as the index key and the storage value corresponding to the sampling weight of the neighboring node as the index value. The stored values contained in the same sub-block constitute the sampling weights of each neighboring node recorded sequentially.
5. The method according to claim 4, characterized in that, The index structure includes an index tree; The step of constructing an index structure for the target node based on each of the sub-blocks includes: Based on the key-value pairs stored in each of the sub-blocks, hierarchical index information is generated; Based on each of the sub-blocks and each of the hierarchical index information, an index tree is constructed for the target node.
6. The method according to claim 5, characterized in that, The hierarchical index information is represented by the key-value pairs contained in each hierarchical index sub-block; The generation of hierarchical index information based on the key-value pairs stored in each sub-block includes: For each sub-block, the identifier with the smallest value is selected from the key-value pairs stored in the sub-block, and the sum of the sampling weights of each neighboring node assigned to the sub-block is determined; Using the identifier and the sum of the weights as key-value pairs, new level hierarchical index sub-blocks are generated sequentially until the new level hierarchical index sub-block becomes the top-level index sub-block.
7. The method according to claim 6, characterized in that, The number of identifiers in the same sub-block meets the sub-block capacity threshold condition; The step of generating new-level hierarchical index sub-blocks sequentially using the identifier and the sum of the weights as key-value pairs, until the new-level hierarchical index sub-blocks become the top-level index sub-blocks, includes: Using the identifier and the sum of the weights as key-value pairs, new level hierarchical index sub-blocks are generated sequentially according to the sub-block capacity threshold condition, until the number of new level hierarchical index sub-blocks is 1.
8. The method according to any one of claims 5 to 7, characterized in that, The method further includes: In response to the sampling weight update request for the first neighbor node in the target node, the updated sampling weight of the first neighbor node and the first identifier representing the first neighbor node are determined. Based on the first identifier, the first sub-block where the first neighbor node is located is found through the index tree, and the first key-value pair of the first neighbor node is found from the first sub-block; Based on the position of the first stored value of the first key-value pair in the binary indexed tree, an associated stored value that is related to the first stored value is determined; Update the first stored value and the associated stored value according to the updated sampling weight.
9. The method according to claim 8, characterized in that, The step of updating the first stored value and the associated stored value according to the updated sampling weight includes: Based on the target weight data matched by the first stored value, determine the weight difference between the updated sampling weight and the target weight data; Update the first stored value and the associated stored value according to the weight difference.
10. The method according to any one of claims 5 to 7, characterized in that, The method further includes: In response to a deletion request for the second neighbor node in the target node, determine the second identifier of the second neighbor node; Based on the second identifier, the second sub-block where the second neighbor node is located is found through the index tree, and the second key-value pair of the second neighbor node is found from the second sub-block; Swap the contents of the second key-value pair with the last key-value pair of the second character block; The sampling weight is updated based on the second key-value pair after the content swap, and the last key-value pair after the content swap is deleted.
11. The method according to any one of claims 5 to 7, characterized in that, The method further includes: In response to an insertion request for a third neighbor node in the target node, the target sampling weight of the third neighbor node is obtained, and a target sub-block that meets the insertion conditions is searched from the index tree corresponding to the target node; If the target sub-block satisfies the key-value pair addition condition, the sampling weight accumulation interval corresponding to the third neighbor node is determined according to the tree array formed by the stored values in the target sub-block. The sampling weight of the third neighbor node is summed with the sampling weights corresponding to the sampling weight accumulation interval to obtain the storage value corresponding to the sampling weight of the third neighbor node. The third identifier and the corresponding storage value of the third neighbor node are used as key-value pairs and inserted at the end of the target sub-block.
12. The method according to any one of claims 5 to 7, characterized in that, The method further includes: In response to a sampling request for neighboring nodes of a target node, the sum of the sampling weights of each neighboring node of the target node is obtained; The sampling parameters are randomly obtained with the sum of the weights as the upper limit of the sampling parameter values. From the leaf nodes of the index tree matched by the target node, find the sampling sub-block that matches the sampling parameters, and determine the node sampling parameters for the sampling sub-block; Based on the key-value pairs contained in the sampling sub-block, sample neighboring nodes that match the node sampling parameters.
13. The method according to any one of claims 4 to 7, characterized in that, The identifier is represented using a binary number; the method further includes: For each sub-block at each level in the index tree, based on the common prefix of each identifier contained in the sub-block, each identifier is split into a prefix and a suffix; The identifiers contained in the same sub-block are stored according to a combination of one of the prefixes and multiple of the suffixes.
14. A graph data processing apparatus, characterized in that, The device includes: The target node determination module is used to identify target nodes with neighboring nodes from graph data; The sub-block partitioning module is used to partition each target node into sub-blocks according to the identifiers of each neighboring node, and determine the sub-block to which each neighboring node belongs; wherein the identifiers in the same sub-block are arranged in no order, and the largest identifier in the previous sub-block is less than or equal to the smallest identifier in the next sub-block. The weight storage module is used to store the sampling weights of each neighboring node belonging to the same sub-block in the form of a tree array to the sub-block; An index building module is used to build an index structure for the target node based on each of the sub-blocks, and the index structure is used to perform weighted sampling of the neighboring nodes of the target node.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 13.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 13.