Data processing method and device, equipment, medium and program product
By using the message propagation method of graph neural network in high-end complex industrial fields, the information transmission path between any two nodes in the data is determined, and the massive relational data management problem is solved, and efficient query and management of data relationships is achieved.
Patent Information
- Application Number
- CN202411809328.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-05-16
AI Technical Summary
In the high-end complex industrial field, massive relational data is difficult to effectively manage, resulting in the inability to meet users' query needs for relationships between data.
Through a message propagation method based on graph neural network, all information delivery paths between any two nodes in the to be processed data are determined, thereby indicating the relationship between nodes.
It realizes effective management of large-scale relational data, meets users' query needs for relationships between data, and provides conditions that facilitate data use and user decision-making.
Smart Images

Figure CN120011597A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device, equipment, medium and program product. Background Art
[0002] In high-end complex industrial fields such as large civil aircraft design and manufacturing, large ship design and manufacturing, and large machinery and equipment manufacturing, design and manufacturing activities generate massive amounts of data. These data not only cover information on design parameters, material selection, manufacturing processes, etc., but also have complex relationships between data, such as assembly relationships between parts, dependencies in the design iteration process, and timing relationships in the manufacturing process. However, there is currently no effective way to manage these relational data, which makes it impossible to better meet users' query needs for relationships between data. Summary of the invention
[0003] The present application provides a data processing method, apparatus, device, medium and program product. Through the scheme of the present application, all information transmission paths between any two nodes in the to-be-processed data of any graph structure can be determined. Since the information transmission path between two nodes can represent the relationship between the two nodes, the present application can better meet the user's query needs for the relationship between data.
[0004] The present application provides a data processing method, comprising the following steps: Acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent relationships between different entity objects; Determine the starting nodes corresponding to all the edges in the data to be processed; Through the message propagation method based on graph neural network, the starting node is used as the starting point of message propagation, and the storage information inside each node is propagated to other nodes to obtain processed data. The storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
[0005] According to a data processing method provided by the present application, the message propagation method based on graph neural network takes the starting node as the message propagation starting point, and propagates the storage information inside each node to other nodes, including: Determine an end node connected to the start node; The stored information in the starting node is propagated to the end node through the message propagation method based on the graph neural network; The end node is re-determined as the start node, and the above steps are repeated until a preset stop condition is met.
[0006] According to a data processing method provided by the present application, the message propagation method based on graph neural network is used to propagate the stored information inside the starting node to the end node, including: By using the message propagation method based on graph neural network, the storage information inside the starting node is sent from the starting node to the ending node; Aggregating the storage information inside the starting node and the information transmission path between the starting node and the end node to obtain aggregated information; The aggregated information is used as storage information inside the end point node.
[0007] According to a data processing method provided by the present application, the message propagation method based on graph neural network takes the starting node as the message propagation starting point, and propagates the storage information inside each node to other nodes, including: If it is determined that the starting node is not a loop node, the message propagation method based on the graph neural network is used to propagate the storage information inside each node to other nodes with the starting node as the message propagation starting point, and the loop node is a node on any loop in the data to be processed; If it is determined that the starting node is a loop node, when the starting node is used as the starting point of message propagation for the first time, the stored information inside each node is propagated to other nodes through the message propagation method based on graph neural network.
[0008] According to a data processing method provided by the present application, the determining that the starting node is not a loop node includes: If the preset loop node set does not include the starting node, determining that the starting node is not a loop node; Among them, the preset loop node set is obtained through the following steps: through the message propagation method based on graph neural network, taking the starting node as the starting point of message propagation, propagating the test message from the starting node to other nodes, and adding the nodes that send the test message more than the preset number of times during the propagation process to the preset loop node set.
[0009] According to a data processing method provided by the present application, the message propagation method based on graph neural network takes the starting node as the message propagation starting point, and propagates the storage information inside each node to other nodes, including: Divide the starting nodes into first-type nodes and second-type nodes, wherein the first-type nodes have connection relationships with at least a preset number of other nodes, and the second-type nodes are nodes other than the first-type nodes among all nodes of the data to be processed; For the first-type nodes, propagating the storage information inside each of the first-type nodes to other nodes according to a first preset strategy; For the second-type nodes, the storage information inside each of the second-type nodes is propagated to other nodes according to a second preset strategy, where the second preset strategy is different from the first preset strategy.
[0010] According to a data processing method provided by the present application, after obtaining the processed data, the method further includes: Receiving a path query request sent by a terminal device, wherein the path query request includes identity information of a target start node and a target end node; According to the identity information of the target starting node and the target ending node, in the processed data, all information transmission paths between the target starting node and the target ending node are determined; In response to the path query request, all the information transmission paths are sent to the terminal device.
[0011] The present application also provides a data processing device, comprising the following modules: An acquisition module is used to acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent relationships between different entity objects; A first determination module is used to determine the starting nodes corresponding to all the edges in the data to be processed; The processing module is used to propagate the storage information inside each node to other nodes through a message propagation method based on a graph neural network, taking the starting node as the starting point of message propagation, to obtain processed data, wherein the storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
[0012] The present application also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above-described data processing methods when executing the computer program.
[0013] The present application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements any of the data processing methods described above.
[0014] The present application also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the data processing methods described above.
[0015] The present application provides a data processing method, device, equipment, medium and program product. The method of the present application is implemented by first obtaining the data to be processed. The data structure of the data to be processed is a directed graph structure. The data to be processed includes multiple nodes. Each node is connected to other nodes through edges. The node represents the entity object in the data to be processed, and the edge represents the relationship between different entity objects; then, the starting node corresponding to each of the edges in the data to be processed is determined; finally, through a message propagation method based on a graph neural network, the starting node is used as the starting point of the message propagation, and the storage information inside each node is propagated to other nodes to obtain the processed data. The storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes. In the present application, all information transmission paths between any two nodes in the data to be processed of any graph structure can be determined, which can significantly optimize the management strategy of large-scale relational data. Since the information transmission path between two nodes can represent the relationship between the two nodes, the present application can meet the user's query needs for the relationship between data, thereby providing great convenience for the subsequent use of data and user decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present application or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 is a flow chart of a data processing method shown in one embodiment of the present application; Figure 2 is a schematic diagram of a structure of data to be processed shown in an embodiment of the present application; Figure 3 This is a schematic diagram of a message propagation process shown in an embodiment of the present application; Figure 4 is a schematic diagram of another structure of data to be processed shown in an embodiment of the present application; Figure 5is a schematic diagram of a process for determining a loop node according to an embodiment of the present application; Figure 6 This is a data query flow diagram shown in an embodiment of the present application; Figure 7 is a structural schematic diagram of a data processing device shown in an embodiment of the present application; Figure 8 It is a schematic diagram of the physical structure of an electronic device shown in an embodiment of the present application. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0019] The data processing method of the present application is executed by a data processing device provided by the present application, or any type of electronic device.
[0020] Figure 1 1 is a flow chart of a data processing method according to an embodiment of the present application. Figure 1 The data processing method of the present application specifically includes the following steps: Step 101, obtaining the data to be processed. The data structure of the data to be processed is a directed graph structure. The data to be processed includes multiple nodes. Each node is connected to other nodes through an edge. The node represents an entity object in the data to be processed, and the edge represents the relationship between different entity objects.
[0021] In this embodiment, the data to be processed can be any type of data, such as manufacturing data of mechanical equipment in the industrial field, medical data in the medical field, financial data in the financial field, etc. It can be set according to actual needs, and this embodiment does not make specific restrictions on this.
[0022] The data structure of the data to be processed is a directed graph structure, for example Figure 2 As shown. Figure 2In the figure, each node (node 1-node 8) represents an entity object in the data to be processed. For example, when the data to be processed is manufacturing data of mechanical equipment, the entity object can represent the data generated in various activities (design activities, simulation activities, testing activities, production activities, exhibition activities, etc.) in the manufacturing process of mechanical equipment. There is a certain relationship between different entity objects. For example, simulation activities usually use data in design activities, so there is a certain relationship between simulation activities and design activities. This relationship can be reflected by edges. For example, a design activity can be connected to a simulation activity through an edge, and the direction is from the design activity to the simulation activity, indicating that the data in the design activity can be used in the simulation activity. Figure 2 It is a schematic diagram of the structure of data to be processed shown in an embodiment of the present application.
[0023] exist Figure 2 In the example, there are 8 nodes and 8 edges. The 8 nodes include node 1 to node 8, and the 8 edges include (1, 2), (2, 5), (5, 6), (1, 3), (3, 4), (4, 6), (3, 7), and (7, 8). In this embodiment, each edge can be represented by (a, b), where a represents the starting node of the edge, b represents the end node of the edge, and the direction is from a to b.
[0024] Step 102: Determine the starting nodes corresponding to all the edges in the data to be processed.
[0025] By executing step 102, the starting node of each edge in the data to be processed can be determined, and these starting nodes are recorded in the starting node set.
[0026] For example, Figure 2 The determined starting node set is {1, 2, 5, 1, 3, 4, 3, 7}.
[0027] Step 103: Through the message propagation method based on graph neural network, with the starting node as the starting point of message propagation, the storage information inside each node is propagated to other nodes to obtain processed data. The storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
[0028] Among them, the Graph Neural Network (GNN) is a neural network that can process graph structured data and updates the state or representation of nodes by simulating the propagation of information between nodes. The propagation and aggregation of messages are the core operations in graph neural networks. In the message propagation method based on graph neural networks, each node generates messages to be passed to neighboring nodes based on its own characteristics and the characteristics of the edges connected to it (if any). This process may use the current characteristics of the node itself, the current characteristics of the neighboring nodes, and the characteristics of the edges between the node and the neighbors. The generated messages are passed to the neighboring nodes through the edges, which simulates the flow of information on the graph. Secondly, each node receives messages from its neighboring nodes and aggregates these messages. The aggregation function can be sum, mean, maximum, etc., depending on the design of the graph neural network.
[0029] In this embodiment, the storage information inside each node includes not only the information transmission path between the current node and other nodes, but also the identity information (ID) of the current node, and other information (which will be described in detail later).
[0030] For example, in Figure 2 In the example, taking node 1 as the starting node, the information stored in node 1 includes the identity information of node 1 and the information transmission path between node 1 and other nodes (empty at this time). After node 1 sends its own storage information to node 3, node 3 aggregates the storage information of node 1 and the information transmission path between node 1 and node 3, and then stores the aggregated information (including the identity information of node 1, the identity information of node 3, and the information transmission path between node 1 and node 3) as the storage information of node 3. After node 3 sends its own storage information to node 4, node 4 aggregates the storage information of node 3 and the information transmission path between node 3 and node 4, and then stores the aggregated information (including the identity information of node 1, the identity information of node 3, the identity information of node 4, the information transmission path between node 1 and node 3, the information transmission path between node 3 and node 4, and the information transmission path between node 1 and node 3 and node 4) as the storage information of node 4.
[0031] According to the above message propagation logic, starting from each starting node, the storage information inside each node is propagated to other nodes until the preset stop condition is met. After the propagation is completed, the data to be processed becomes the processed data.
[0032] In this embodiment, if the number of starting nodes in the starting node set determined in step 102 is large, a grouping and batching processing method can be used to propagate the message. Specifically, first, all the starting nodes in the starting node set are divided into multiple groups. For example, all the starting nodes in the starting node set can be divided into multiple groups in a manner that each preset number (the specific value is set according to actual needs) of starting nodes is a group (all the starting nodes remaining less than the preset number are directly regarded as a group). Then, message propagation is performed for each group in turn until all groups have completed message propagation. When message propagation is performed for a group, all the starting nodes in the group perform message propagation at the same time.
[0033] For example, all the starting nodes in the starting node set can be divided into 5 groups, including group 1 to group 5, each of which includes 3 starting nodes. Then, the message is propagated to the 3 starting nodes in group 1 first, and then to the 3 starting nodes in group 2, until all the starting nodes in the 5 groups have completed the message propagation.
[0034] This embodiment supports group and batch processing for message propagation, which can effectively deal with the situation where the hardware conditions of the execution entity are difficult to meet large-scale data calculations.
[0035] The data processing method of the present application is essentially a method for calculating the full path of graph data. Through the processed data, all information transmission paths between any two nodes can be determined, that is, the present application can determine all information transmission paths between any two nodes in any graph data.
[0036] The method for implementing the present application first obtains the data to be processed, the data structure of the data to be processed is a directed graph structure, the data to be processed includes multiple nodes, each node is connected to other nodes through edges, the node represents the entity object in the data to be processed, and the edge represents the relationship between different entity objects; then, determine the starting node corresponding to each of the edges in the data to be processed; finally, through the message propagation method based on the graph neural network, take the starting node as the starting point of the message propagation, propagate the storage information inside each node to other nodes, and obtain the processed data, the storage information includes the information transmission path between the node to which it belongs and other nodes, and the processed data includes the information transmission path between any two different nodes. In the present application, all information transmission paths between any two nodes in the data to be processed of any graph structure can be determined, which can significantly optimize the management strategy of large-scale relational data. Since the information transmission path between two nodes can represent the relationship between the two nodes, the present application can meet the user's query needs for the relationship between data, thereby providing great convenience for the subsequent use of data and user decision-making.
[0037] In combination with the above embodiments, in one implementation, a message propagation method based on a graph neural network is used to propagate the stored information in each node to other nodes with the starting node as the starting point of the message propagation, which may include: Determine the end node connected to the start node; The stored information in the starting node is propagated to the end node through the message propagation method based on graph neural network; The end node is re-determined as the start node, and the above steps are repeated until the preset stop condition is met.
[0038] In one embodiment, propagating the stored information in the starting node to the end node through a message propagation method based on a graph neural network may include: The stored information in the starting node is sent from the starting node to the ending node through the message propagation method based on graph neural network; Aggregate the storage information inside the starting node and the information transmission path between the starting node and the end node to obtain aggregated information; The aggregated information is stored in the terminal node as storage information inside the terminal node.
[0039] In this embodiment, each node in the data to be processed can use a serial number to represent an ID. After obtaining the starting node set, firstly, each starting node in the starting node set is deduplicated, and then arranged in order of ID from small to large.
[0040] Next, for each starting node in the set of starting nodes after deduplication, the activation state of the starting node is modified to true, and then all the terminal nodes connected to the starting node are determined. Then, based on the message propagation mechanism of the graph neural network, a propagation message is sent to each terminal node, and the propagation message includes the storage information inside the starting node. After receiving the propagation message, each terminal node aggregates the storage information inside the starting node, its own identity information, and the information transmission path between the starting node and itself in the propagation message, and then uses the aggregated information as its own storage information. The above process is the first propagation. Then the second propagation is carried out, the activation state of the starting node in the above process is modified to false, and the activation state of all terminal nodes is modified to true, and then these terminal nodes with an activation state of true are used as new propagation starting points, and propagation messages are sent to the terminal nodes connected to them. The propagation principle is the same as the principle of the first propagation mentioned above. Repeat the above process for multiple propagations until the preset stop condition is met, and the processed data can be obtained.
[0041] In this application, only nodes whose activation state is true can send propagation messages.
[0042] In this embodiment, the propagation message also includes the activation status of the node.
[0043] Figure 3 It is a schematic diagram of a message propagation process shown in an embodiment of the present application. Figure 3 Only the process of the first message propagation and the second message propagation is illustrated.
[0044] For example, combined with Figure 2 and Figure 3 , after deduplication and arranging the starting nodes in ascending order of ID, the set is {1, 2, 3, 4, 5, 7}. Then, for starting node 1, first change the activation state of the starting node to true, and then determine the end nodes connected to starting node 1, including node 2 and node 3. Then, starting node 1 sends propagation messages to node 2 and node 3 respectively. After receiving the propagation message, node 2 aggregates the storage information of starting node 1, the information transmission path from starting node 1 to node 2, the identity information of node 2, and the activation state of node 2 in the propagation message, and uses the aggregated information (identity information of starting node 1, identity information of node 2, information transmission path from node 1 to node 2, activation state of starting node 1, activation state of node 2) as its own storage information, where the storage information of starting node 1 includes the identity information of starting node 1, the activation state of starting node 1, and the information transmission path between starting node 1 and other nodes (empty). After receiving the propagation message, node 3 aggregates the storage information of the starting node 1, the information transmission path from the starting node 1 to the node 3, the identity information of node 3, and the activation state of node 3 in the propagation message, and uses the aggregated information (the identity information of the starting node 1, the identity information of node 3, the information transmission path from node 1 to node 3, the activation state of the starting node 1, and the activation state of node 3) as its own storage information, wherein the storage information of the starting node 1 includes the identity information of the starting node 1, the activation state of the starting node 1, and the information transmission path (empty) between the starting node 1 and other nodes. The above process is the first message propagation. Next, the second message propagation is performed, the activation state of the starting node 1 is modified to false, and the activation states of node 2 and node 3 are both modified to true, and then the end node corresponding to node 2 (i.e. node 5) and the end node corresponding to node 3 (i.e. node 4 and node 7) are determined. Next, node 2 sends a propagation message to node 5, and node 3 sends a propagation message to node 4 and node 7 respectively. The specific propagation principle is the same as that of the first message propagation. Reference Figure 3 Repeat the above process for multiple propagations until the preset stop condition is met, and the processed data can be obtained.
[0045] In this embodiment, the preset stop condition may be that the number of times the message is propagated is greater than a set number (eg, 50 times), or that the message propagation is completed, which may be specifically set according to actual needs.
[0046] Through this embodiment, the storage information inside each node in the data to be processed can be propagated to other nodes, providing technical support for ultimately obtaining the full path data between any nodes in the data to be processed.
[0047] In combination with the above embodiments, in one implementation, a message propagation method based on a graph neural network is used to propagate the stored information in each node to other nodes with the starting node as the starting point of the message propagation, which may include: If it is determined that the starting node is not a loop node, the message propagation method based on graph neural network is used to propagate the stored information in each node to other nodes, with the starting node as the starting point of message propagation. The loop node is a node on any loop in the data to be processed. If the starting node is determined to be a loop node, when the starting node is used as the starting point of message propagation for the first time, the stored information inside each node is propagated to other nodes through the message propagation method based on graph neural network.
[0048] In this embodiment, the loop in the data to be processed refers to a directed loop. If a path starts from a certain node and eventually returns to the node, then the path is a loop. Figure 3 In the example, node 3-node 4-node 6-node 7-node 3 is a loop. Figure 4 It is a schematic diagram of another structure of data to be processed shown in an embodiment of the present application.
[0049] In this embodiment, all nodes located on the ring are ring nodes.
[0050] In actual business data graphs, loops are difficult to avoid. For example, in the industrial field, loops mean repeating certain activities, which leads to stagnation of business processes and waste of resources. Therefore, loops are invalid in most cases, and loops will generate a large amount of duplicate and redundant data, affecting the efficiency of full path calculation. Therefore, in this application, the influence of loops needs to be eliminated when calculating the full path.
[0051] During the message propagation process, if a starting node is not a loop node, there will only be one opportunity to send a propagation message, that is, its activation state will only be modified to true once. Therefore, if a starting node is determined to be a loop node, only when the starting node is used as the starting point for message propagation for the first time (the activation state is modified to true for the first time), the message propagation method based on the graph neural network will be used to propagate its internal storage information to other nodes, and the subsequent message propagation will no longer repeat the loop point. If a starting node is not a loop node, the message propagation method based on the graph neural network can be used directly to propagate the storage information inside each node to other nodes, taking the starting node as the starting point for message propagation. The specific propagation process can be referred to the previous text.
[0052] Through this embodiment, the influence of loops in the message propagation process can be eliminated, and the efficiency of full path calculation can be significantly improved.
[0053] In combination with the above embodiment, in one implementation manner, determining that the starting node is not a loop node may include: If the preset loop node set does not include the start node, determining that the start node is not a loop node; Among them, the preset loop node set is obtained through the following steps: through the message propagation method based on graph neural network, taking the starting node as the starting point of message propagation, propagating the test message from the starting node to other nodes, and adding the nodes that send test messages more than the preset number of times during the propagation process to the preset loop node set.
[0054] In this embodiment, before calculating the full path data, it is necessary to first determine the loop node set so that when actually calculating the full path data, it is possible to quickly determine whether a certain starting node is a loop node. Specifically, if a certain starting node is not in the loop node set, it means that the starting node is not a loop node. On the contrary, if a certain starting node is in the loop node set, it means that the starting node is a loop node.
[0055] Figure 5 This is a schematic diagram of a process for determining a loop node according to an embodiment of the present application. Figure 4 and Figure 5 , the process of obtaining the loop node set is described, which specifically includes the following steps: Step 1: Get the starting node set consisting of all starting nodes, that is, {1, 2, 3, 4, 5, 6, 7}; Step 2: Mark the depth of all starting nodes. Specifically, the depth value of each starting node is assigned to 0. At this time, the information stored on each starting node is: its identity information, activation status, and depth value (initialized to 0); Step 3: According to all known edges, from the starting node to the end node of the edge, a message is propagated. The content of the message is 1, and the message is used as a test message. After sending the test message, the activation state of all nodes that send the test message becomes false, that is, not activated. Only nodes in the activated state are eligible to send test messages; Step 4: Update all nodes that receive the test message. Specifically, update the information on all nodes that receive the test message to: their own identity information, activation status (true, indicating activated), and depth value +1; Step 5: After the first round of message propagation, nodes 2, 3, 4, 5, 6, 7, and 8 receive the test message (content 1), and the depth values of these nodes become 1. After the second round of propagation, nodes 3, 4, 5, 6, 7, and 8 receive the propagation message, and the depth values of these nodes become 2. After the third round of propagation, nodes 3, 4, 6, 7, and 8 receive the propagation message, and the depth values of these nodes become 3. After the fourth round of propagation, nodes 3, 4, 6, 7, and 8 receive the propagation message, and the depth values of these nodes become 4. Repeat the above steps until the depth value of the node is greater than the set threshold N (determined according to the scale of the data and the path depth value required by the user, different data scales have different thresholds). Taking the threshold N as 50 as an example, after 50 rounds of propagation, all nodes with a depth value of 50 and an activated state are screened out (at this time, only nodes 3, 4, 6, 7, and 8 have a depth value of 50 and are in an activated state), forming a set S1 (at this time S1={3, 4, 6, 7, 8}).
[0056] Step 6: Next Build Figure 4 The transposed graph of the graph shown in is as follows: Figure 4 The directions of all the edges in the graph shown are reversed, and the nodes remain unchanged. Then, the message is propagated again according to the implementation process of the above steps 1 to 5. Taking the threshold N as 50 as an example, after 50 rounds of propagation, all nodes with a depth value of 50 and an activated state are screened out (at this time, the depth value of nodes 1, 2, 3, 4, 5, 6, 7 is 50 and is in an activated state), forming a set S2 (at this time S2={1, 2, 3, 4, 5, 6, 7}). Finally, the intersection of set S1 and set S2 is taken, and the final intersection {3, 4, 6, 7} is determined as the preset loop node set.
[0057] The method for obtaining the preset loop node set provided in the above steps 1 to 6 is an approach taken based on the estimation of the scale and style of the data to be processed (such as business data) in engineering practice. In actual implementation, if there is no accurate grasp of the graph structure of the entire data to be processed, the loop node screening method based on the Kosaraju (i.e., Kosaraju-Sharir, an algorithm for finding strongly connected components of directed graphs. In a directed graph, if there is a path from a vertex s to t, and there is also a path from t to s, that is, s and t are reachable to each other, then s and t are said to be strongly connected. The component composed of vertices that are strongly connected to each other is called a strongly connected component) algorithm provided below can be adopted to obtain an accurate preset loop node set. The loop node screening method based on the Kosaraju algorithm specifically includes the following steps: Step 1: Right Figure 4 Perform a depth-first search (DFS) on the nodes in the graph shown. Create a stack S to store the completion time of each node. Create a Boolean array visited[] to mark whether each node has been visited, and the initial visit status is all set to false. Then traverse each node in the graph. If a node has not been visited (that is, visited[node] == false), start a depth-first search with that node as the starting point. During the depth-first search process, when all adjacent nodes of a node have been visited, the node is pushed into the stack S, indicating that the node has completed the depth-first search.
[0058] Step 2: Create a transposed graph. Figure 4 The directions of all edges in the graph shown are reversed, and the nodes remain unchanged, forming a transposed graph relative to the original graph. Reinitialize the Boolean array to false because all nodes need to be revisited. Traverse the nodes in the stack S: pop the node from the top of the stack and check whether the node has been visited. If the node has not been visited, start a new depth-first search with the node as the starting point, this time on the transposed graph. Construct a set of all the nodes visited by this depth-first search. If the set contains more than one node, filter out this set.
[0059] Step 3: Output all sets containing more than one node, and form a new set S from all these sets. The new set S is the preset loop node set.
[0060] In this embodiment, after determining whether each node is a loop node, a description of whether it is a loop node can also be added to the propagation message. In other words, the content of the propagation message sent between nodes includes: the node's own identity information, the information transmission path between other nodes, the activation state, and whether it is a loop node.
[0061] Of course, in actual implementation, the loop node set may also be determined by other means, and this embodiment does not specifically limit this.
[0062] Through this embodiment, the loop node set can be quickly determined to ensure smooth implementation of subsequent full path calculation.
[0063] In combination with the above embodiments, in one implementation, a message propagation method based on a graph neural network is used to propagate the stored information in each node to other nodes with the starting node as the starting point of the message propagation, which may include: The starting nodes are divided into first-class nodes and second-class nodes, wherein the first-class nodes have connection relationships with at least a preset number of other nodes, and the second-class nodes are nodes other than the first-class nodes among all nodes of the data to be processed; For the first type of nodes, the storage information inside each first type of node is propagated to other nodes according to the first preset strategy; For the second type of nodes, the storage information inside each second type of node is propagated to other nodes according to a second preset strategy, where the second preset strategy is different from the first preset strategy.
[0064] In this embodiment, the first type of node is a super node among all nodes, and the super node has established a connection relationship with at least a preset number (which can be set according to actual needs, usually a larger value, such as 1000) of other nodes. The second type of node is other nodes among all nodes except the super node.
[0065] In this embodiment, when super nodes and ordinary nodes are calculated in the same batch, a serious data skew problem will occur, that is, most of the computing resources are occupied by a few super nodes, while most ordinary nodes can only use very few computing resources. Therefore, in this application, the first preset strategy is used to process super nodes, and the second preset strategy is used to process ordinary nodes.
[0066] Among them, the first preset strategy is to process according to smaller data batches, that is, to divide all super nodes into more combinations, each combination contains fewer super nodes. The second preset strategy is to process according to larger data batches, that is, to divide all ordinary nodes into smaller combinations, each combination contains more ordinary nodes. For example, assuming that all nodes include 10 super nodes and 10 ordinary nodes, then the 10 super nodes can be divided into 10 groups, each containing 1 super node, and the 10 ordinary nodes can be divided into 2 groups, each with 5 ordinary nodes. The first preset strategy and the second preset strategy can be set according to actual needs.
[0067] After the first type of nodes are divided into multiple groups of nodes according to the first preset strategy, and the second type of nodes are divided into multiple groups of nodes according to the second preset strategy, each group of nodes is distributed to a corresponding computing module for message propagation in a random normal or random uniform distribution manner.
[0068] The specific process of the computing module performing message propagation on each group of nodes can be referred to the foregoing description, and will not be elaborated in detail in this embodiment.
[0069] In this embodiment, different strategies are adopted to process super nodes and ordinary nodes respectively, which can avoid the problem that most computing resources are occupied by a small number of super nodes.
[0070] In combination with the above embodiments, in one implementation manner, after obtaining the processed data, the method of the present application further includes: Receiving a path query request sent by a terminal device, the path query request including identity information of a target start node and a target end node; According to the identity information of the target start node and the target end node, all information transmission paths between the target start node and the target end node are determined in the processed data; In response to the path query request, all information delivery paths are sent to the terminal device.
[0071] In this embodiment, after the processed data (the calculated full path data) is obtained, the data can be stored in a distributed file system (Hadoop Distributed File System, HDFS). HDFS is one of the core components of Hadoop and is a highly fault-tolerant distributed file system that can provide high-throughput data access and automatically perform fault-tolerant processing when hardware failures occur.
[0072] Since HDFS is difficult to use as a business library for query systems, this application imports the full path data stored in the distributed file system into the Hive data warehouse. Next, create an index in Elasticsearch (ES, a distributed, scalable, real-time search and data analysis engine built on Apache Lucene that provides fast search capabilities and supports large-scale data analysis) and define field mappings for it to ensure that the data in Hive can be correctly mapped to the corresponding fields in the Elasticsearch index. Finally, start a scheduled task that will be responsible for batch synchronization or mapping of data in Hive to the index that has been created in Elasticsearch.
[0073] Elasticsearch creates an index for each word in the processed data by scanning it, and indicates the number and position of each word in the processed data. When a user queries, the index system searches based on the pre-created index and feeds back the search results to the user. The inverted index is the core of the entire Elasticsearch, which consists of two parts: the word dictionary and the inverted table. The word dictionary records the words of all documents and the association between words and the inverted table; the inverted table records the document set corresponding to the word, which is composed of the inverted index.
[0074] In this application, a high-performance computing cluster is specifically used to implement full path data calculation. The high-performance computing cluster includes components such as Hadoop, Spark, ElasticSearch, Hive, HDFS, etc., with more than 500 cluster cores, more than 2T running memory, and more than 200TB hard disk memory. The functional configuration of the high-performance computing cluster can also be set according to actual needs, and this embodiment does not impose specific restrictions on this. The data to be processed involved in this application is stored in a high-performance distributed graph database, such as Neo4j, Galaxy Base, Nebula Graph, etc. When calculating the full path data, a message propagation method based on a graph neural network is used and a computing service is developed based on spark GraphX. The processed data is stored in HDFS, and full path data storage is provided through Hive and ElasticSearch. A data query system is established to open a path data query service between any two nodes.
[0075] Figure 6 This is a data query flow diagram shown in an embodiment of the present application. Figure 6 When querying data, the user first enters the query instruction in the query system front end, including the ID of the starting node and the end node. Then, the query system front end sends the query instruction to the query system back end. Then, the query system back end retrieves all the path data marked with the ID from Elasticsearch based on the ID of the starting node and the end node. Then, the query system back end obtains the business data corresponding to each node on the path from the graph database that stores the corresponding graph data based on the path data. Finally, the query system returns the obtained business data to the query system front end to display it to the user so that the user can make selections and further operations.
[0076] In this embodiment, the execution subject (for example Figure 6 The query system backend in the can receive and process terminal devices (such as Figure 6After obtaining the path query request, the identity information of the target start node and the target end node in the path query request is first determined, and then all information transmission paths between the target start node and the target end node are searched in the processed data through Elasticsearch based on the identity information. Finally, all the searched information transmission paths are sent to the terminal device for viewing by the user on the terminal device side, or the business data corresponding to each node on all the searched information transmission paths are extracted from the graph database, and the business data is returned to the terminal device.
[0077] In summary, the technical solution of this application has at least the following technical effects: (1) The solution of this application has good feasibility and wide applicability. The full path calculation method based on the message propagation mechanism of graph neural network can be applied to computing scenarios of various data scales and does not require the computing cluster hardware to have a high configuration.
[0078] (2) This application provides a method for eliminating loop nodes during the calculation of full path data, which can avoid loop calculations and problems such as a surge in data volume and a decrease in data quality due to excessive repetition.
[0079] (3) This application provides a data skew processing strategy in cluster parallel computing, optimizes the efficiency of full-path computing, reduces the time used for parallel computing, and reduces the load on the computing cluster.
[0080] The data processing device provided by the present application is described below. The data processing device described below and the data processing method described above can be referenced to each other.
[0081] Figure 7 is a structural diagram of a data processing device shown in an embodiment of the present application. Figure 7 , the data processing device 700 of the present application includes: The acquisition module 701 is used to acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent the relationships between different entity objects; A first determination module 702 is used to determine the starting nodes corresponding to all the edges in the data to be processed; Processing module 703 is used to propagate the storage information inside each node to other nodes through a message propagation method based on a graph neural network, taking the starting node as the starting point of message propagation, to obtain processed data, wherein the storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
[0082] According to a data processing device 700 provided by the present application, the processing module 703 includes: A first determination submodule, used to determine an end node connected to the start node; A first processing submodule is used to propagate the stored information inside the starting node to the end node through the message propagation method based on the graph neural network; The second determination submodule is used to re-determine the end node as the start node and repeat the above steps until a preset stop condition is met.
[0083] According to a data processing device 700 provided by the present application, the first processing submodule includes: A sending submodule, used to send the storage information inside the starting node from the starting node to the end node through the message propagation method based on the graph neural network; An aggregation submodule, used to aggregate the storage information inside the starting node and the information transmission path between the starting node and the end node to obtain aggregated information; The third determining submodule is used to use the aggregated information as storage information inside the endpoint node.
[0084] According to a data processing device 700 provided by the present application, the processing module 703 includes: The second processing submodule is used for, if it is determined that the starting node is not a loop node, using the message propagation method based on the graph neural network, taking the starting node as the message propagation starting point, propagating the storage information inside each of the nodes to other nodes, and the loop node is a node on any loop in the data to be processed; The third processing submodule is used to propagate the storage information inside each node to other nodes through the message propagation method based on the graph neural network when the starting node is determined to be a loop node and the starting node is used as the starting point of message propagation for the first time.
[0085] According to a data processing device 700 provided by the present application, the second processing submodule includes: A fourth determination submodule, configured to determine that the starting node is not a loop node if the starting node is not included in the preset loop node set; Among them, the preset loop node set is obtained through the following steps: through the message propagation method based on graph neural network, taking the starting node as the starting point of message propagation, propagating the test message from the starting node to other nodes, and adding the nodes that send the test message more than the preset number of times during the propagation process to the preset loop node set.
[0086] According to a data processing device 700 provided by the present application, the processing module 703 includes: A division submodule, used for dividing the starting node into a first type of node and a second type of node, wherein the first type of node has a connection relationship with at least a preset number of other nodes, and the second type of node is the node other than the first type of node among all nodes of the data to be processed; a fourth processing submodule, configured to propagate the storage information inside each of the first-type nodes to other nodes according to a first preset strategy for the first-type nodes; The fifth processing submodule is used to propagate the storage information inside each second-type node to other nodes according to a second preset strategy for the second-type nodes, where the second preset strategy is different from the first preset strategy.
[0087] According to a data processing device 700 provided in the present application, the device 700 further includes: A receiving module, configured to receive a path query request sent by a terminal device, wherein the path query request includes identity information of a target start node and a target end node; A second determination module is used to determine all information transmission paths between the target starting node and the target end node in the processed data according to the identity information of each of the target starting node and the target end node; The sending module is used to respond to the path query request and send all the information transmission paths to the terminal device.
[0088] Figure 8 is a schematic diagram of the physical structure of an electronic device shown in an embodiment of the present application, such as Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the data processing method, which includes: Acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent relationships between different entity objects; Determine the starting nodes corresponding to all the edges in the data to be processed; Through the message propagation method based on graph neural network, the starting node is used as the starting point of message propagation, and the storage information inside each node is propagated to other nodes to obtain processed data. The storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
[0089] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0090] On the other hand, the present application further provides a computer program product, the computer program product comprising a computer program, the computer program may be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the data processing method provided by the above methods, the method comprising: Acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent relationships between different entity objects; Determine the starting nodes corresponding to all the edges in the data to be processed; Through the message propagation method based on graph neural network, the starting node is used as the starting point of message propagation, and the storage information inside each node is propagated to other nodes to obtain processed data. The storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
[0091] In another aspect, the present application further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the data processing method provided by the above methods is implemented, and the method includes: Acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent relationships between different entity objects; Determine the starting nodes corresponding to all the edges in the data to be processed; Through the message propagation method based on graph neural network, the starting node is used as the starting point of message propagation, and the storage information inside each node is propagated to other nodes to obtain processed data. The storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
[0092] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A data processing method, characterized in that: include: Acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent relationships between different entity objects; Determine the starting nodes corresponding to all the edges in the data to be processed; Through the message propagation method based on graph neural network, the starting node is used as the starting point of message propagation, and the storage information inside each node is propagated to other nodes to obtain processed data. The storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
2. The data processing method according to claim 1, characterized in that: The message propagation method based on graph neural network, taking the starting node as the message propagation starting point, propagates the storage information inside each node to other nodes, including: Determine an end node connected to the start node; The stored information in the starting node is propagated to the end node through the message propagation method based on the graph neural network; The end node is re-determined as the start node, and the above steps are repeated until a preset stop condition is met.
3. The data processing method according to claim 2, characterized in that: The method of propagating the stored information in the starting node to the end node through the message propagation method based on the graph neural network includes: By using the message propagation method based on graph neural network, the storage information inside the starting node is sent from the starting node to the ending node; Aggregating the storage information inside the starting node and the information transmission path between the starting node and the end node to obtain aggregated information; The aggregated information is used as storage information inside the end point node.
4. The method according to claim 1, characterized in that: The message propagation method based on graph neural network, taking the starting node as the message propagation starting point, propagates the storage information inside each node to other nodes, including: If it is determined that the starting node is not a loop node, the message propagation method based on the graph neural network is used to propagate the storage information inside each node to other nodes with the starting node as the message propagation starting point, and the loop node is a node on any loop in the data to be processed; If it is determined that the starting node is a loop node, when the starting node is used as the starting point of message propagation for the first time, the stored information inside each node is propagated to other nodes through the message propagation method based on graph neural network.
5. The method according to claim 4, characterized in that The determining that the starting node is not a loop node includes: If the preset loop node set does not include the starting node, determining that the starting node is not a loop node; Among them, the preset loop node set is obtained through the following steps: through the message propagation method based on graph neural network, taking the starting node as the starting point of message propagation, propagating the test message from the starting node to other nodes, and adding the nodes that send the test message more than the preset number of times during the propagation process to the preset loop node set.
6. The method according to any one of claims 1 to 5, characterized in that: The message propagation method based on graph neural network, taking the starting node as the message propagation starting point, propagates the storage information inside each node to other nodes, including: Divide the starting nodes into first-type nodes and second-type nodes, wherein the first-type nodes have connection relationships with at least a preset number of other nodes, and the second-type nodes are nodes other than the first-type nodes among all nodes of the data to be processed; For the first-type nodes, propagating the storage information inside each of the first-type nodes to other nodes according to a first preset strategy; For the second-type nodes, the storage information inside each of the second-type nodes is propagated to other nodes according to a second preset strategy, where the second preset strategy is different from the first preset strategy.
7. The method according to claim 6, characterized in that After obtaining the processed data, the method further includes: Receiving a path query request sent by a terminal device, wherein the path query request includes identity information of a target start node and a target end node; According to the identity information of the target starting node and the target ending node, in the processed data, all information transmission paths between the target starting node and the target ending node are determined; In response to the path query request, all the information transmission paths are sent to the terminal device.
8. A data processing device, characterized in that: include: An acquisition module is used to acquire data to be processed, wherein the data structure of the data to be processed is a directed graph structure, wherein the data to be processed includes a plurality of nodes, each of the nodes is connected to other nodes through an edge, the nodes represent entity objects in the data to be processed, and the edges represent relationships between different entity objects; A first determination module is used to determine the starting nodes corresponding to all the edges in the data to be processed; The processing module is used to propagate the storage information inside each node to other nodes through a message propagation method based on a graph neural network, taking the starting node as the starting point of message propagation, to obtain processed data, wherein the storage information includes the information transmission path between the node and other nodes, and the processed data includes the information transmission path between any two different nodes.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the data processing method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.
11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.