Graph data processing method and device, equipment and medium

By separating directed graph data into symmetric output matrix and incoming matrix, the problem that traditional graph convolutional neural networks cannot handle directed graphs is solved, and the effective application of symmetric graph neural networks is realized, and the accuracy and efficiency of the model are improved.

CN120256997APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410004450.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional graph convolutional neural networks cannot effectively process directed graph data because their asymmetry leads to the loss of directional information of edges, affecting model performance and result accuracy.

Method used

The matrix data of directed graph data is separated into a symmetrical outgoing matrix and incoming matrix, which characterizes the outgoing relationship and incoming relationship of nodes, and inputs the symmetrical graph neural network for data analysis, so as to preserve edge direction information while reducing model complexity.

Benefits of technology

It improves the accuracy and comprehensiveness of information expression of graph data processing, reduces model complexity and storage requirements, and improves training efficiency and result accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256997A_ABST
    Figure CN120256997A_ABST
Patent Text Reader

Abstract

The invention provides a graph data processing method and device, equipment and a medium, relates to the technical field of artificial intelligence, can be applied to scenes such as cloud technology, artificial intelligence, intelligent traffic and auxiliary driving, and comprises the following steps: if a data structure of to-be-processed graph data is a directed graph, obtaining first matrix data and first graph attribute characteristics of the to-be-processed graph data; the first matrix data comprises an out-degree matrix and an in-degree matrix, elements of the out-degree matrix are used for representing an outward connection relationship between a corresponding node in the to-be-processed graph data and other nodes, and elements of the in-degree matrix are used for representing a connection relationship between the corresponding node in the to-be-processed graph data and the other nodes; the out-degree matrix and the in-degree matrix are symmetrical matrixes; performing data analysis of the first matrix data and the first graph attribute features based on the target graph neural network to obtain a graph analysis result; by means of the conversion method, direction information of the directed graph data can be reserved, and meanwhile complex conversion is not introduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, apparatus, device, and medium for graph data processing. Background Art

[0002] With the development of artificial intelligence technology, graph convolutional neural networks introduce spatial convolution operations on graph data, which can successfully process structured data and achieve excellent performance on undirected graphs. However, traditional graph convolutional neural networks require the input graph data to have symmetry, mainly from undirected graphs. In practical problems, many scenarios need to construct directed graphs, and their asymmetric graph data cannot be used as the input of the above graph convolutional neural networks. Therefore, it is necessary to add transpose matrices of the directed graph data to obtain symmetric graph data, but this processing method will cause the loss of edge direction information, affecting the model performance and result accuracy. Summary of the Invention

[0003] This application provides a method, apparatus, device, and medium for graph data processing, which can significantly improve the accuracy and comprehensiveness of information expression in graph data processing.

[0004] On the one hand, this application provides a method for graph data processing, and the method includes:

[0005] If the data structure of the graph data to be processed is a directed graph, obtain the first matrix data and the first graph attribute features of the graph data to be processed; the first matrix data includes an out-degree matrix and an in-degree matrix, the elements of the out-degree matrix are used to represent the outward connection relationship between the corresponding node and other nodes in the graph data to be processed, the elements of the in-degree matrix are used to represent the connection relationship that the corresponding node in the graph data to be processed receives from other nodes, and both the out-degree matrix and the in-degree matrix are symmetric matrices; the first graph attribute features are used to represent the node features of each node in the graph data to be processed;

[0006] Perform data analysis on the first matrix data and the first graph attribute features based on a target graph neural network to obtain a graph analysis result.

[0007] On the other hand, a graph data processing apparatus is provided, and the apparatus includes:

[0008] Acquisition module: used to obtain the first matrix data and the first graph attribute features of the graph data to be processed if the data structure of the graph data to be processed is a directed graph; the first matrix data includes an out-degree matrix and an in-degree matrix, the elements of the out-degree matrix are used to represent the outward connection relationship between the corresponding node and other nodes in the graph data to be processed, the elements of the in-degree matrix are used to represent the connection relationship of the corresponding node in the graph data to be processed receiving connections from other nodes, and both the out-degree matrix and the in-degree matrix are symmetric matrices; the first graph attribute features are used to represent the node features of each node in the graph data to be processed.

[0009] Analysis module: used to perform data analysis on the first matrix data and the first graph attribute features based on a target graph neural network to obtain a graph analysis result.

[0010] On the other hand, a computer device is provided, the device includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the graph data processing method as described above.

[0011] On the other hand, a computer-readable storage medium is provided, and at least one instruction or at least one program segment is stored in the storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the graph data processing method as described above.

[0012] On the other hand, a server is provided, the server includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the graph data processing method as described above.

[0013] On the other hand, a terminal is provided, the terminal includes a processor and a memory, and at least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the graph data processing method as described above.

[0014] On the other hand, a computer program product or a computer program is provided, the computer program product or the computer program includes computer instructions, and when the computer instructions are executed by a processor, the graph data processing method as described above is implemented.

[0015] The graph data processing method, device, equipment, storage medium, server, terminal, computer program and computer program product provided by this application have the following technical effects:

[0016] When the data structure of the graph data to be processed in the technical solution of this application is a directed graph, the first matrix data and the first graph attribute features of the graph data to be processed are obtained; the first matrix data includes an out-degree matrix and an in-degree matrix. The elements of the out-degree matrix are used to represent the outward connection relationship between the corresponding node and other nodes in the graph data to be processed, and the elements of the in-degree matrix are used to represent the connection relationship that the corresponding node in the graph data to be processed receives from other nodes. Both the out-degree matrix and the in-degree matrix are symmetric matrices; the first graph attribute features are used to represent the node features of each node in the graph data to be processed; the first matrix data and the first graph attribute features are input into a target graph neural network for data analysis to obtain a graph analysis result; in this way, the matrix data of the directed graph data is separated into an out-degree matrix and an in-degree matrix that both have symmetric properties to respectively represent the out-degree relationship and the in-degree relationship of the nodes. Furthermore, while not losing the edge direction information, existing symmetric graph neural networks can be used for data analysis, improving the result accuracy and avoiding affecting the model performance. And, the above solution does not need to use an asymmetric graph neural network that can distinguish out-edges and in-edges, thereby reducing the model complexity and the parameter quantity requirements, ensuring the model training efficiency, and reducing the storage and memory occupancy requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] To more clearly illustrate the technical solutions and advantages in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is a schematic diagram of an application environment provided by an embodiment of this application;

[0019] Figure 2 is a schematic flowchart of a graph data processing method provided by an embodiment of this application;

[0020] Figure 3 is an example diagram of the graph data to be processed in the form of a directed graph provided by an embodiment of this application;

[0021] Figure 4 is a schematic flowchart of another graph data processing method provided by an embodiment of this application;

[0022] Figure 5 is a schematic flowchart of another graph data processing method provided by an embodiment of this application;

[0023] Figure 6 is a schematic flowchart of another graph data processing method provided by an embodiment of this application;

[0024] Figure 7 It is a schematic flowchart of another method for processing graph data provided by an embodiment of the present application;

[0025] Figure 8 It is a schematic framework diagram of a target graph neural network provided by an embodiment of the present application;

[0026] Figure 9 It is a schematic framework diagram of a graph data processing device provided by an embodiment of the present application;

[0027] Figure 10 It is a hardware structure block diagram of an electronic device for executing the method for processing graph data provided by an embodiment of the present application. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0029] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or sub-modules does not necessarily have to be limited to those steps or sub-modules clearly listed, but may include other steps or sub-modules that are not clearly listed or are inherent to these processes, methods, products or devices.

[0030] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0031] Graph Convolution Neural Networks (GCN): It is a neural network framework for processing graph-structured data, which generalizes the convolution operation from traditional data (images or grids) to graph data.

[0032] Adjacency Matrix: It is a matrix representing the adjacent relationship between vertices.

[0033] Directed graph neural network: A graph neural network where the connections between nodes have directionality, with an asymmetric adjacency matrix.

[0034] Undirected graph neural network: A graph neural network where the connections between nodes do not have directionality, with a symmetric adjacency matrix.

[0035] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0036] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, large graph data processing technology, pre-trained model technology, operating / interactive systems, and mechatronics. Among them, the pre-trained model, also known as the large model or the foundation model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0037] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. The pre-trained model is the latest development result of deep learning, integrating the above technologies.

[0038] Deep learning: The concept of deep learning originated from the research of artificial neural networks. The multi-layer perceptron with multiple hidden layers is a deep learning structure. Deep learning forms more abstract high-level representations (attribute categories or features) by combining low-level features to discover the distributed feature representations of data.

[0039] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0040] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application environment provided by an embodiment of this application. As Figure 1 shown, this application environment may at least include a terminal 01 and a server 02. In practical applications, the terminal 01 and the server 02 may be directly or indirectly connected through wired or wireless communication methods, and this application does not make any restrictions here.

[0041] The server 02 in the embodiment of this application may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0042] Specifically, cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. Cloud technology can be applied in various fields, such as medical cloud, cloud IoT, cloud security, cloud education, cloud conferencing, artificial intelligence cloud services, cloud applications, cloud calling, and cloud social networking. Cloud technology is applied based on the cloud computing business model. It distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to be infinitely expandable to users, and can be obtained at any time, used on demand, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool (referred to as a cloud platform, generally called IaaS (Infrastructure as a Service)) platform will be established, and various types of virtual resources will be deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.

[0043] According to the logical function division, the PaaS (Platform as a Service) layer can be deployed on the IaaS layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. Alternatively, the SaaS can be directly deployed on the IaaS. PaaS is the platform for software operation, such as databases, web containers, etc. SaaS includes various business software, such as web portals, mass text message senders, etc. Generally speaking, SaaS and PaaS are the upper layers relative to IaaS.

[0044] Specifically, the server 02 involved above can include physical devices, which can specifically include network communication sub-modules, processors, memories, etc., or can include software running on physical devices, which can specifically include application programs, etc.

[0045] Specifically, the terminal 01 can include physical devices such as smartphones, desktop computers, tablet computers, laptop computers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, intelligent voice interaction devices, smart home appliances, intelligent wearable devices, in-vehicle terminal devices, etc., or can include software running on physical devices, such as application programs, etc.

[0046] In the embodiments of the present application, the terminal 01 can send the to-be-processed graph data to the server 02, so that the server 02 can obtain the first matrix data and the first graph attribute features of the to-be-processed graph data, and then perform data analysis; or, the terminal 01 can send the interaction data between nodes, including interaction operation data, etc., to the server 02, so that the server 02 can construct the to-be-processed graph data based on the interaction data, and then perform matrix data and feature acquisition, as well as corresponding data analysis; or, the terminal 01 can also run the target graph neural network to directly perform matrix data, feature acquisition, and corresponding data analysis on the to-be-processed graph data, obtain the graph analysis result and send it to the server 02 for storage.

[0047] In addition, it can be understood that Figure 1 The shown is only an application environment of a graph data processing method. This application environment can include more or fewer nodes, and the present application does not limit this here.

[0048] The application environment involved in the embodiments of the present application, or terminals 01 and servers 02 in the application environment, etc. can be a distributed system formed by connecting a client and multiple nodes (any form of computing device accessing the network, such as a server, a terminal) through network communication. The distributed system can be a blockchain system, and the blockchain system can provide the above-mentioned graph data processing service, model training service, data storage service, etc.

[0049] The following introduces the technical solution of the present application based on the above application environment and dialogue system. The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, intelligent transportation, assisted driving, etc. Please refer to Figure 2 , Figure 2 is a schematic flowchart of a graph data processing method provided by an embodiment of the present application. This specification provides method operation steps such as in the embodiment or flowchart, but based on routine or non-creative labor, it can include more or fewer operation steps. The step order listed in the embodiment is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual system or server product executes, it can be executed in the order shown in the embodiment or the drawing or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Specifically, as Figure 2 shown, the method can include the following steps S201 - S203:

[0050] S201: If the data structure of the graph data to be processed is a directed graph, obtain the first matrix data and the first graph attribute features of the graph data to be processed.

[0051] Specifically, the data structure of the graph data to be processed can be a directed graph or an undirected graph. A directed graph refers to a graph data structure in which the edges connecting nodes have directional information, and an undirected graph refers to a graph data structure in which the edges connecting nodes do not have directional information. Graph data includes nodes and the edges between nodes. Nodes can refer to the subjects to be analyzed, such as accounts, terminal devices, etc., and edges can refer to the interactions between subjects, such as interaction operations, etc. It can be understood that in application scenarios, many graph data exist in the form of directed graphs. Taking the payment network scenario as an example, the nodes in the graph data are payment accounts, and the connection relationships between the nodes represent the collection relationships such as transfer payments between accounts and have directions. Refer to Figure 3 , Figure 3 shows a simplified directed graph between payment accounts. In this example, there are a total of 6 payment account nodes, and the node sorting is represented by numbers 1 - 6. Some nodes have both out-edges pointing to other nodes and in-edges coming from other nodes, some nodes only have out-edges pointing to other nodes, and some nodes only have in-edges coming from other nodes.

[0052] Understandably, the adjacency matrix of directed graph data is usually an asymmetric matrix. For example, Figure 3 The adjacency matrix of the graph data in the figure is shown in the following table. The adjacency matrix represents the matrix of the adjacent relationship between vertices in the graph. The adjacency matrix of the above directed graph can be expressed as (the part in the dotted box). Among them, both the rows and columns are sorted based on the node numbers. The value of the i-th row and j-th column represents whether there is an edge pointing from node i to node j (for example, this edge can represent whether account i pays to account j). For example, the value of the element (2,1) in the second row and the first column is 0, indicating that there is no edge from node 2 to node 1. The value of the element (1,3) in the first row and the third column is 1, indicating that there is an edge from node 1 to node 3.

[0053]

[0054] Traditional graph neural networks only receive symmetric matrix data. For example, Graph Convolutional Networks (GCN) introduce spatial convolution operations on graph data and can successfully process structured data. It uses a symmetric normalized Laplacian matrix for convolution operations and requires the adjacency matrix to be symmetric, that is, the graph structure needs to be an undirected graph. For a directed graph, the generated adjacency matrix data is an asymmetric matrix and cannot be eigen-decomposed. Specifically, in GCN, the graph convolution operation of each layer depends on the symmetric normalized Laplacian matrix. Given the adjacency matrix A and degree matrix D (the diagonal elements are the degrees of the nodes, that is, the number of edges connected to the node) of an undirected graph, the symmetric normalized Laplacian matrix can be calculated as Since the adjacency matrix A of an undirected graph is itself symmetric, the calculated symmetric normalized Laplacian matrix L is also symmetric. In the convolution operation of the GCN layer, the symmetric normalized Laplacian matrix L can make the model more stable and is beneficial to calculating gradients during backpropagation. When processing a directed graph, the adjacency matrix is usually asymmetric and cannot be eigen-decomposed in a symmetric graph neural network. Applying it to GCN will cause inaccurate information propagation, thereby affecting the model performance and unable to obtain effective analysis results.

[0055] In the related art, the adjacency matrix A of a directed graph is added to its transpose matrix to obtain a symmetric adjacency matrix. That is, let the adjacency matrix of the directed graph be A, and its (i,j) element represents the weight of the edge from node i to node j. The conversion steps include: 1. Calculate the element-wise sum of the adjacency matrix A and its transpose matrix A T : A sym = A + A T ; 2. Use A sym to replace the original adjacency matrix A of the directed graph. Through the above conversion process, the asymmetric adjacency matrix is transformed into a symmetric matrix, but this simplification process will cause the loss of edge direction information, thereby affecting the performance of the model.

[0056] In related technologies, an asymmetric graph neural network is also used to directly receive asymmetric matrix data for data analysis. For example, GAT (Graph Attention Networks) is used. GAT is a graph neural network based on the self-attention mechanism, which can better propagate information on a directed graph. By introducing attention weights, GAT can distinguish incoming edges and outgoing edges to handle the directional features of a directed graph. However, compared with symmetric graph neural networks (such as GCN), its model complexity and the number of parameters are greatly increased, the computing resources required for the training process are also significantly increased, it has high storage and memory occupancy requirements, and the consumption of computing resources is also relatively high.

[0057] Combined with the above, the first matrix data of the present application includes an out-degree matrix and an in-degree matrix. Both the out-degree matrix and the in-degree matrix are symmetric matrices. The out-degree matrix and the in-degree matrix are obtained by splitting the adjacency matrix of the graph data to be processed according to the out-degree relationship and the in-degree relationship of the nodes in the graph data to be processed. The elements of the out-degree matrix are used to represent the outward connection relationship between the corresponding node and other nodes in the graph data to be processed, that is, to represent its out-degree relationship (such as an outward payment relationship), and the elements of the in-degree matrix are used to represent the connection relationship of the corresponding node in the graph data to be processed receiving other nodes, that is, to represent its in-degree relationship (such as an inward collection relationship). Here, the out-degree relationship can represent the relevant connection information of the edges sent from the current node to other nodes, and the in-degree relationship can represent the relevant connection information of the edges sent from other nodes to the current node. Exemplarily, referring to Figure 3 , the out-degree relationship of node 2 is that there are three outgoing edges pointing to node 3, node 4, and node 5, and the in-degree relationship is that there are no incoming edges; the out-degree relationship of node 4 is that there is an outgoing edge pointing to node 5, and the in-degree relationship is that there are two incoming edges pointing to its own node from node 1 and node 2. Through the above matrix splitting, the problem of asymmetry between the out-degree connection relationship and the in-degree connection relationship is eliminated, and then symmetric matrix data is obtained, which not only retains the connection direction information of the original graph data, but also does not introduce a conversion method with complex transformations. It can be used as the input of a symmetric graph neural network, reducing the model occupancy and resource consumption requirements, and improving the model generalization and analysis accuracy.

[0058] Specifically, the first graph attribute feature is used to characterize the node features of each node in the graph data to be processed, and can be obtained by splicing the node features of each node in the graph data to be processed. Each node here can refer to all nodes in the graph data to be processed. The node feature is used to characterize the information of the corresponding node, including the node basic feature to characterize the basic attributes of the node. Exemplarily, in the scenario where the node is a payment account, the node basic feature can be, but is not limited to, registration time feature, balance feature, device identification feature, region feature where it belongs, etc. In some cases, the node feature can also include the out-degree feature or in-degree feature, etc. The out-degree feature refers to the feature associated with the out-degree relationship of the node, such as features related to payment actions such as payment frequency and payment amount of the payment account. The in-degree feature refers to the feature associated with the in-degree relationship of the node, such as features related to receipt actions such as receipt frequency and receipt amount of the payment account. In some cases, the out-degree feature can also include the out-degree value or in-degree value of the node, such as Figure 3 in Figure 3 , the out-degree value of node 2 is 3 and the in-degree value is 0.

[0059] S203: Perform data analysis on the first matrix data and the first graph attribute feature based on the target graph neural network to obtain a graph analysis result.

[0060] In some embodiments, the first matrix data and the first graph attribute feature can be directly used as the input of the target graph neural network to obtain a graph analysis result. In other embodiments, input data can be generated based on the first matrix data and the first graph attribute feature, and then input into the target graph neural network for data analysis to obtain a graph analysis result.

[0061] Specifically, the target graph neural network is a symmetric graph neural network, such as GCN, etc.

[0062] It can be understood that the data analysis method can be set based on the scenario and the output requirements of the target graph neural network. For example, if the output requirement is the overall category index data of the graph data to be processed, correspondingly, the graph analysis result is used to characterize the category to which the graph data to be processed belongs. For example, in the binary classification recognition scenario of a risk account group, the graph analysis result characterizes that the node network corresponding to the graph data to be processed is a risk account group or a normal account group. Or, the graph analysis result can also characterize the node category information of each node in the graph data to be processed, such as characterizing each node as a risk account or a normal account.

[0063] In summary, the present application separates the matrix data of the directed graph data into an out-degree matrix and an in-degree matrix both having symmetric properties to respectively represent the out-degree relationship and the in-degree relationship of nodes. Furthermore, while not losing the edge direction information, existing symmetric graph neural networks can be used for data analysis, improving the result accuracy and avoiding affecting the model performance. Moreover, the above solution does not require an asymmetric graph neural network that can distinguish out-edges and in-edges, thereby reducing the model complexity and the parameter quantity requirements, ensuring the model training efficiency, and reducing the storage and memory occupancy requirements.

[0064] Based on some or all of the above embodiments, in some embodiments, referring to Figure 4 , the method for obtaining the first matrix data includes S301-S307:

[0065] S301: Obtain the node sorting of each node in the graph data to be processed, and the out-degree relationship and the in-degree relationship of each node;

[0066] S303: When the out-degree relationship indicates that there is an out-edge between its corresponding node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the out-edge to the first value. When the out-degree relationship indicates that there is an in-edge or no edge between the node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the in-edge or the element formed by the two nodes with no edge to the second value to obtain the out-degree matrix;

[0067] S305: When the in-degree relationship indicates that there is an in-edge between its corresponding node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the in-edge to the first value. When the in-degree relationship indicates that there is an out-edge or no edge between the node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the out-edge or the element formed by the two nodes with no edge to the second value to obtain the in-degree matrix;

[0068] S307: Generate the first matrix data based on the out-degree matrix and the in-degree matrix.

[0069] Specifically, the node sorting can be set according to the scenario requirements and is not limited herein. For example, based on the account registration time, the earlier the registration time, the more forward the sorting position of the node, and vice versa. Taking Figure 3 as an example, the smaller the node number, the more forward its sorting, and vice versa.

[0070] Specifically, the out-degree relationship represents the relevant connection information of the edge sent from the current node to other nodes, and the in-degree relationship represents the relevant connection information of the edge sent from other nodes to the current node. Obtain the out-degree relationship and the in-degree relationship of each node in the graph data to be processed to clarify the out-edge, in-edge, and connected nodes of each node. The out-edge refers to the edge executed by the current node to other nodes, and the in-edge refers to the edge pointed from other nodes to the current node.

[0071] Specifically, the first value represents the existence of a target out-edge in the out-degree matrix and the existence of a target in-edge in the in-degree matrix; the second value represents the non-existence of a target out-edge in the out-degree matrix and the non-existence of a target in-edge in the in-degree matrix. A target out-edge refers to an edge from the current node to another node with a later node order. If there is no edge from the current node to another node with a later node order, it is determined that there is no target out-edge. A target in-edge is an edge from another node with a later node order to the current node. If there is no edge from another node with a later node order to the current node, it is determined that there is no target in-edge.

[0072] Specifically, based on the out-degree relationship, each target out-edge corresponding to the out-degree matrix is determined, and the values of the elements (i1, j1) and (j1, i1) formed by its two end nodes are determined as the first value (such as 1), and the values of other elements are determined as the second value (such as 0); based on the in-degree relationship, each target in-edge corresponding to the in-degree matrix is determined, and the values of the elements (i2, j2) and (j2, i2) formed by its two end nodes are determined as the first value (such as 1), and the values of other elements are determined as the second value (such as 0).

[0073] Take Figure 3 as an example. Only considering the out-degree relationship of the nodes, traverse from node 1 to node N (N = 6 in the figure) in the order of node numbers. For node i, if there is a directed edge from node i to node j (j > i), the elements (i, j) and (j, i) of the matrix are both assigned the value 1, otherwise 0. That is, the following out-degree matrix is obtained, and this out-degree matrix is a symmetric matrix.

[0074]

[0075] Similarly, only considering the in-degree relationship of the nodes, traverse from node 1 to node N in the order of node numbers. For node i, if there is a directed edge from node j (j > i) to node i, the elements (i, j) and (j, i) of the matrix are both assigned the value 1, otherwise 0. That is, the following out-degree matrix is obtained, and this in-degree matrix is a symmetric matrix.

[0076]

[0077] Combining the above, the out-degree and in-degree of the connection relationship are separated. The out-degree represents the connection relationship from the node outward (such as the payment relationship), and the in-degree represents the connection relationship of the node receiving connections from other nodes (such as the receipt relationship). Furthermore, based on the splitting of the out-degree relationship and the in-degree relationship, the connection relationship representing between nodes is separated into a symmetric out-degree matrix and an in-degree matrix. Furthermore, the two can be combined to obtain an updated adjacency matrix E edge , this updated adjacency matrix E edgeIt has symmetry, can be convolved and plugged in, and can be directly applied to the target graph convolutional neural network; it eliminates the problem of asymmetry between the out-degree connection relationship and the in-degree connection relationship in directed graph data. Compared with existing methods for processing directed graphs such as GCN, it not only retains the connection direction information of the original graph but also does not introduce complex transformation methods. Exemplarily, the updated adjacency matrix E edge can be expressed as follows: E 出 represents the out-degree matrix, and E 入 represents the in-degree matrix.

[0078]

[0079] In some embodiments, the node features include node basic features, which are used to characterize the attributes of the corresponding nodes and are independent of the directions of the in-degree and out-degree of the corresponding nodes. The node basic features are obtained by feature representation based on node basic data, which characterize the basic attributes of the nodes and are independent of the out-actions (such as payment actions) and in-actions (such as receipt actions) of the nodes. The node basic data can be, for example, registration time, account balance, device identifier, region of belonging, etc., and correspondingly obtain registration time features, balance features, device identifier features, region of belonging features, etc. Training the target graph neural network through the node basic features to construct the association between node attributes, node operations, and graph network characteristics can improve the accuracy of model analysis.

[0080] In some embodiments, the node features include at least one of the out-degree feature and the in-degree feature. The out-degree feature is associated with the out-edge connection relationship of the corresponding node, that is, related to the out-action of the corresponding node, and is obtained by feature representation of the out-degree data of the corresponding node. The out-degree data can be, for example, payment frequency, payment amount, etc. related to payment actions of the payment account; the in-degree feature is associated with the in-edge connection relationship of the corresponding node, that is, related to the in-action of the corresponding node, and is obtained by feature representation of the in-degree data of the corresponding node. The in-degree data can be, for example, receipt frequency, receipt amount, etc. related to receipt actions of the payment account. Associating the out-action and in-action of the node through the out-degree feature or the in-degree feature is convenient for the target graph neural network to master the association relationship between operation data and analysis results based on the graph connection relationship, and further improve the model performance and result reliability.

[0081] In some embodiments, each node only includes node basic features. The method for obtaining the first graph attribute feature includes: using the node basic features as the first feature and the second feature, and splicing the first feature and the second feature to generate node features; and then generating the first graph attribute feature based on the node features of each node.

[0082] In one embodiment, the first graph attribute feature can be represented in matrix form. Exemplarily, the node feature X i' and the first graph attribute feature X node The expression is as follows, where X i1 represents the first feature, and X i2 represents the second feature. Understandably, the node feature X i ' is a feature matrix, and X node is the feature matrix obtained by splicing each node matrix.

[0083]

[0084] X node = [X1', X2', …, X' N

[0085] Assume that the node base feature of each node i is Xi, that is, the feature of each node is a feature independent of the direction of in-degree and out-degree. Then, the first feature X i1 and the second feature X i2 of the node are both the node base feature, which is Xi.

[0086] In some other embodiments, in addition to the node base feature, each node may further include at least one of the out-degree feature and the in-degree feature. Correspondingly, in some embodiments, referring to Figure 5 , the acquisition method of the first graph attribute feature includes S401 - S409:

[0087] S401: Obtain the node base feature, out-degree feature, and in-degree feature of each node in the graph data to be processed;

[0088] S403: Fuse the node base feature and the out-degree feature to obtain the first feature of the node;

[0089] S405: Fuse the node base feature and the in-degree feature to obtain the second feature of the node;

[0090] S407: Generate the node feature of the node based on the first feature and the second feature;

[0091] S409: Splice the node features of each node based on the node sorting to obtain the first graph attribute feature.

[0092] Specifically, the fusion here may refer to splicing, etc. The node feature and the first graph attribute feature may exist in matrix form, and may be specifically consistent with the expressions of the above node feature and the first graph attribute feature.

[0093] ​Understandably, when there is only an out-degree feature for a single node, the second feature can be the node basic feature, or the feature segments of the in-degree feature can be filled with blank features, such as filling with 0. Or, when there is only an in-degree feature for a single node, the first feature can be the node basic feature, or the feature segments of the out-degree feature can be filled with blank features. By separating the out-degree feature and the in-degree feature, the features of the nodes are classified into two categories according to the action type, further improving the model's learning of the directional information in the directed graph and enhancing the accuracy of the training effect and the analysis result.

[0094] In some other embodiments, after obtaining the node basic feature of the node and the out-degree feature or the in-degree feature, they can be directly fused to generate a node feature, and then a first graph attribute feature is generated.

[0095] In some embodiments, S401 may include S4011 - S4012:

[0096] S4011: Obtain the node basic data, out-degree data, and in-degree data of the nodes in the to-be-processed graph data;

[0097] S4012: Respectively perform feature representation on the node basic data, node data, and in-degree data to obtain the node basic feature, out-degree feature, and in-degree feature.

[0098] Specifically, the node basic data is data that has nothing to do with the in-out degree direction of the corresponding node, the out-degree data is data associated with the out-degree connection relationship of the corresponding node, that is, data associated with the out-action, and the in-degree data is data associated with the in-degree connection relationship of the corresponding node, that is, data associated with the in-action. Here, the feature representation can be feature encoding, such as one-hot encoding, or, in some cases, the feature representation can also be implemented in a numerical mapping manner, such as mapping the account balance to a specified numerical interval, and other feature representation methods can also be used, which are not specifically limited here.

[0099] Based on some or all of the above embodiments, in some embodiments, the data format of the to-be-processed graph data may also be an undirected graph, and its corresponding adjacency matrix is itself a symmetric matrix. Correspondingly, referring to Figure 6 , the method may further include S501 - S503:

[0100] S501: If the data structure of the to-be-processed graph data is an undirected graph, obtain the second matrix data and the second graph attribute feature of the to-be-processed graph data;

[0101] S503: Input the second matrix data and the second graph attribute feature into the target graph neural network for data analysis to obtain a graph analysis result.

[0102] Specifically, the second matrix data includes an out-degree matrix and an in-degree matrix, both of which are adjacency matrices of the graph data to be processed; the second graph attribute feature is used to characterize the node features of each node in the graph data to be processed. The second matrix data is a symmetric matrix data composed of two adjacency matrices, which is aligned with the input data of the directed graph, so that the generalization application of the model in the directed graph data and the undirected graph data can be realized without changing the model structure and parameters, and it can be applied to complex data scenarios.

[0103] Correspondingly, in some embodiments, referring to Figure 7 , the acquisition method of the second graph attribute feature includes S601-S609:

[0104] S601: Obtain the node basic features of each node in the graph data to be processed;

[0105] S603: Determine the node basic feature as the first feature of the node;

[0106] S605: Determine the node basic feature as the second feature of the node;

[0107] S607: Generate the node feature of the node based on the first feature and the second feature;

[0108] S609: Concatenate the node features of each node based on the node sorting to obtain the second graph attribute feature.

[0109] It can be understood that in an undirected graph, each feature of a node has nothing to do with the in-out degree direction, that is, it can be used as a node basic feature. By determining it as the first feature and the second feature, the node feature in matrix format can be further generated, and then the second graph attribute feature in matrix format can be obtained to align with the first graph attribute feature of the directed graph, so as to further improve the model generalization and the accuracy of result analysis.

[0110] Based on some or all of the above embodiments, in some embodiments, the target graph neural network includes a feature network and a classifier, and the above S203 may include:

[0111] S701: Perform feature extraction on the first matrix data and the first graph attribute feature based on the feature network to obtain graph extraction features;

[0112] S703: Input the graph extraction features into the classifier for classification and recognition to obtain a graph analysis result, which is used to characterize the graph category of the graph data to be processed and / or to characterize the node category of each node in the graph data to be processed.

[0113] In some embodiments, for a directed graph, the first matrix data and the first graph attribute features are directly input into the feature network to output graph extraction features; for an undirected graph, the second matrix data and the second graph attribute features are directly input into the feature network to output graph extraction features; and the graph extraction features are input into a classifier to obtain a graph analysis result. In this embodiment, the node out-degree feature may include an out-degree value, and the node in-degree feature may include an in-degree value, so as to cover the in-out degree information of the nodes in the graph attribute feature matrix.

[0114] In other embodiments, for a directed graph, the first degree matrix and the second degree matrix of the graph data to be processed may also be obtained. The first degree matrix and the second degree matrix are diagonal matrices. The elements on the diagonal of the first degree matrix are the out-degrees of the respective nodes, and the elements on the diagonal of the second degree matrix are the in-degrees of the respective nodes. The out-degree refers to the number of edges that the node points to other nodes, and the in-degree refers to the number of edges that other nodes point to this node. Or, in the case where the weights of each edge are inconsistent, the out-degree refers to the result of multiplying each edge that the node points to other nodes by their respective weight values and adding them up, and the in-degree refers to the result of multiplying each edge that other nodes point to this node by their respective weight values and adding them up.

[0115] Further, a first out matrix is generated based on the first degree matrix and the aforementioned out-degree matrix, a first in matrix is generated based on the second degree matrix and the aforementioned in-degree matrix, and the first out matrix and the second in matrix are combined to generate an input matrix, so as to be used as the input of the target graph neural network together with the first graph attribute feature matrix.

[0116] In one example, the first degree matrix D1 and the out-degree matrix A1 can generate the first out matrix L in the following manner 1出 .

[0117]

[0118] In one example, the second degree matrix D2 and the in-degree matrix A2 can generate the first in matrix L in the following manner 1入 .

[0119]

[0120] The input matrix is

[0121] For an undirected graph, the degree matrix of the graph data to be processed can also be obtained. The degree matrix is a diagonal matrix, and the elements on the diagonal of the degree matrix are the degrees of the nodes, indicating the number of edges associated with the node. Or, in the case where the weights of each edge are inconsistent, the degree of a node refers to the result of multiplying each edge associated with the node by its respective weight value and summing them. Further, a second out-degree matrix and a second in-degree matrix are generated based on the degree matrix and the aforementioned adjacency matrix, and an input matrix is generated based on the second out-degree matrix and the second in-degree matrix to be used as the input of the target graph neural network together with the second graph attribute feature matrix. The second out-degree matrix and the second in-degree matrix are the same.

[0122] In one example, the second out-degree matrix L 2出 and the second in-degree matrix L 2入 can both be generated based on the degree matrix D3 and the adjacency matrix A3 in the following manner.

[0123]

[0124] The input matrix is

[0125] In this way, superimposing the degree matrix increases the model's learning of the degree information of the graph data and establishes its correlation with the analysis result, further improving the reliability of the result.

[0126] It can be understood that the graph analysis result is strongly related to the application scenario and can be set based on the scenario requirements, such as being used to represent the category to which the graph data to be processed belongs, or representing the node category information of each node in the graph data to be processed. The former represents the overall network category of the graph data to be processed, and the latter represents the category of each individual node. Accordingly, different classifiers can be used to implement this. In some embodiments, the same feature network can be used to extract features and different task classifiers can be connected to achieve multi-task result output.

[0127] In one embodiment, referring to Figure 8, the feature network may include M layers of graph convolutional networks (such as GCN), and M can be, for example, 5 - 10 layers. The constructed network weight parameters are randomly initialized. The target graph convolutional network can be obtained by constraining and training the initial graph neural network with sample data as input and sample labels as the expected output. The sample data can include the sample matrix data of the sample graph data and the sample graph attribute features of the sample graph data. The sample graph data is similar to the graph data to be processed, the sample matrix data is similar to the aforementioned first matrix data or second matrix data, and the sample graph attribute features are similar to the aforementioned first graph attribute features or second graph attribute features, which will not be elaborated here. The sample labels can have different forms based on different tasks, such as representing the category to which the sample graph data belongs, or the category to which each node in the sample graph data belongs, etc. Taking the binary classification of the entire graph as an example, such as fraud account mining in payment, the output result is a value between [0, 1]. The entire model uses the standard neural network training method, that is, optimizing the objective function L loss to learn the network parameters. In one example, where S is the number of all samples, p i is the output of the initial graph neural network for sample i, and y i is the label of sample i.

[0128] In summary, the graph data processing method of the present application can transform the directed graph into an undirected symmetric graph, strip the out - degree connection matrix and in - degree connection matrix, and out - degree features and in - degree features of the directed graph nodes, eliminating the problem of asymmetry between the out - degree connection relationship and the in - degree connection relationship. Compared with the existing method of using GCN to process directed graphs, it simply retains the connection direction information of the original graph, and at the same time does not introduce a complex transformation method. It merges the out - degree - related matrix and features, and the in - degree - related matrix and features for unified training, improving the training and prediction efficiency, and enhancing the network performance and generalization ability.

[0129] The embodiment of the present application also provides a graph data processing device 800, as Figure 9 shown, Figure 9 shows a schematic structural diagram of a graph data processing device provided by the embodiment of the present application. The device may include the following modules.

[0130] Obtaining module 10: used to obtain the first matrix data and the first graph attribute features of the graph data to be processed if the data structure of the graph data to be processed is a directed graph; the first matrix data includes an out - degree matrix and an in - degree matrix. The elements of the out - degree matrix are used to represent the outward connection relationship between the corresponding node and other nodes in the graph data to be processed, and the elements of the in - degree matrix are used to represent the connection relationship that the corresponding node in the graph data to be processed receives from other nodes. Both the out - degree matrix and the in - degree matrix are symmetric matrices; the first graph attribute features are used to represent the node features of each node in the graph data to be processed;

[0131] Analysis module 20: configured to perform data analysis on the first matrix data and the first graph attribute features based on a target graph neural network to obtain a graph analysis result.

[0132] In some embodiments, the obtaining module 10 may include:

[0133] The first obtaining sub-module: configured to obtain the node sorting of each node in the graph data to be processed, and the out-degree relationship and in-degree relationship of each node;

[0134] The out-degree matrix sub-module: configured to, when the out-degree relationship indicates that there is an out-edge between its corresponding node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the out-edge to a first value, and when the out-degree relationship indicates that there is an in-edge or no edge between the node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the in-edge or the element formed by the two nodes with no edge to a second value, to obtain an out-degree matrix;

[0135] The in-degree matrix sub-module: configured to, when the in-degree relationship indicates that there is an in-edge between its corresponding node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the in-edge to a first value, and when the in-degree relationship indicates that there is an out-edge or no edge between the node and another node with a later node sorting, assign the element formed by the two nodes at both ends of the out-edge or the element formed by the two nodes with no edge to a second value, to obtain an in-degree matrix;

[0136] The first matrix generation sub-module: configured to generate first matrix data based on the out-degree matrix and the in-degree matrix.

[0137] In some embodiments, the node features include node basic features, and the node basic features are used to characterize the attributes of the corresponding nodes and are independent of the directions of the in-degree and out-degree of the corresponding nodes.

[0138] In some embodiments, the node features include at least one of an out-degree feature and an in-degree feature. The out-degree feature is associated with the out-edge connection relationship of the corresponding node, and the in-degree feature is associated with the in-edge connection relationship of the corresponding node.

[0139] In some embodiments, the obtaining module 10 may include:

[0140] The second obtaining sub-module: configured to obtain the node basic features, out-degree features, and in-degree features of each node in the graph data to be processed;

[0141] The first feature sub-module: configured to fuse the node basic features and the out-degree features to obtain the first feature of the node;

[0142] The second feature sub-module: configured to fuse the node basic features and the in-degree features to obtain the second feature of the node;

[0143] Node feature sub-module: used to generate the node features of nodes based on the first feature and the second feature;

[0144] Graph feature sub-module: used to splice the node features of each node based on the node sorting to obtain the first graph attribute feature.

[0145] In some embodiments, the second acquisition sub-module may include:

[0146] Data acquisition unit: used to acquire the node basic data, out-degree data and in-degree data of the nodes in the graph data to be processed. The node basic data is data independent of the in-out degree direction of the corresponding node. The out-degree data is data associated with the out-degree connection relationship of the corresponding node. The in-degree data is data associated with the in-degree connection relationship of the corresponding node;

[0147] Feature representation unit: used to respectively perform feature representation on the node basic data, node data and in-degree data to obtain the node basic feature, out-degree feature and in-degree feature.

[0148] In some embodiments, the acquisition module is further used to, if the data structure of the graph data to be processed is an undirected graph, acquire the second matrix data and the second graph attribute feature of the graph data to be processed. The second matrix data includes an out-degree matrix and an in-degree matrix, and both the out-degree matrix and the in-degree matrix are the adjacency matrices of the graph data to be processed; the second graph attribute feature is used to characterize the node features of each node in the graph data to be processed;

[0149] The analysis module is further used to input the second matrix data and the second graph attribute feature into the target graph neural network for data analysis to obtain the graph analysis result.

[0150] In some embodiments, the acquisition module may further include:

[0151] Third acquisition sub-module: used to acquire the node basic features of each node in the graph data to be processed;

[0152] The first feature sub-module is further used to determine the node basic feature as the first feature of the node;

[0153] The second feature sub-module is further used to determine the node basic feature as the second feature of the node;

[0154] The node feature sub-module is further used to generate the node features of the nodes based on the first feature and the second feature;

[0155] The graph feature sub-module is further used to splice the node features of each node based on the node sorting to obtain the second graph attribute feature.

[0156] In some embodiments, the target graph neural network includes a feature network and a classifier, and the analysis module may include:

[0157] Feature extraction sub-module: configured to perform feature extraction on the first matrix data and the first graph attribute features based on a feature network to obtain graph extraction features;

[0158] Graph analysis sub-module: configured to input the graph extraction features into a classifier for classification and recognition to obtain a graph analysis result, where the graph analysis result is used to characterize the graph category of the graph data to be processed and / or used to characterize the respective node categories of each node in the graph data to be processed.

[0159] It should be noted that the above device embodiments and method embodiments are based on the same implementation manner.

[0160] An embodiment of the present application provides a device, which can be a terminal or a server, including a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the graph data processing method provided by the above method embodiment.

[0161] The memory can be used to store software programs and modules. The processor runs the software programs and modules stored in the memory to perform various functional applications and anomaly detections. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for functions, etc.; the data storage area can store data created according to the use of the device, etc. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory can also include a memory controller to provide the processor with access to the memory.

[0162] The method embodiments provided by the embodiments of the present application can be executed in electronic devices such as mobile terminals, computer terminals, servers, or similar computing devices. Figure 10 It is a hardware structure block diagram of an electronic device for a graph data processing method provided by an embodiment of the present application. As Figure 10As shown, the electronic device 900 can vary significantly due to different configurations or performances. It may include one or more central processing units (CPUs) 910 (the processor 910 may include, but is not limited to, processing devices such as a microcontroller unit (MCU) or a field-programmable gate array (FPGA)), a memory 930 for storing data, and one or more storage media 920 for storing application programs 923 or data 922 (such as one or more mass storage devices). Among them, the memory 930 and the storage media 920 can be transient storage or persistent storage. The program stored in the storage media 920 may include one or more modules, and each module may include a series of instruction operations for the electronic device. Further, the central processor 910 can be configured to communicate with the storage media 920 and execute a series of instruction operations in the storage media 920 on the electronic device 900. The electronic device 900 may also include one or more power supplies 960, one or more wired or wireless network interfaces 950, one or more input / output interfaces 940, and / or one or more operating systems 921, such as Windows Server TM , Mac OS X TM , Unix TM , LinuxTM, FreeBSDTM, etc.

[0163] The input / output interface 940 can be used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the electronic device 900. In one example, the input / output interface 940 includes a network interface controller (NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one example, the input / output interface 940 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0164] Those of ordinary skill in the art can understand that Figure 10 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the electronic device 900 may also include more or fewer components than those shown Figure 10 in the figure, or have a different configuration from that shown Figure 10 in the figure.

[0165] An embodiment of the present application further provides a computer-readable storage medium, which can be disposed in an electronic device to store at least one instruction or at least one program related to an anomaly detection method in a method embodiment. The at least one instruction or the at least one program is loaded and executed by the processor to implement the anomaly detection method provided in the above method embodiment.

[0166] Optionally, in this embodiment, the above storage medium may be located in at least one network server among multiple network servers of a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs.

[0167] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above various optional implementation manners.

[0168] With the graph data processing method, device, equipment, storage medium, server, terminal, and program product provided by the present application, when the data structure of the graph data to be processed is a directed graph, the first matrix data and the first graph attribute features of the graph data to be processed are obtained; the first matrix data includes an out-degree matrix and an in-degree matrix. The elements of the out-degree matrix are used to represent the outward connection relationship between the corresponding node and other nodes in the graph data to be processed, and the elements of the in-degree matrix are used to represent the connection relationship that the corresponding node in the graph data to be processed receives from other nodes. Both the out-degree matrix and the in-degree matrix are symmetric matrices; the first graph attribute features are used to represent the node features of each node in the graph data to be processed; the first matrix data and the first graph attribute features are input into a target graph neural network for data analysis to obtain a graph analysis result; in this way, the matrix data of the directed graph data is separated into an out-degree matrix and an in-degree matrix both having symmetric properties to respectively represent the out-degree relationship and the in-degree relationship of the nodes, so that while not losing the edge direction information, the existing symmetric graph neural network can be used for data analysis, improving the result accuracy and avoiding affecting the model performance. Moreover, the above solution does not need to use an asymmetric graph neural network that can distinguish out-edges and in-edges, thereby reducing the model complexity and the parameter quantity requirements, ensuring the model training efficiency, and reducing the storage and memory occupancy requirements.

[0169] It should be noted that the above order of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. Moreover, the above specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0170] The various embodiments in the present application are all described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0171] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware or can be completed by a program instructing the relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.

[0172] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A method for processing graph data, characterized in that, The method includes: If the data structure of the graph data to be processed is a directed graph, obtaining first matrix data and first graph attribute features of the graph data to be processed; the first matrix data includes an out-degree matrix and an in-degree matrix, the elements of the out-degree matrix are used to represent the outward connection relationship between the corresponding node and other nodes in the graph data to be processed, the elements of the in-degree matrix are used to represent the connection relationship that the corresponding node in the graph data to be processed receives from other nodes, and both the out-degree matrix and the in-degree matrix are symmetric matrices; the first graph attribute features are used to represent the node features of each node in the graph data to be processed; Performing data analysis on the first matrix data and the first graph attribute features based on a target graph neural network to obtain a graph analysis result.

2. The method according to claim 1, wherein The obtaining method of the first matrix data includes: Obtaining the node sorting of each node in the graph data to be processed and the out-degree relationship and in-degree relationship of each node; When the out-degree relationship indicates that there is an out-edge between the corresponding node and another node with a later node sorting, assigning the element composed of the two nodes at both ends of the out-edge to a first value, and when the out-degree relationship indicates that there is an in-edge or no edge between the node and another node with a later node sorting, assigning the element composed of the two nodes at both ends of the in-edge or the element composed of the two nodes with no edge to a second value to obtain the out-degree matrix; When the in-degree relationship indicates that there is an in-edge between the corresponding node and another node with a later node sorting, assigning the element composed of the two nodes at both ends of the in-edge to a first value, and when the in-degree relationship indicates that there is an out-edge or no edge between the node and another node with a later node sorting, assigning the element composed of the two nodes at both ends of the out-edge or the element composed of the two nodes with no edge to a second value to obtain the in-degree matrix; Generating the first matrix data based on the out-degree matrix and the in-degree matrix.

3. The method according to claim 1, characterized in that, The node features include node basic features, and the node basic features are used to represent the attributes of the corresponding nodes and are independent of the directions of the in-degree and out-degree of the corresponding nodes.

4. The method according to claim 1, characterized in that The node features include at least one of an out-degree feature and an in-degree feature, the out-degree feature is associated with the out-edge connection relationship of the corresponding node, and the in-degree feature is associated with the in-edge connection relationship of the corresponding node.

5. The method according to any one of claims 1-4, characterized in that The obtaining method of the first graph attribute features includes: Obtaining the node basic features, out-degree features, and in-degree features of each node in the graph data to be processed; Fusing the node basic features and the out-degree features to obtain the first feature of the node; Fusing the node basic features and the in-degree features to obtain the second feature of the node; Generating the node features of the node based on the first feature and the second feature; Based on the node sorting, splicing the node features of each node to obtain the first graph attribute features.

6. The method according to claim 5, wherein The obtaining of the node basic features, out-degree features, and in-degree features of the nodes in the graph data to be processed includes: Obtain the node basic data, out-degree data, and in-degree data of the nodes in the to-be-processed graph data. The node basic data is data independent of the in-out degree direction of the corresponding node. The out-degree data is data associated with the out-degree connection relationship of the corresponding node. The in-degree data is data associated with the in-degree connection relationship of the corresponding node. Characterize the node basic data, the node data, and the in-degree data respectively to obtain the node basic feature, the out-degree feature, and the in-degree feature.

7. The method according to any one of claims 1-4, characterized in that, The method further includes: If the data structure of the to-be-processed graph data is an undirected graph, obtain the second matrix data and the second graph attribute feature of the to-be-processed graph data. The second matrix data includes an out-degree matrix and an in-degree matrix, and both the out-degree matrix and the in-degree matrix are the adjacency matrix of the to-be-processed graph data. The second graph attribute feature is used to characterize the node features of each node in the to-be-processed graph data. Input the second matrix data and the second graph attribute feature into a target graph neural network for data analysis to obtain a graph analysis result.

8. The method according to claim 7, wherein The obtaining method of the second graph attribute feature includes: Obtain the node basic features of each node in the to-be-processed graph data. Determine the node basic feature as the first feature of the node. Determine the node basic feature as the second feature of the node. Generate the node feature of the node based on the first feature and the second feature. Concatenate the node features of each node based on node sorting to obtain the second graph attribute feature.

9. The method according to any one of claims 1 to 4, characterized in that, The target graph neural network includes a feature network and a classifier. The data analysis of the first matrix data and the first graph attribute feature based on the target graph neural network to obtain a graph analysis result includes: Extract features of the first matrix data and the first graph attribute feature based on the feature network to obtain graph extraction features. Input the graph extraction features into the classifier for classification recognition to obtain the graph analysis result. The graph analysis result is used to characterize the graph category of the to-be-processed graph data and / or to characterize the node categories of each node in the to-be-processed graph data.

10. A graph data processing device, characterized in that, The device includes: An obtaining module: used to obtain the first matrix data and the first graph attribute feature of the to-be-processed graph data if the data structure of the to-be-processed graph data is a directed graph. The first matrix data includes an out-degree matrix and an in-degree matrix. The elements of the out-degree matrix are used to characterize the outward connection relationship between the corresponding node and other nodes in the to-be-processed graph data. The elements of the in-degree matrix are used to characterize the connection relationship that the corresponding node in the to-be-processed graph data receives from other nodes. Both the out-degree matrix and the in-degree matrix are symmetric matrices. The first graph attribute feature is used to characterize the node features of each node in the to-be-processed graph data. An analysis module: used to perform data analysis of the first matrix data and the first graph attribute feature based on a target graph neural network to obtain a graph analysis result.

11. A computer-readable storage medium, characterized in that, At least one instruction or at least one program segment is stored in the storage medium, and the at least one instruction or the at least one program segment is loaded and executed by a processor to implement the graph data processing method according to any one of claims 1-9.

12. A computer device, characterized in that, The device includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the graph data processing method according to any one of claims 1-9.