Graph data representation method and device, equipment and storage medium
By extracting edge features, node features and adjacency matrices from graph structure data, generating connection optimization problems and solving them, updating edge feature matrix using relative distance tensors, and calculating attention characteristics of nodes and edges, the problem that the graph Transformer architecture failed to integrate structural information before calculating attention scores, achieving a more accurate feature expression of graph structure data.
Patent Information
- Application Number
- CN202510006349.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-30
AI Technical Summary
The existing graph Transformer architecture fails to integrate structural information before calculating attention scores, and cannot effectively capture the global structure of the graph, resulting in insufficient feature expression capabilities of graph structure data.
By obtaining edge feature matrix, node feature matrix and adjacency matrix from the graph structure data, generating connection optimization problems and solving them to obtain optimized position parameters, updating edge feature matrix using relative distance tensors, calculating attention characteristics of nodes and edges, and finally obtaining the characterization characteristics of graph data.
The graph Transformer architecture has improved the feature representation ability of graph structure data, and can more accurately evaluate the correlation between nodes and improve the calculation accuracy of attention characteristics.
Smart Images

Figure CN120068926A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of graph data processing, and in particular, to a method, apparatus, device, and storage medium for graph data representation. Background Art
[0002] Graph-structured data plays a crucial role in many application fields such as drug discovery, protein design, and social networks. The graph Transformer architecture constructed for graph data is a neural network architecture based on the self-attention mechanism, which integrates the Graph Neural Network (GNN) and the Transformer architecture. In this architecture, the GNN aims to capture local structures, while the Transformer architecture is used to integrate global entity information.
[0003] In the related art, before calculating the attention scores, the graph Transformer architecture does not integrate structural information, does not consider the structure between any two nodes in the graph, and cannot explicitly encode the structure of the graph, resulting in the model having a weak ability to perceive the global structure and insufficient feature expression ability for graph-structured data. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a method, apparatus, device, and storage medium for graph data representation, so as to improve the feature representation ability of the graph Transformer architecture for graph-structured data.
[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a method for graph data representation, including:
[0006] Obtain an edge feature matrix, a node feature matrix, and an adjacency matrix from the graph-structured data;
[0007] Generate a connection optimization problem according to the graph-structured data, and solve the connection optimization problem to obtain an optimized position parameter;
[0008] Obtain a relative distance tensor corresponding to the graph-structured data according to the adjacency matrix and the decay exponent, and update the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix;
[0009] For any node, obtain the attention distance corresponding to the node based on the hidden edge feature matrix, obtain the attention scores between the node and other nodes according to the attention distance, obtain the single-head attention feature corresponding to the node at least according to the attention scores and the optimized position parameter, and obtain the node attention feature according to the single-head attention feature;
[0010] For any edge, after linear mapping according to the hidden edge feature matrix, an edge attention feature is obtained;
[0011] According to the edge attention feature and the node attention feature, a graph data representation feature is obtained.
[0012] In some embodiments, the generating of the connection optimization problem according to the graph structure data includes:
[0013] For a node pair formed by any two nodes, a first node and a second node are obtained, a node connection label is acquired, and the first node is positionally encoded according to the initial position parameter to obtain a first node position encoding, and the second node is positionally encoded to obtain a second node position encoding, where the initial position parameter includes a first initial position parameter and a second initial position parameter;
[0014] The first node position encoding is mapped according to the mapping transformation vector, and a mapping result is obtained, where the mapping transformation vector includes a third initial position parameter and a fourth initial position parameter;
[0015] A difference vector between the mapping result and the second node position encoding is acquired, and a target function is obtained based on the norm value of the difference vector and the node connection label, and the connection optimization problem is generated by minimizing the target function.
[0016] In some embodiments, the solving of the connection optimization problem to obtain an optimized position parameter includes:
[0017] The node pair is divided into positive samples and negative samples according to the node connection label;
[0018] Based on the connection optimization problem, a scoring function for each node pair is acquired, and a positive sample scoring function and a negative sample scoring function are respectively obtained;
[0019] An optimized loss function is generated according to the positive sample scoring function and the negative sample scoring function, and the initial parameters are optimized according to the optimized loss function to satisfy the connection optimization problem until the optimized position parameters corresponding to the initial position parameter, the third initial position parameter, and the fourth initial position parameter are obtained, where the optimized position parameters include: a first position parameter, a second position parameter, a third position parameter, and a fourth position parameter.
[0020] In some embodiments, the obtaining of the relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation index includes:
[0021] The degree matrix and the matrix trace of the adjacency matrix are acquired, and a normalized adjacency matrix is obtained according to the degree matrix and the matrix trace;
[0022] Obtain the product result of the attenuation index and the normalized adjacency matrix, subtract the product result from the identity matrix to obtain a difference matrix, and obtain the relative distance tensor according to the inverse matrix of the difference matrix.
[0023] In some embodiments, the updating the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix includes:
[0024] Apply a size adjustment function to the relative distance tensor and then apply a first activation function to obtain an activated distance tensor;
[0025] Multiply the activated distance tensor and the relative distance tensor to obtain a projected distance tensor;
[0026] Obtain the hidden edge feature matrix based on the projected distance tensor and the edge feature matrix.
[0027] In some embodiments, the obtaining the attention distance corresponding to the node based on the hidden edge feature matrix includes:
[0028] Take the node as the target node, and select calculation nodes one by one from other nodes outside the target node;
[0029] Obtain the element value corresponding to the edge between the target node and the calculation node in the hidden edge feature matrix as the attention distance.
[0030] In some embodiments, the optimizing position parameters at least include: a first position parameter and a second position parameter, and the obtaining the attention score between the node and other nodes according to the attention distance includes:
[0031] Based on the optimizing position parameters, obtain an updated target query vector of the target node and an updated calculation key vector of the calculation node, and calculate the vector sum of the updated target query vector and the updated calculation key vector to obtain an initial sum vector;
[0032] Take the products of the first weight matrix and the second weight matrix and the attention score as a first attention value and a second attention value respectively, calculate the inner product of the initial sum vector and the first attention value to obtain a third attention value, and input the sum of the third attention value and the second attention value into a second activation function to output an intermediate attention distance corresponding to the calculation node;
[0033] Obtain the attention score according to the intermediate attention distance, the updated target query vector and the updated calculation key vector.
[0034] In some embodiments, the obtaining the updated target query vector of the target node and the updated calculation key vector of the calculation node includes:
[0035] Obtain the target node feature of the target node and the computing node feature of the computing node from the node feature matrix;
[0036] Obtain the initial target query vector of the target node, and obtain the updated target query vector based on the optimized position parameter, the initial target query vector, and the target node feature;
[0037] Obtain the initial computing key vector of the computing node, and obtain the updated computing key vector based on the optimized position parameter, the initial computing key vector, and the computing node feature.
[0038] In some embodiments, obtaining the updated target query vector based on the optimized position parameter, the initial target query vector, and the target node feature includes:
[0039] Obtain the product of the initial target query vector and the target node feature to obtain a query intermediate value;
[0040] Calculate the target position encoding corresponding to the target node according to the first position parameter and the second position parameter;
[0041] Calculate the inner product of the query intermediate value and the target position encoding to obtain the updated target query vector.
[0042] In some embodiments, obtaining the single-head attention feature corresponding to the node based on at least the attention score and the optimized position parameter includes:
[0043] Obtain the computed value vector of the computing node, use the product of the third weight matrix and the intermediate attention distance as the fourth attention value, obtain the fifth attention value according to the sum of the computed value vector and the fourth attention value, and calculate the product of the attention score and the fifth attention value to obtain the attention feature corresponding to the computing node;
[0044] Accumulate all the attention features to obtain the single-head attention feature of the target node, and take each node as the target node one by one to calculate the single-head attention features corresponding to all nodes.
[0045] In some embodiments, obtaining the node attention feature according to the single-head attention feature includes:
[0046] Obtain the single-head attention features corresponding to each attention head, and splice them in order to obtain the corresponding spliced vector;
[0047] After performing a linear transformation on the splicing vector, normalizing it to obtain a normalization result, and calculating the sum of the normalization result and the node feature matrix, the node attention feature is obtained.
[0048] In some embodiments, obtaining the graph data representation feature according to the edge attention feature and the node attention feature includes:
[0049] Obtain the edge input data and node input data at the current moment, where the edge input data is the edge output data at the previous moment, and the node input data is the node output data at the previous moment;
[0050] Use the edge input data to update the hidden edge feature matrix, and obtain the updated edge attention feature as the edge output data at the current moment;
[0051] Use the node input data to update the node feature matrix, and obtain the updated node attention feature as the first output intermediate value;
[0052] Add the first output intermediate value to the node input data to obtain a second output intermediate value;
[0053] Send the second output intermediate value to a feed-forward neural network for data processing to obtain a feed-forward output result, and add the feed-forward output result and the second output intermediate value to obtain the node output data at the current moment;
[0054] Use the node output data of the current node as the node input data at the next moment, and use the edge output data as the edge input data at the next moment, and perform at least one iteration process until a target node feature matrix and a target edge feature matrix are obtained as the graph data representation feature;
[0055] At the initial moment, the edge input data is the edge attention feature, the node input data is the node attention feature, the node output data at the last moment is the target node feature matrix, and the edge output data is the target edge feature matrix.
[0056] To achieve the above object, a second aspect of the embodiments of the present application proposes a graph data representation device, including:
[0057] A graph data acquisition module: used to acquire an edge feature matrix, a node feature matrix, and an adjacency matrix from graph structure data;
[0058] An optimization solving module: used to generate a connection optimization problem according to the graph structure data, and solve the connection optimization problem to obtain an optimized position parameter;
[0059] Relative distance calculation module: used to obtain the relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation exponent, and update the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix;
[0060] Node attention calculation module: for any node, based on the hidden edge feature matrix, obtain the attention distance corresponding to the node, obtain the attention score between the node and other nodes according to the attention distance, obtain the single-head attention feature corresponding to the node at least according to the attention score and the optimized position parameter, and obtain the node attention feature according to the single-head attention feature;
[0061] Edge attention calculation module: for any edge, perform a linear mapping according to the hidden edge feature matrix to obtain an edge attention feature;
[0062] Data representation module: used to obtain the graph data representation feature according to the edge attention feature and the node attention feature.
[0063] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0064] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a storage medium, the storage medium is a storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0065] The graph data representation method, device, equipment, and storage medium proposed in the embodiments of this application obtain an edge feature matrix, a node feature matrix, and an adjacency matrix from graph structure data, generate a connection optimization problem based on the graph structure data, solve the connection optimization problem to obtain optimized position parameters, obtain a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation exponent, update the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix, for any node, obtain the attention distance corresponding to the node based on the hidden edge feature matrix, obtain an attention score according to the attention distance, obtain the single-head attention feature corresponding to the node at least according to the attention score and the optimized position parameters, obtain the node attention feature according to the single-head attention feature, for any edge, obtain the edge attention feature after linear mapping according to the hidden edge feature matrix, and finally obtain the graph data representation feature according to the edge attention feature and the node attention feature. The embodiments of this application first generate a relative distance tensor based on the adjacency matrix corresponding to the graph structure data and the attenuation exponent. Among them, the attenuation exponent is used to quantify the range of the perception field. This relative distance tensor not only considers the direct connections between nodes, but also comprehensively considers the indirect connections and distance relationships between nodes with the help of the attenuation exponent, providing richer and more comprehensive global structure information for the model in the subsequent feature expression process, thereby enhancing the model's understanding and expression ability of graph data. In addition, when calculating the attention feature, by combining the connection relationships between nodes, the model can more accurately evaluate the correlation between nodes and avoid only considering the local features of nodes when calculating the attention. In this way, the attention weights can be more accurately allocated, and the calculation accuracy of the attention feature can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a flowchart of the graph data representation method provided by the embodiments of this application.
[0067] Figure 2 is an overall flowchart of the graph data representation method provided by the embodiments of this application.
[0068] Figure 3 is a flowchart of generating a connection optimization problem according to the graph structure data provided by the embodiments of this application.
[0069] Figure 4 is a flowchart of solving the connection optimization problem to obtain optimized position parameters provided by the embodiments of this application.
[0070] Figure 5 is a flowchart of obtaining a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation exponent provided by the embodiments of this application.
[0071] Figure 6 is a schematic diagram of node connections provided by the embodiments of this application.
[0072] Figure 7 It is a flowchart for updating an edge feature matrix according to a relative distance tensor to obtain a hidden edge feature matrix provided by an embodiment of the present application.
[0073] Figure 8 It is a flowchart for obtaining an attention score between a node and other nodes according to an attention distance provided by an embodiment of the present application.
[0074] Figure 9 It is a flowchart for obtaining an updated target query vector of a target node and calculating an updated calculation key vector of a node provided by an embodiment of the present application.
[0075] Figure 10 It is a flowchart for obtaining an updated target query vector according to an initial target query vector and a target node feature provided by an embodiment of the present application.
[0076] Figure 11 It is a schematic diagram of an attention calculation process provided by an embodiment of the present application.
[0077] Figure 12 It is a flowchart for obtaining a single-head attention feature corresponding to a node according to at least an attention score and an optimized position parameter provided by an embodiment of the present application.
[0078] Figure 13 It is a flowchart for obtaining a node attention feature according to a single-head attention feature provided by an embodiment of the present application.
[0079] Figure 14 It is a flowchart for obtaining a graph data representation feature according to an edge attention feature and a node attention feature provided by an embodiment of the present application.
[0080] Figure 15 It is a schematic diagram of verification data for an embodiment of the present application.
[0081] Figure 16 It is a structural block diagram of a graph data representation device provided by another embodiment of the present application.
[0082] Figure 17 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0083] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0084] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown or described can be executed in a different module division from that in the device or a different order from that in the flowchart.
[0085] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0086] First, several nouns involved in this application are analyzed:
[0087] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information processes of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.
[0088] Graph-structured data plays a crucial role in many application fields such as drug discovery, protein design, and social networks. In the field of drug discovery, drug molecules are represented as graph structures. Among them, nodes can represent atoms, and edges represent chemical bonds between atoms, thereby enabling a more intuitive analysis of the structural characteristics of drug molecules. In the field of protein design, graph-structured data can be used to describe the relationship between the three-dimensional structure of proteins and amino acid sequences. By constructing a graph model of the protein structure, with nodes representing amino acids and edges representing the spatial distance or interaction between amino acids, the folding rules and functional mechanisms of proteins can be deeply analyzed. In the field of social networks, the nodes of graph-structured data represent users, and the edges represent the relationships between users, such as friendship relationships, following relationships, etc. By analyzing the graph-structured data of social networks, the behavior patterns, interest preferences, and information dissemination rules of users can be deeply understood.
[0089] Among them, the Transformer architecture, as a neural network architecture based on the self-attention mechanism, is widely used in sequence data modeling. On this basis, the development of the Transformer architecture for graph data has sparked great interest.
[0090] The graph Transformer architecture constructed for graph data is a neural network architecture based on the self-attention mechanism. The graph Transformer architecture mainly has two branches. One branch is to fuse the Graph Neural Network (GNN) with the Transformer architecture. In this architecture, the GNN aims to capture local structures, while the Transformer architecture is used to integrate global entity information. The other branch is to use the statistical features of the graph as positional encoding and embed the structural information in the graph into the Transformer architecture, such as node degree and Laplacian matrix.
[0091] In the related art, before calculating the attention scores, neither of the two branches of the graph Transformer architecture integrates the structural information, does not consider the structure between any two nodes in the graph, and cannot explicitly encode the structure of the graph, resulting in the model having a weak ability to perceive the global structure and insufficient feature expression ability for graph structure data.
[0092] Based on this, the embodiments of the present application provide a graph data representation method, apparatus, device, and storage medium. First, a relative distance tensor is generated according to the adjacency matrix and the attenuation exponent corresponding to the graph structure data. The attenuation exponent is used to quantify the range of the perception field. The relative distance tensor not only considers the direct connections between nodes but also comprehensively considers the indirect connections and distance relationships between nodes with the help of the attenuation exponent, providing richer and more comprehensive global structure information for the model in the subsequent feature expression process, thereby enhancing the model's understanding and expression ability for graph data. In addition, when calculating the attention features, by combining the connection relationships between nodes, the model can more accurately evaluate the correlation between nodes and avoid only considering the local features of nodes when calculating the attention. In this way, the attention weights can be more precisely allocated, and the calculation accuracy of the attention features can be improved.
[0093] The embodiments of the present application provide a graph data representation method, apparatus, device, and storage medium, which will be specifically described through the following embodiments. First, the graph data representation method in the embodiments of the present application will be described.
[0094] Embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results of theory, methods, techniques, and application systems. In other words, artificial intelligence is a comprehensive technology in computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0095] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0096] The graph data representation method provided by the embodiments of this application relates to the field of graph data processing technology. The graph data representation method provided by the embodiments of this application can be applied to a terminal, or to a server, or can be a computer program running on a terminal or a server. For example, the computer program can be a native program or software module in an operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in an operating system to run, such as a client that supports graph data representation, that is, a program that only needs to be downloaded to a browser environment to run; it can also be a small program that can be embedded in any APP. In short, the above computer program can be any form of application program, module, or plug-in. Among them, the terminal communicates with the server through a network. This graph data representation method can be executed by the terminal or the server, or jointly executed by the terminal and the server.
[0097] In some embodiments, the terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc. In addition, the terminal may also be an intelligent vehicle-mounted device. The intelligent vehicle-mounted device applies the graph data representation method of this embodiment to provide relevant services and enhance the driving experience. The server may be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; it may also be a service node in a blockchain system, and the service nodes in the blockchain system form a Peer To Peer (P2P) network, and the P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). The terminal and the server can be connected through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not limit this.
[0098] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0099] The graph data representation method in the embodiments of this application is described below.
[0100] Figure 1 is an optional flowchart of the graph data representation method provided by the embodiments of this application, Figure 1 The method in may include but is not limited to steps 110 to 160. At the same time, it can be understood that this embodiment does not specifically limit the Figure 1 order of steps 110 to 160 in, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.
[0101] Step 110: Obtain an edge feature matrix, a node feature matrix, and an adjacency matrix from the graph structure data.
[0102] In one embodiment, referring to Figure 2 , Figure 2 is the overall flowchart of the graph data representation method provided by the embodiments of the present application. First, for a graph structure data, it can be determined by an edge representation ε and a node representation v, denoted as where the node representation is each independent element included in the graph structure data, and these elements are the basic units constituting the graph. Each node can carry specific information. For example, in a social network graph, a node may represent a user, and each user node may contain basic information of the user, such as name, age, gender, etc. In a biomolecular graph, a node may represent an atom, carrying attributes such as the type of atom and charge. The edge representation is the connection relationship between different nodes, and the nature and weight of the edge can vary according to specific application scenarios and problems. For example, in a social network, an edge can represent a friendship relationship, a following relationship, or a common interest between users. Through the combination of nodes and edges, the graph structure data can not only display the characteristics of individual elements (nodes) but also clearly present the mutual connections (edges) between the elements.
[0103] In one embodiment, referring to Figure 2 , the edge representation is encoded into hidden features using an edge feature encoder to obtain an edge feature matrix The node representation is encoded into hidden features using a node feature encoder to obtain a node feature matrix X e . Among them, n represents the number of nodes, d model represents the hidden dimension, and the edge feature encoder f edge and the node feature encoder f node are both linear encoders.
[0104] In one embodiment, the graph structure data also includes an adjacency matrix. The edge information describes the connection relationship between different nodes in the graph structure data. For example, in a simple undirected graph, the edge information may only indicate that there is a connection between nodes; while in a weighted directed graph, the edge information not only needs to indicate that there is an edge from one node to another node but also needs to give the weight value of the edge to represent a certain strength of the connection. Based on this edge information, the adjacency matrix A of the graph structure data can be generated. The adjacency matrix A is a two-dimensional matrix, where the rows and columns of the matrix respectively correspond to the nodes in the graph.
[0105] For an undirected graph, if there is an edge between node i and node j, then A ij = 1; if there is no edge, then A ij= 0. Since the edges in an undirected graph have no direction, that is, the connection between node i and node j is equivalent to the connection between node j and node i, the adjacency matrix of an undirected graph is a symmetric matrix, that is, A ij = A ji . For a directed graph, if there is an edge between node i and node j, then A ij = 1; if there is no edge, then A ij = 0. The adjacency matrix of a directed graph is not necessarily a symmetric matrix because the connection between nodes does not mean that there is also an edge between the nodes. Thus, it can be seen that this adjacency matrix generated based on edge information can represent the connection relationship between nodes in graph structure data.
[0106] Step 120: Generate a connection optimization problem according to the graph structure data, and solve the connection optimization problem to obtain the optimized position parameters.
[0107] In one embodiment, referring to Figure 3 , Figure 3 is the flowchart of generating a connection optimization problem according to the graph structure data provided by the embodiments of the present application, which specifically includes the following steps:
[0108] Step 310: For any node pair composed of two nodes, take it as the first node and the second node, obtain the node connection label, and perform position encoding on the first node according to the initial position parameters to obtain the first node position encoding, and perform position encoding on the second node to obtain the second node position encoding.
[0109] In one embodiment, the Attention mechanism in the Transformer framework essentially calculates the attention weights of each token in the input sequence with respect to the entire sequence. With the help of this mechanism, the feature expression ability can be improved. However, for any two word vectors, regardless of their position changes, the attention weights between them will not change, which indicates that the attention calculation result is independent of the position of the word vectors. Therefore, by using Rotary Position Embedding (RoPE), position indices can be added to the input sequence, and then position encoding information can be introduced into the attention mechanism. However, this kind of rotary position encoding is only applicable to chain sequences and cannot be applied to graph structure data. Based on this, the embodiments of the present application have improved the position encoding.
[0110] During the process of rotary position encoding, if the rotation angle difference between two adjacent nodes n and n - 1 is θ, at this time, the input sequence is equivalently reformulated as a chain graph, then the position embedding p n-1 of node n - 1 can be realized by multiplying the position encoding p n of node n by the rotation transformation r, that is, e inθ ·e -iθ - e i(n-1)θ= 0, where e inθ ·e -iθ is equivalent to rotating the position encoding p of node n n by nθ first and then rotating it back by θ, and the result is equal to the position encoding p of node n - 1 n-1 Rotating by (n - 1)θ.
[0111] From this, it can be seen that if the chain graph is to be extended to a general graph structure, at this time, if there is an edge connection from node m to node n in the graph structure, the position encoding p of node m m and the position encoding p of node n n have the following relationship: p m ·r - p n ≈ 0, where r represents the mapping transformation vector.
[0112] In one embodiment, in order to capture all connections in the graph structure data, a node pair (u, v) composed of any two nodes is selected to obtain a first node and a second node, where the first node can be node u and the second node can be node v. The node connection label of this node pair is Y u,v , where Y u,v ∈ {1, -1}. If there is a connection between node u and node v, then Y u,v = 1. If there is no connection between node u and node v, then Y u,v = -1.
[0113] Next, position encoding is performed, and the position encoding can also be understood as position embedding. Taking the first node as node u as an example, the first node is position-encoded according to the initial position parameters to obtain the first node position encoding p u , expressed as:
[0114]
[0115] where d = d model is the hidden dimension, ω i refers to the i-th index of ω, where i in is an imaginary number, represents the first initial position parameter θ corresponding to node u,
[0116] Similarly, the second node position encoding p v is obtained and expressed as:
[0117]
[0118] where represents the first initial position parameter θ corresponding to node v, Represents the second initial position parameter corresponding to node v
[0119] It can be understood that the initial position parameter includes the first initial position parameter and the second initial position parameter, and both of these parameters are learnable parameters.
[0120] Step 320: Map the first node position encoding according to the mapping transformation vector to obtain a mapping result.
[0121] In one embodiment, the mapping transformation vector is expressed as:
[0122]
[0123] Among them, Represents the third initial position parameter, Represents the fourth initial position parameter.
[0124] Therefore, the mapping result is expressed as:
[0125] f(p u ,r)
[0126] Among them, f represents the mapping function, which is an abstract function.
[0127] Step 330: Obtain the difference vector between the mapping result and the second node position encoding, and obtain an objective function based on the norm value of the difference vector and the node connection label. Minimize the objective function to generate a connection optimization problem.
[0128] In one embodiment, the difference vector is expressed as:
[0129] f(p u ,r)-p v
[0130] The objective function is expressed as:
[0131] Y u,v ·||f(p u ,r)-p v || 2
[0132] Therefore, the connection optimization problem can be described as follows:
[0133]
[0134] Among them, Represents an unobserved connection, Represents the set of all nodes. That is to say, ε′ represents the set of nodes without connection relationships, and ε represents the set of nodes with connection relationships.
[0135] As described above, the goal of the connection optimization problem in the embodiments of this application is to find the most suitable positions for embedding p and the corresponding mapping transformation vectors r for each node, such that when Y u,v = 1, for all the observed connections in the graph, the mapping results are as close as possible to the position embeddings, while for the unobserved connections, their mapping results should be as far away from the position embeddings as possible. This is because Y u,v = -1. By minimizing this objective function, the position embeddings and mapping transformations of each node can be learned, thereby better capturing the connection structure in the graph.
[0136] To solve this optimization problem, the embodiments of this application propose a method for learning the position embedding p and the corresponding mapping transformation vector r of each node based on contrastive learning for spatial absolute position embedding.
[0137] In one embodiment, referring to Figure 4 , Figure 4 is the flowchart for solving the connection optimization problem to obtain the optimized position parameters provided by the embodiments of this application, which specifically includes the following steps:
[0138] Step 410: Divide the node pairs into positive samples and negative samples according to the node connection labels.
[0139] In one embodiment, if there is a connection relationship between two nodes, then this node pair is a positive sample; otherwise, it is a negative sample. Hereinafter, the node pair (u, v) is used to represent a positive sample, i.e., (u, v) ∈ ε, and the node pair is used to represent a negative sample, i.e., (l, m) ∈ ε'.
[0140] Step 420: Based on the connection optimization problem, obtain the score functions for each node pair, and respectively obtain the positive sample score function and the negative sample score function.
[0141] In one embodiment, taking the given edge (u, v) ∈ ε as an example, its score function g(·, ·) is defined as:
[0142]
[0143] where ξ is a margin variable to avoid overfitting.
[0144] From the above formula, θ u + δθ - θ v represents the difference between node u and node v on the θ-related vector, represents the difference between node u and node v on the -related vector. Next, the scalar value obtained by calculating the Euclidean norm represents the distance between the node position encodings corresponding to node u and node v. Finally, subtracting this distance from the margin variable ξ can obtain the score function g ξ(u, v). The purpose of the scoring function is to measure the closeness of the relationship between node u and node v. If nodes u and v are relatively close in the position encoding, the corresponding scoring function is larger. If nodes u and v are quite different in the position encoding, then the scoring function is smaller.
[0145] According to the above formula, the positive sample scoring function is expressed as:
[0146]
[0147] The negative sample scoring function is expressed as:
[0148]
[0149] Step 430: Generate an optimized loss function based on the positive sample scoring function and the negative sample scoring function, and optimize the initial parameters according to the optimized loss function to satisfy the connection optimization problem until the optimized position parameters corresponding to the initial position parameter, the third initial position parameter, and the fourth initial position parameter are obtained.
[0150] In one embodiment, the optimized loss function L is expressed as:
[0151]
[0152] where σ represents the Sigmoid activation function.
[0153] It can be seen that the purpose of the optimized loss function L is to maximize the value of the positive sample scoring function, and at the same time, minimize the value of the negative sample scoring function. If the positive sample scoring function is larger, σ(g ξ (u, v)) is closer to 1, and logσ(g ξ (u, v)) is closer to zero. And the smaller the negative sample scoring function, -g ξ (l, m) is larger, σ(-g ξ (l, m)) is closer to 1, and logσ(-g ξ (l, m)) is closer to zero. Then, it is optimized by minimizing the logarithmic likelihood loss through a negative sign, enabling the model to correctly distinguish positive samples (node pairs with edge connections) and negative samples (node pairs without edge connections) in the graph structure data.
[0154] Since the initial position parameter, the third initial position parameter, and the fourth initial position parameter are all learnable parameters, after the above optimization process, the corresponding optimized position parameters are obtained. Among them, the optimized position parameters include: the first position parameter θ * 、the second position parameter the third position parameter δθ * and the fourth position parameter are expressed as:
[0155]
[0156] In the embodiment of the present application, the process of calculating the optimized position parameter mentioned above is called Spatial Absolute Position Encoding (SAPE). That is to say, if it is constructed into a module, the output of the Spatial Absolute Position Encoding module at this time is the optimized position parameter.
[0157] Step 130: Obtain the relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation exponent, and update the edge feature matrix according to the relative distance tensor to obtain the hidden edge feature matrix.
[0158] In one embodiment, combined with Figure 2 , next, calculate the hidden edge feature matrix according to the adjacency matrix and the edge feature matrix corresponding to the edge representation. Refer to Figure 5 , Figure 5 is the flowchart for obtaining the relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation exponent provided by the embodiment of the present application, which specifically includes the following steps:
[0159] Step 510: Obtain the degree matrix and the matrix trace of the adjacency matrix, and obtain the normalized adjacency matrix according to the degree matrix and the matrix trace.
[0160] In one embodiment, the degree matrix D is expressed as: D = diag(A), the matrix trace is expressed as TrA, and the normalized adjacency matrix is expressed as:
[0161]
[0162] Step 520: Obtain the product result of the attenuation exponent and the normalized adjacency matrix, use the identity matrix to subtract the product result to obtain the difference matrix, and obtain the relative distance tensor according to the inverse matrix of the difference matrix.
[0163] In one embodiment, the difference matrix is expressed as:
[0164]
[0165] where, I n represents the identity matrix, γ represents the attenuation exponent, and the attenuation exponent is a learnable parameter.
[0166] Therefore, the relative distance tensor D A (γ) is expressed as:
[0167]
[0168] where, K represents the number of hops between two nodes. When K approaches infinity, converges to It can be regarded as the path weight from one node to another through i hops. When i = 0, it means that the distance of the node itself is 0. When i = 1, it means the distance between nodes through one step (direct connection). When i = 2, it means the distance between nodes through two steps (passing through one node in the middle), and so on. Since γ ∈ (0, 1), as i increases, the value of will gradually decrease, thus ensuring the convergence of the infinite series. Finally, the sum of this infinite series can be represented by the matrix to represent.
[0169] In one embodiment, the relative distance tensor can represent the relative distances between different nodes in the graph structure data. Referring to Figure 6 , Figure 6 is the node connection schematic diagram provided by the embodiment of the present application. Taking the calculation of the relative distance from node u to node 1 as an example, there are two paths at this time. One path is: u -> 2 -> 1, and the other path is: u -> 1. Assuming that each hop is multiplied by a γ, so the relative distance corresponding to node u to node 1 is γ + γ 2 . It can be seen that the embodiment of the present application uses the attenuation exponent γ to represent the relative distances between any two nodes. When γ is small, the root node will focus on nearby nodes, and when γ approaches 1, it will consider information from farther nodes. This process in this embodiment is called the spatial residence mechanism.
[0170] In one embodiment, in order to enable the relevant models represented by the graph data to have spatial preservation within a multi-scale range, the embodiment of the present application introduces a set of attenuation exponents to adaptively control the range of the model and achieve a multi-scale spatial range. At this time, the values of the attenuation exponents can have multiple different values, and the specific attenuation exponents can be determined according to the actual learning process, that is However, the embodiment of the present application notices that during the optimization process, the Frobenius norm of D A (γ i ) will become larger as γ → 1. Therefore, considering numerical stability, it is necessary to adjust the size of the relative distance tensor, and use the Rescale module shown in Figure 2 to adjust D A (γ i ) according to the given γ.
[0171] In one embodiment, referring to Figure 7 , Figure 7 is the flowchart for updating the edge feature matrix according to the relative distance tensor to obtain the hidden edge feature matrix provided by the embodiment of the present application, which specifically includes the following steps:
[0172] Step 710: After applying the size adjustment function to the relative distance tensor, apply the first activation function to obtain the activation distance tensor.
[0173] In one embodiment, considering numerical stability, referring to Figure 2 , after obtaining the relative distance tensor D A (γ), it can also be input into the Rescale module for size adjustment, and the size adjustment function is applied to the relative distance tensor.
[0174] Next, the first activation function can be the Sigmoid activation function. Input the result after size adjustment into the Sigmoid activation function to calculate the activation distance tensor, which belongs to the normalization process.
[0175] Step 720: Multiply the activation distance tensor and the relative distance tensor to obtain the projection distance tensor.
[0176] In one embodiment, the projection distance tensor is expressed as:
[0177]
[0178] Among them, represents the projection distance tensor, (σ。f rescale )(γ) represents the activation distance tensor, ° is the function composition operator, indicating that the Rescale function is applied first, and then the Sigmoid function is applied. is a linear mapping function, which can be set according to the actual scenario, and σ is the Sigmoid function.
[0179] Step 730: Obtain the hidden edge feature matrix based on the projection distance tensor and the edge feature matrix.
[0180] In one embodiment, first project the projection distance tensor into the hidden space corresponding to the edge feature matrix, and then add it to the edge feature matrix X e to obtain the hidden edge feature matrix, which is expressed as:
[0181]
[0182] Among them, X e ′ represents the hidden edge feature matrix. is a learnable parameter, which can be pre-trained. This embodiment refers to this process as multi-scale spatial retention relative position encoding (MSSR).
[0183] In one embodiment, referring to Figure 2 , the above process obtains the node feature matrix X and the edge feature matrix X e' and the optimized position parameters, the execution process of the following steps is calculated by L stacked GRN Transformer layers. Among them, each GRN Transformer layer includes a spatial retention head SpRet and a feed-forward neural network FFN of the Transformer. Different GRN Transformer layers are cascaded, and the input data of the latter GRN Transformer layer is the output data of the previous GRN Transformer layer. Assuming that different moments correspond to different GRN Transformer layers, at the initial moment, that is, the input of the first GRN Transformer layer is the node feature matrix X, the edge feature matrix X e ' and the optimized position parameters, through calculation, perform attention calculation on the node feature matrix X to obtain node attention features, and perform attention calculation on the edge feature matrix X e ' to obtain edge attention features. Next, send the node attention features, edge attention features, and optimized position parameters into the next GRN Transformer layer for similar operations, and iterate multiple times until the last GRN Transformer layer outputs node attention features and edge attention features. Take the last node attention feature as the target node feature matrix and the last edge attention feature as the target edge feature matrix.
[0184] Step 140: For any node, obtain the attention distance corresponding to the node based on the hidden edge feature matrix, obtain the attention scores between the node and other nodes according to the attention distance, obtain the single-head attention feature corresponding to the node at least according to the attention scores and the optimized position parameters, and obtain the node attention feature according to the single-head attention feature.
[0185] In one embodiment, it is necessary to calculate the node attention feature corresponding to each node. First, obtain the attention distance corresponding to the node based on the hidden edge feature matrix, and the specific description is as follows: Take the node as the target node, and select and calculate nodes one by one from other nodes outside the target node. Then, obtain the element value corresponding to the edge between the target node and the calculated node in the hidden edge feature matrix as the attention distance.
[0186] Among them, select a node as the target node, and then select and calculate nodes one by one from other nodes outside the target node. Take node u as the target node and node v as the calculated node as an example for illustration. For the hidden edge feature matrix X e ', the element value in its u-th row and v-th column is the attention distance e corresponding to the edge between node u and node v u,v .
[0187] Next, obtain the attention scores corresponding to the target node and each calculated node. Refer to Figure 8 ,Figure 8 This is a flowchart for obtaining the attention score between a node and other nodes based on the attention distance provided by an embodiment of the present application, which specifically includes the following steps:
[0188] Step 810: Based on the optimized position parameter, obtain the updated target query vector of the target node and the updated calculation key vector of the calculation node, and calculate the vector sum of the updated target query vector and the updated calculation key vector to obtain the initial sum vector.
[0189] In one embodiment, referring to Figure 9 , Figure 9 This is a flowchart for obtaining the updated target query vector of the target node and the updated calculation key vector of the calculation node provided by an embodiment of the present application, which specifically includes the following steps:
[0190] Step 910: Obtain the target node feature of the target node and the calculation node feature of the calculation node from the node feature matrix.
[0191] In one embodiment, the target node feature is represented as x u , and the calculation node feature is represented as x v .
[0192] Step 920: Obtain the initial target query vector of the target node, and obtain the updated target query vector based on the optimized position parameter, the initial target query vector, and the target node feature.
[0193] In one embodiment, referring to Figure 10 , Figure 10 This is a flowchart for obtaining the updated target query vector based on the initial target query vector and the target node feature provided by an embodiment of the present application, which specifically includes the following steps:
[0194] Step 1010: Obtain the product of the initial target query vector and the target node feature to obtain the query intermediate value.
[0195] In one embodiment, according to the attention mechanism in the Transform framework, the initial target query vector of the target node is represented as W Q , and at this time, the query intermediate value is represented as: (W Q x u ).
[0196] Step 1020: Calculate the target position encoding corresponding to the target node according to the first position parameter and the second position parameter.
[0197] In one embodiment, the target position encoding is represented as:
[0198]
[0199] Wherein, Represents the first position parameter corresponding to node u, Represents the second position parameter corresponding to node u.
[0200] Step 1030: Calculate the inner product of the query intermediate value and the target position encoding to obtain the updated target query vector.
[0201] In one embodiment, the updated target query vector is represented as:
[0202]
[0203] where e represents the inner product operation, Q u = f Q (x u ) represents the updated target query vector.
[0204] Step 930: Obtain the initial computational key vector of the computational node, and obtain the updated computational key vector based on the optimized position parameter, the initial computational key vector, and the computational node feature.
[0205] In one embodiment, the updated computational key vector obtained according to the above similar calculation process is represented as:
[0206]
[0207] where W K represents the initial computational key vector, represents the first position parameter corresponding to node v, represents the second position parameter corresponding to node v, K v = f K (x v ) represents the updated computational key vector.
[0208] In one embodiment, referring to Figure 11 , Figure 11 is a schematic diagram of the attention calculation process provided by the embodiment of the present application. After obtaining the updated target query vector and the updated computational key vector, it is also necessary to calculate the vector sum of the updated target query vector and the updated computational key vector using the summation operator to obtain the initial sum vector. Therefore, the initial vector sum is represented as (Q u + K v ).
[0209] Step 820: Use the products of the first weight matrix and the second weight matrix and the attention score as the first attention value and the second attention value, calculate the inner product of the initial sum vector and the first attention value to obtain the third attention value, and input the sum of the third attention value and the second attention value into the second activation function to output the intermediate attention distance corresponding to the computational node.
[0210] In one embodiment, the first weight matrix is represented as Wew The first attention value is denoted as W ew e u,v The second weight matrix is denoted as W eb The second attention value is denoted as W eb e u,v Among them, both the first weight matrix and the second weight matrix are learnable parameters, which are obtained through training in advance.
[0211] Therefore, the third attention value is denoted as:
[0212] (Q u +K v )e W ew e u,v
[0213] Assume that the second activation function is the Sigmoid activation function, and the intermediate attention distance is denoted as:
[0214]
[0215] Among them, ρ is a linear function to ensure numerical stability, σ is the Sigmoid activation function, and d h represents the dimension of the h-th attention head. Combining Figure 11 , the attention distance e in the hidden edge feature matrix is introduced using a linear layer u,v , and the intermediate attention distance is calculated using a linear layer
[0216] Step 830: Obtain the attention score according to the intermediate attention distance, the updated target query vector, and the updated calculation key vector.
[0217] In one embodiment, first, an intermediate value of the attention score is obtained according to the intermediate attention distance, the updated target query vector, and the updated calculation key vector. The intermediate value of the attention score is denoted as:
[0218]
[0219] Among them, i is the imaginary unit, and Re represents extracting the real part from the complex number result to enable subsequent calculations. Here, f Q (x u ) T f K (x v ) is calculated to characterize the distance information between the position encodings corresponding to nodes u and v. Since the position encoding is represented by an exponent, the conjugate of the updated calculation key vector is taken so that the exponent part can perform a difference operation in a computable manner.
[0220] Next, the softmax function is used to calculate the intermediate value of the attention scores to obtain the attention scores, which are expressed as:
[0221]
[0222] where α u,v represents the attention scores, and W A represents the learnable weights. Combining Figure 11 , the aggregation layer is used to aggregate according to the intermediate value of the attention scores to obtain the attention scores.
[0223] According to the above process, the attention scores between node u and each computing node are obtained.
[0224] Next, referring to Figure 12 , Figure 12 is a flowchart for obtaining the single-head attention features corresponding to nodes at least according to the attention scores and the optimized position parameters provided by the embodiments of the present application, which specifically includes the following steps:
[0225] Step 1210: Obtain the computed value vector of the computing node, take the product of the third weight matrix and the intermediate attention distance as the fourth attention value, obtain the fifth attention value according to the sum of the computed value vector and the fourth attention value, and calculate the product of the attention scores and the fifth attention value to obtain the attention features corresponding to the computing node.
[0226] In one embodiment, the computed value vector of computing node v is V v , the third weight matrix is expressed as W ev , this matrix is a learnable parameter, and the fourth attention value is expressed as: The fifth attention value is expressed as Therefore, combining Figure 11 , the attention features obtained by matrix multiplication are expressed as:
[0227]
[0228] Step 1220: Accumulate all the attention features to obtain the single-head attention features of the target node, and take each node as the target node one by one to calculate the single-head attention features corresponding to all nodes.
[0229] In one embodiment, when node u is used as the target node, its single-head attention features are expressed as:
[0230]
[0231] From the calculation process, it can be seen that the single-head attention features are related to the node feature matrix X, the edge feature matrix X e ′, the first position parameter θ * and the second position parameter is related, and the first position parameter θ * and the second position parameter can indicate the position encoding related information of node u.
[0232] In one embodiment, the process of calculating the single-head attention feature is implemented by using the self-attention module, which is the spatial retention head SpRet mentioned above. According to the above process, each node is used as the target node by using the spatial retention head, and the single-head attention feature corresponding to each node is obtained.
[0233] Next, referring to Figure 13 , Figure 13 is the flowchart for obtaining the node attention feature according to the single-head attention feature provided by the embodiment of the present application, which specifically includes the following steps:
[0234] Step 1310: Obtain the single-head attention feature corresponding to each attention head, and splice them in order to obtain the corresponding spliced vector.
[0235] In one embodiment, referring to Figure 11 , there are h attention heads in total. Combining the previous calculation process, each attention head can calculate the corresponding single-head attention feature, denoted as {head 1 ,..., head h}. At this time, the spliced vector is denoted as: Concat(head 1 ,..., head h ).
[0236] Step 1320: After linearly transforming the spliced vector, normalize it to obtain the normalization result, and calculate the sum of the normalization result and the node feature matrix to obtain the node attention feature.
[0237] In one embodiment, use the learnable weight W o to linearly transform the spliced vector, and use BatchNorm to normalize it to obtain the normalization result. Calculate the sum of the normalization result and the node feature matrix to obtain the node attention feature, denoted as:
[0238]
[0239] where represents the node attention feature.
[0240] It can be understood that the node attention feature can be calculated in this way for each node in the graph structure data.
[0241] Step 150: For any edge, linearly map it according to the hidden edge feature matrix to obtain the edge attention feature.
[0242] In one embodiment, for each edge in the graph-structured data, a learnable matrix is used to linearly map the hidden edge feature matrix X e ′, and then the result is normalized using BatchNorm to obtain the edge attention feature, expressed as:
[0243]
[0244] where represents the edge attention feature. From the calculation process described above, it can be seen that this edge attention feature is related to the node feature matrix X, the edge feature matrix X e ′, and the projection distance tensor is relevant.
[0245] Step 160: Obtain the graph data representation feature based on the edge attention feature and the node attention feature.
[0246] In one embodiment, referring to Figure 14 , Figure 14 is the flowchart for obtaining the graph data representation feature based on the edge attention feature and the node attention feature provided by the embodiments of the present application, specifically including the following steps:
[0247] Step 1410: Obtain the edge input data and the node input data at the current moment.
[0248] In one embodiment, the edge input data is the edge output data at the previous moment, and the node input data is the node output data at the previous moment. Assume that the current moment corresponds to the (l + 1) th th GRN Transformer layer, and the previous moment corresponds to the l th th GRN Transformer layer. At this time, the edge input data at the current moment is the edge output data at the previous moment The node input data at the current moment is the node output data X l . At the initial moment, that is, the first GRN Transformer layer, its edge input data is the edge attention feature calculated previously, and the node input data is the node attention feature calculated previously.
[0249] Step 1420: Update the hidden edge feature matrix using the edge input data to obtain the updated edge attention feature as the edge output data at the current moment.
[0250] In one embodiment, using the edge input data as the edge feature matrix, according to the previous calculation method, the hidden edge feature matrix is recalculated, and then the edge attention feature is updated and calculated based on the hidden edge feature matrix. The updated edge attention feature is used as the edge output data at the current moment, expressed as:
[0251]
[0252] Among them, represents the edge output data at the current moment.
[0253] Step 1430: Update the node feature matrix using the node input data to obtain the updated node attention feature as the first output intermediate value.
[0254] In one embodiment, using the node input data as the node feature matrix, according to the previous calculation method, update the attention distance based on the recalculated hidden edge feature matrix, then calculate the attention score based on the updated attention distance, calculate the single-head attention feature according to the attention score, and finally calculate the updated node attention feature according to the single-head attention feature. The updated node attention feature is used as the first output intermediate value, which is expressed as:
[0255]
[0256] Among them, represents the first output value.
[0257] Step 1440: Add the first output intermediate value to the node input data to obtain the second output intermediate value.
[0258] In one embodiment, the second output intermediate value is expressed as:
[0259]
[0260] Step 1450: Send the second output intermediate value into the feed-forward neural network for data processing to obtain the feed-forward output result, and add the feed-forward output result and the second output intermediate value to obtain the node output data at the current moment.
[0261] In one embodiment, the node output data X at the current moment l+1 is expressed as:
[0262]
[0263] Among them, FFN represents the feed-forward neural network, represents the feed-forward output result.
[0264] Step 1460: Use the node output data of the current node as the node input data for the next moment, and use the edge output data as the edge input data for the next moment, and perform at least one iteration process until the target node feature matrix and the target edge feature matrix are obtained as the graph data representation features.
[0265] In one embodiment, multiple iterative operations are performed. The node output data of the current node is used as the node input data for the next moment, and the edge output data is used as the edge input data for the next moment, until the node output data at the last moment is obtained as the target node feature matrix, and the edge output data is obtained as the target edge feature matrix.
[0266] In one embodiment, with reference to Figure 2 , after obtaining the graph data representation features, the structural features of the graph are encoded into a low-dimensional numerical vector space, and numerical values are used to represent the connection relationships between nodes in the graph. The learned structural features of the graph can be used for downstream tasks. The downstream tasks here can be classification tasks or regression tasks, etc. For example, in the regression learning of the chemical properties of chemical molecules, the possible chemical properties of the molecule can be predicted based on the known molecular structure, particle types, and edge types. This embodiment does not make any limitations in this regard.
[0267] The graph data representation method provided by the embodiments of this application proposes a new graph Transformer architecture, named Graph Retention Network (GRN).
[0268] First, the embodiments of this application note that polynomials can reflect the distance between any two nodes, and thus can capture the global relative distance. For example, γ is used to represent the one-hop distance, and γ 2 is used to represent the two-hop distance. Based on this, the embodiments of this application follow this idea to design a relative distance tensor for the given graph structure data, and then construct a new module for modeling the global structure to achieve Multi-Scale Spatial Retention Relative Position Encoding (MSSR). In addition, the value of γ determines the range of the perception field, because the greater the distance, the greater the power of the polynomial, which in turn leads to stronger weight decay. Subsequently, a set of learnable exponential decay factors are introduced to adaptively adjust different ranges perceived by the model. Finally, the embodiments of this application combine the relative distance tensor with the attention map to inject global structure information.
[0269] Secondly, the embodiment of the present application integrates relative distance information into the inner product of the query and the key, and proposes Spatial Absolute Position Encoding (SAPE). Drawing on the idea of Rotary Position Embedding (RoPE) that models each position in a sequence as a rotation, the embodiment of the present application extends it to graph-structured data. Different from the manually fixed rotation angle in RoPE, when performing position encoding, the embodiment of the present application proposes a new optimization strategy based on contrastive learning to learn the rotation angle, and adds an additional phase term to enhance the expressiveness of the module and learn the optimized position parameters.
[0270] Finally, the embodiment of the present application combines SAPE and the MSSR module with the transformer layer to encode relative distance information and absolute position information, and creates a Spatial Retention Head (SpRet), thereby implementing a brand-new single-head attention mechanism. The final graph retention network contains multiple paradigms and can effectively learn the features of graph data.
[0271] The verification process of the graph data representation method of the embodiment of the present application is described below.
[0272] In one embodiment, refer to Figure 15 , Figure 15 which is a schematic diagram of the verification data of the embodiment of the present application. According to Figure 15 the data, the graph data representation method of the embodiment of the present application can effectively learn and capture the representation of the relative information of nodes in graph-structured data. In addition, the graph data representation method of this embodiment achieves state-of-the-art performance in graph tasks (such as node classification and graph classification / regression).
[0273] Specifically, the embodiments of the present application and similar models in the related art were respectively verified for performance on 12 graph datasets. These 12 datasets include two categories. The first category is seven single-graph node classification datasets, namely Chameleon, Squirrel, Cornell, Texas, Wisconsin, F-Squirrel, and F-Chameleon. "F-Chameleon" and "F-Squirrel" are the filtered Chameleon dataset and the filtered Squirrel dataset respectively, and some duplicate edges are deleted in these two datasets. The second category is other datasets, which are: ZINC~\{zinc_paper}, PATTERN~\{Dwivedi2023BenchmarkingGN}, CLUSTER, MNIST, CIFAR10. Among them, ZINC~\{zinc_paper} is a dataset for multi-graph regression, MNIST and CIFAR10 are graph classification datasets for multi-graphs, and the remaining two are inductive node classification datasets for multi-graphs.
[0274] Specifically, Figure 15 Table 1 in [reference] shows the performance of the classification benchmark test on the single-graph node classification dataset. Shown in Table 1 are the average accuracies (%) ± standard deviations of 10 different data splits. "OOM" indicates out of memory. The best results are shown in bold, and the second-best results are shown in italics.
[0275] Figure 15 Table 2 in [reference] shows the performance on five other datasets. These five benchmarks include a graph regression task, two inductive node classification tasks, and two graph classification tasks. Shown in Table 2 are the mean absolute errors (MAE) accuracies (%) ± standard deviations under different methods (including the models in the related art and the graph retention network GRN in the embodiments of the present application). Among all the indicators, the best results are shown in bold, and the second-best results are shown in italics.
[0276] The technical solution provided by the embodiment of the present application obtains an edge feature matrix, a node feature matrix, and an adjacency matrix from the graph structure data, generates a connection optimization problem based on the graph structure data, solves the connection optimization problem to obtain optimized position parameters, obtains a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation exponent, and updates the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix. For any node, an attention distance corresponding to the node is obtained based on the hidden edge feature matrix, an attention score is obtained according to the attention distance, at least the attention score and the optimized position parameters are used to obtain a single-head attention feature corresponding to the node, and a node attention feature is obtained according to the single-head attention feature. For any edge, an edge attention feature is obtained after linear mapping according to the hidden edge feature matrix, and finally a graph data representation feature is obtained according to the edge attention feature and the node attention feature. The embodiment of the present application first generates a relative distance tensor based on the adjacency matrix corresponding to the graph structure data and the attenuation exponent. Among them, the attenuation exponent is used to quantify the range of the perception field. The relative distance tensor not only considers the direct connection between nodes, but also comprehensively considers the indirect connection and distance relationship between nodes with the help of the attenuation exponent, providing richer and more comprehensive global structure information for the model in the subsequent feature expression process, thereby enhancing the model's understanding and expression ability of graph data. In addition, when calculating the attention feature, by combining the connection relationship between nodes, the model can more accurately evaluate the correlation between nodes, avoiding only considering the local features of nodes when calculating the attention. In this way, the attention weight can be more accurately allocated, and the calculation accuracy of the attention feature can be improved.
[0277] The embodiment of the present application also provides a graph data representation device, which can implement the above graph data representation method. Refer to Figure 16 , the device includes:
[0278] A graph data acquisition module 1610: configured to obtain an edge feature matrix, a node feature matrix, and an adjacency matrix from the graph structure data.
[0279] An optimization solving module 1620: configured to generate a connection optimization problem based on the graph structure data and solve the connection optimization problem to obtain optimized position parameters.
[0280] A relative distance calculation module 1630: configured to obtain a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation exponent, and update the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix.
[0281] Node attention calculation module 1640: For any node, based on the hidden edge feature matrix, obtain the attention distance corresponding to the node, obtain the attention scores between the node and other nodes according to the attention distance, obtain the single-head attention feature corresponding to the node at least according to the attention scores and the optimized position parameters, and obtain the node attention feature according to the single-head attention feature.
[0282] Edge attention calculation module 1650: For any edge, after performing linear mapping according to the hidden edge feature matrix, obtain the edge attention feature.
[0283] Data representation module 1660: For obtaining the graph data representation feature according to the edge attention feature and the node attention feature.
[0284] The specific implementation manner of the graph data representation device in this embodiment is basically the same as that of the above graph data representation method, and will not be elaborated here.
[0285] This application embodiment also provides an electronic device, including: at least one memory; at least one processor; at least one program; the program is stored in the memory, and the processor executes the at least one program to implement the above graph data representation method of this application. This electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0286] Please refer to Figure 17 , Figure 17 which shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0287] The processor 1701 can be implemented in the form of a general - purpose central processing unit (CPU), a microprocessor, an application - specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application; the memory 1702 can be implemented in the form of a read - only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 1702 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1702 and are called by the processor 1701 to execute the graph data representation method of the embodiments of the present application; the input / output interface 1703 is used to implement information input and output; the communication interface 1704 is used to implement communication interaction between this device and other devices, and can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and the bus 1705 is used to transmit information between various components of the device (such as the processor 1701, the memory 1702, the input / output interface 1703, and the communication interface 1704); among them, the processor 1701, the memory 1702, the input / output interface 1703, and the communication interface 1704 are communicatively connected to each other inside the device through the bus 1705.
[0288] Embodiments of the present application also provide a storage medium. The storage medium is a storage medium that stores a computer program, and when the computer program is executed by a processor, it implements the above - mentioned graph data representation method.
[0289] As a non - transitory storage medium, the memory can be used to store non - transitory software programs and non - transitory computer - executable programs. In addition, the memory can include high - speed random access memory, and can also include non - transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non - transitory solid - state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above - mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0290] The graph data representation method, apparatus, device, and storage medium proposed in the embodiments of this application obtain an edge feature matrix, a node feature matrix, and an adjacency matrix from graph structure data, generate a connection optimization problem based on the graph structure data, solve the connection optimization problem to obtain optimized position parameters, obtain a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the attenuation index, and update the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix. For any node, based on the hidden edge feature matrix, obtain the attention distance corresponding to the node, obtain an attention score according to the attention distance, obtain the single-head attention feature corresponding to the node at least according to the attention score and the optimized position parameters, and obtain the node attention feature according to the single-head attention feature. For any edge, obtain the edge attention feature after linear mapping according to the hidden edge feature matrix, and finally obtain the graph data representation feature according to the edge attention feature and the node attention feature. The embodiments described in the embodiments of this application are to more clearly illustrate the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0291] The embodiments described in the embodiments of this application are for more clearly explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0292] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than those shown in the figures, or combine some steps, or different steps.
[0293] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0294] Those of ordinary skill in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof.
[0295] As used in the specification of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0296] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0297] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned unit division is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical, or other forms.
[0298] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0299] In addition, each functional unit in various embodiments of the present application may be integrated into a processing unit, may exist separately as individual physical units, or two or more units may be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0300] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0301] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. This does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.
Claims
1. A graph data representation method, characterized in that: include: Obtain edge feature matrix, node feature matrix and adjacency matrix from graph structure data; Generate a connection optimization problem according to the graph structure data, and solve the connection optimization problem to obtain optimized position parameters; Obtaining a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the decay index, and updating the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix; For any node, obtaining the attention distance corresponding to the node based on the hidden edge feature matrix, obtaining the attention score between the node and other nodes according to the attention distance, obtaining the single-head attention feature corresponding to the node according to at least the attention score and the optimized position parameter, and obtaining the node attention feature according to the single-head attention feature; For any edge, linear mapping is performed according to the hidden edge feature matrix to obtain the edge attention feature; The graph data representation feature is obtained according to the edge attention feature and the node attention feature.
2. The graph data characterization method according to claim 1, characterized in that: The generating of connection optimization problems according to the graph structure data includes: For a node pair consisting of any two nodes, a first node and a second node are obtained, a node connection label is obtained, and a position encoding is performed on the first node according to an initial position parameter to obtain a first node position encoding, and a position encoding is performed on the second node to obtain a second node position encoding, wherein the initial position parameter includes a first initial position parameter and a second initial position parameter; Mapping the first node position code according to a mapping transformation vector to obtain a mapping result, wherein the mapping transformation vector includes a third initial position parameter and a fourth initial position parameter; A difference vector between the mapping result and the second node position code is obtained, an objective function is obtained based on a norm value of the difference vector and the node connection label, and the connection optimization problem is generated by minimizing the objective function.
3. The graph data characterization method according to claim 2, characterized in that: The step of solving the connection optimization problem to obtain optimized position parameters includes: Dividing the node pairs into positive samples and negative samples according to the node connection labels; Based on the connection optimization problem, a score function of each node pair is obtained to obtain a positive sample score function and a negative sample score function respectively; An optimized loss function is generated according to the positive sample score function and the negative sample score function, and the initial parameters are optimized according to the optimized loss function to meet the connection optimization problem until the optimized position parameters corresponding to the initial position parameters, the third initial position parameters and the fourth initial position parameters are obtained, and the optimized position parameters include: a first position parameter, a second position parameter, a third position parameter and a fourth position parameter.
4. The graph data characterization method according to claim 1, characterized in that: The step of obtaining a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the decay index includes: Obtaining a degree matrix and a matrix trace of the adjacency matrix, and obtaining a normalized adjacency matrix according to the degree matrix and the matrix trace; The product result of the attenuation index and the normalized adjacency matrix is obtained, and a difference matrix is obtained by subtracting the product result from a unit matrix, and the relative distance tensor is obtained according to an inverse matrix of the difference matrix.
5. The graph data characterization method according to claim 1, characterized in that: The updating of the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix includes: Applying a resizing function to the relative distance tensor and then applying a first activation function to obtain an activated distance tensor; Multiplying the activation distance tensor and the relative distance tensor to obtain a projected distance tensor; The hidden edge feature matrix is obtained based on the projected distance tensor and the edge feature matrix.
6. The graph data characterization method according to claim 1, characterized in that: The obtaining the attention distance corresponding to the node based on the hidden edge feature matrix includes: Taking the node as the target node, selecting computing nodes one by one from other nodes other than the target node; The element value corresponding to the edge between the target node and the calculation node in the hidden edge feature matrix is obtained as the attention distance.
7. The graph data characterization method according to claim 6, characterized in that: The optimized position parameters at least include: a first position parameter and a second position parameter, and obtaining the attention score between the node and other nodes according to the attention distance includes: Based on the optimized position parameters, acquiring an updated target query vector of the target node and an updated calculation key vector of the calculation node, and calculating a vector sum of the updated target query vector and the updated calculation key vector to obtain an initial sum vector; The products of the first weight matrix, the second weight matrix and the attention score are respectively used as the first attention value and the second attention value, a third attention value is obtained by calculating the inner product of the initial sum vector and the first attention value, and the sum of the third attention value and the second attention value is input into a second activation function, and an intermediate attention distance corresponding to the calculation node is output; The attention score is obtained according to the intermediate attention distance, the updated target query vector and the updated calculation key vector.
8. The graph data characterization method according to claim 7, characterized in that: The obtaining of the updated target query vector of the target node and the updated computation key vector of the computation node comprises: Acquire a target node feature of the target node and a computing node feature of the computing node from the node feature matrix; Acquire an initial target query vector of the target node, and obtain the updated target query vector based on the optimized position parameter, the initial target query vector and the target node feature; An initial computation key vector of the computation node is obtained, and the updated computation key vector is obtained based on the optimized position parameter, the initial computation key vector and the computation node feature.
9. The graph data characterization method according to claim 8, characterized in that: The obtaining the updated target query vector based on the optimized position parameter, the initial target query vector and the target node feature comprises: Obtaining the product of the initial target query vector and the target node feature to obtain a query intermediate value; Calculate a target position code corresponding to the target node according to the first position parameter and the second position parameter; The inner product of the query intermediate value and the target position code is calculated to obtain the updated target query vector.
10. The graph data characterization method according to claim 8, characterized in that: The obtaining the single-head attention feature corresponding to the node at least according to the attention score and the optimized position parameter includes: Obtaining a calculation value vector of the calculation node, taking the product of the third weight matrix and the intermediate attention distance as a fourth attention value, obtaining a fifth attention value according to the sum of the calculation value vector and the fourth attention value, and calculating the product of the attention score and the fifth attention value to obtain an attention feature corresponding to the calculation node; All the attention features are accumulated to obtain the single-head attention features of the target node, and all nodes are taken as target nodes one by one to calculate the single-head attention features corresponding to all nodes.
11. The graph data characterization method according to claim 1, characterized in that: The obtaining of the node attention feature according to the single-head attention feature comprises: Obtain the single-head attention features corresponding to each attention head, and concatenate them in order to obtain the corresponding concatenation vector; After the concatenated vector is linearly transformed, it is normalized to obtain a normalized result, and the sum of the normalized result and the node feature matrix is calculated to obtain the node attention feature.
12. The graph data characterization method according to claim 11, characterized in that: The obtaining of graph data representation features according to the edge attention features and the node attention features includes: Obtain edge input data and node input data at the current moment, wherein the edge input data is the edge output data at the previous moment, and the node input data is the node output data at the previous moment; Using the edge input data to update the hidden edge feature matrix, obtaining the updated edge attention feature as the edge output data at the current moment; Using the node input data to update the node feature matrix, obtaining the updated node attention feature as the first output intermediate value; Adding the first output intermediate value to the node input data to obtain a second output intermediate value; Sending the second output intermediate value to a feedforward neural network for data processing to obtain a feedforward output result, and adding the feedforward output result and the second output intermediate value to obtain node output data at the current moment; Using the node output data of the current node as the node input data at the next moment, using the edge output data as the edge input data at the next moment, and performing at least one iteration process until a target node feature matrix and a target edge feature matrix are obtained as the graph data representation features; At the initial moment, the edge input data is the edge attention feature, the node input data is the node attention feature, the node output data at the last moment is the target node feature matrix, and the edge output data is the target edge feature matrix.
13. A graph data representation device, characterized in that: include: Graph data acquisition module: used to obtain edge feature matrix, node feature matrix and adjacency matrix from graph structure data; Optimization solution module: used for generating a connection optimization problem according to the graph structure data, and solving the connection optimization problem to obtain optimized position parameters; Relative distance calculation module: used for obtaining a relative distance tensor corresponding to the graph structure data according to the adjacency matrix and the decay index, and updating the edge feature matrix according to the relative distance tensor to obtain a hidden edge feature matrix; A node attention calculation module is used for obtaining, for any node, an attention distance corresponding to the node based on the hidden edge feature matrix, obtaining an attention score between the node and other nodes according to the attention distance, obtaining a single-head attention feature corresponding to the node according to at least the attention score and the optimized position parameter, and obtaining a node attention feature according to the single-head attention feature; Edge attention calculation module: used for obtaining edge attention features for any edge after performing linear mapping according to the hidden edge feature matrix; Data characterization module: used to obtain graph data characterization features based on the edge attention features and the node attention features.
14. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the graph data representation method according to any one of claims 1 to 12 when executing the computer program.
15. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the graph data representation method according to any one of claims 1 to 12 is implemented.