Graph Transform-based linear complexity node classification method and system

By introducing Linear Differential Attention and StartNet into Graph Transformer, combined with improved FNG kernel functions, the problem of high computational cost and difficulty in integrating graph structure information when processing large-scale graph data is solved, and efficient node classification and model expression capabilities are improved.

CN120011859APending Publication Date: 2025-05-16SHANDONG JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510151126.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Graph Transformer is computationally expensive when processing large-scale graph data, making it difficult to effectively integrate graph structure information into the Transformer architecture, and multi-layer stacking may cause excessive smoothing of node representations.

Method used

Linear Differential Attention and StartNet are introduced, combined with the improved FNG kernel function, the graph structure features are linearized and the calculation complexity is reduced, and the node classification accuracy is improved through the linear transformation module and the linear complexity module.

Benefits of technology

It effectively reduces the computational complexity, improves the node classification accuracy, avoids the problem of excessive smooth node representation, and enhances the model's ability to utilize graph structure information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011859A_ABST
    Figure CN120011859A_ABST
Patent Text Reader

Abstract

The invention provides a Graph Transform-based linear complexity node classification method and system, and the method comprises the steps: obtaining a to-be-classified graph and an adjacent matrix of the to-be-classified graph, and inputting the to-be-classified graph into an improved graph classification network; the network comprises a Transform and a GNN branch, the GNN branch is used for extracting graph structure features of a to-be-classified graph based on the to-be-classified graph and an adjacent matrix of the to-be-classified graph, a linear transformation module of the Transform branch is used for linearly transforming the to-be-classified graph into a query vector, linearly transforming the graph structure features into a key vector and a value vector, and linearly transforming the key vector and the value vector into an adjacent matrix of the to-be-classified graph. The linear complexity module locally normalizes a query vector based on an FNG kernel function; and through star operation fusion, obtaining a fusion vector, and finally predicting a node classification result based on the fusion vector. According to the method, the improved graph classification network is applied, commodity recommendation graph data features are accurately extracted, and the relationship between the user and the commodity is effectively mined; and efficient calculation is realized based on the FNG kernel function, the node classification accuracy is improved, and the commodity recommendation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer graph neural network, and in particular to a linear complexity node classification method and system based on Graph Transformer. Background Art

[0002] In the product recommendation system, the graph structure can effectively integrate multiple entities such as users, products, brands, stores and their complex relationships, and the classification of graph nodes plays a key role. By classifying the user nodes and product nodes in the graph, it is possible to mine the user's preference characteristics and the product's attribute characteristics, and then predict the user's potential interest in different products, accurately recommend products that meet their needs to users, and improve the accuracy of the recommendation system and user satisfaction.

[0003] In order to better process graph structured data to optimize applications such as recommendation systems, Graph Transformer has emerged as an innovative deep learning model. It combines the powerful capabilities of the Transformer architecture with the unique structure of graph data, aiming to address the limitations of traditional graph neural networks (GNNs) in processing complex graph data. The core of Transformer is the self-attention mechanism, which can capture the global dependencies between any two elements in sequence or graph data. GraphTransformer further expands this capability to make it suitable for non-Euclidean data structures such as social networks, molecular graphs, and knowledge graphs. Unlike traditional GNNs that mainly rely on local neighborhood information aggregation, Graph Transformer directly models the interactions between all nodes in the graph through the self-attention mechanism, comprehensively captures long-range dependencies and high-order structural features, and also enhances the use of graph topology information by introducing graph structure encoding. It shows higher flexibility when processing complex scenarios such as heterogeneous graphs and dynamic graphs, and demonstrates excellent performance in many fields such as molecular property prediction, drug discovery, social network analysis, recommendation systems, and knowledge graph reasoning.

[0004] However, Graph Transformer also faces some challenges. The computational complexity of the self-attention mechanism is relatively high. , the computational cost is high when processing large-scale graph data; the node connections in the graph structure are based on topological relationships, which are difficult to process directly using the original Transformer method. How to effectively integrate graph structure information into the Transformer architecture is still an open problem; multi-layer stacked Graph Transformers may cause node representations to be overly smooth, and vectors representing different nodes tend to be similar, affecting model performance. Summary of the invention

[0005] In order to solve the above problems, the present invention proposes a linear complexity node classification method and system based on Graph Transformer. By introducing Linear Differential Attention and StartNet into GraphTransformer, and proposing an improved FNG kernel function, it can not only use the global attention of the former to capture implicit relationships and improve classification accuracy, but also reduce computational complexity and improve efficiency with the help of linearization. It can also use its denoising ability to optimize data, greatly enrich the classification basis, and effectively improve the node classification accuracy.

[0006] In order to achieve the above object, the present invention adopts the following technical solution: In a first aspect, the present invention provides a linear complexity node classification method based on Graph Transformer, comprising: Obtaining a graph to be classified and its adjacency matrix; the graph to be classified is a recommended relationship graph of commodities to be classified; The graph to be classified and its adjacency matrix are input into the improved graph classification network to extract classification features and obtain classification results based on the classification features; The improved graph classification network includes a Transformer branch and a GNN branch; the GNN branch is used to extract the graph structure features of the graph to be classified based on the graph to be classified and its adjacency matrix; the Transformer branch includes a linear transformation module and a linear complexity module; the linear transformation module is used to linearly transform the graph to be classified into a query vector, and linearly transform the graph structure features into a key vector and a value vector; the linear complexity module is used to establish a mapping relationship between the query vector and the key vector based on the FNG kernel function, locally normalize the query vector, and obtain a normalized feature vector; the normalized feature vectors are fused through a star operation to obtain an aggregate vector, and a node classification result is predicted based on the aggregate vector.

[0007] Preferably, the Transformer branch and the GNN branch are connected based on a cross attention layer, so as to input the graph structure features obtained by the GNN branch into the Transformer branch.

[0008] Preferably, the linear transformation module is used to linearly transform the graph to be classified into a query vector, and linearly transform the graph structure features into a key vector and a value vector; specifically, it includes:

[0009] in, represents the query vector, represents the key vector, represents a value vector; Represent the original node input features and the output feature matrix after GNN conversion respectively; is the GNN algorithm, is a sparse adjacency matrix, N is the number of nodes; represents the ReLU activation operation, Indicates chunk operation; There are three sets of learnable parameters.

[0010] Preferably, the FNG kernel function is specifically:

[0011] in, Representation Node The set of first-order neighbor nodes of Representation Node The characteristic vector of Indicates the order of the norm, set ; For Node The neighbor nodes of When FNG is calculated, it is regarded as Expanded from neighborhood nodes to global nodes.

[0012] Preferably, the linear complexity module is used to establish a mapping relationship between the query vector and the key vector based on the FNG kernel function, and locally normalize the query vector to obtain a normalized feature vector; specifically including: Calculate the FNG norm of the query vector and the key vector based on the FNG kernel function; Divide the query vector and key vector by their respective FNG norms for preliminary local normalization; Based on the value vector and the preliminary locally normalized query vector and key vector, attention weighted fusion is performed to obtain a fused feature vector; The fused feature vector is divided again by its FNG norm to obtain the final normalized feature vector.

[0013] Preferably, before fusing the normalized feature vectors through the star operation, the method further includes: denoising the normalized feature vectors, avoiding excessive denoising by introducing an adjacency matrix, and achieving a high degree of fusion of graph structure information.

[0014] Preferably, the normalized feature vectors are fused through a star operation, and the fusion method is element-by-element multiplication.

[0015] In a second aspect, the present invention provides a linear complexity node classification system based on Graph Transformer, comprising: A data acquisition unit, used to acquire a graph to be classified and its adjacency matrix; the graph to be classified is a recommended relationship graph of commodities to be classified; A node classification unit is used to input the graph to be classified and its adjacency matrix into the improved graph classification network, extract classification features, and obtain classification results based on the classification features; The improved graph classification network includes a Transformer branch and a GNN branch; the GNN branch is used to extract the graph structure features of the graph to be classified based on the graph to be classified and its adjacency matrix; the Transformer branch includes a linear transformation module and a linear complexity module; the linear transformation module is used to linearly transform the graph to be classified into a query vector, and linearly transform the graph structure features into a key vector and a value vector; the linear complexity module is used to establish a mapping relationship between the query vector and the key vector based on the FNG kernel function, locally normalize the query vector, and obtain a normalized feature vector; the normalized feature vectors are fused through a star operation to obtain an aggregate vector, and a node classification result is predicted based on the aggregate vector.

[0016] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in a linear complexity node classification method based on Graph Transformer described in the first aspect.

[0017] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the linear complexity node classification method based on Graph Transformer described in the first aspect are implemented.

[0018] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention introduces Linear Differential Attention and StartNet, and uses the global attention of the former to capture implicit relationships to improve classification accuracy. Its linearization reduces computational complexity, solving the problem of high computational cost of the self-attention mechanism. Secondly, the improved node classification network includes GNN branches and Transformer branches, which can effectively integrate graph structure information into the architecture and overcome the difficulty in processing graph structure information. Furthermore, this method innovates in multiple processing modules, such as using FNG operators and star operations, to avoid the problem of over-smoothing of node representation caused by multi-layer stacking, maintain the unique characteristics of nodes, and optimize model performance.

[0019] (2) The present invention proposes the FNG kernel function, which constrains the normalized calculation within the multi-order node neighborhood and uses the topological structure information of the neighborhood relationship of the nodes in the graph to calculate and process the node features, instead of adopting a global perspective like the traditional Frobenius norm. It realizes the transition from traditional global unified linear transformation to local adaptive nonlinear transformation, thereby cleverly combining the graph structure information, improving the feature representation and model expression capabilities, and optimizing the computing efficiency and stability.

[0020] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their description are used to explain the present invention but do not constitute a limitation of the present invention.

[0022] Figure 1 A main flow chart of a linear complexity node classification method based on Graph Transformer provided by an embodiment of the present invention; Figure 2 A schematic diagram of an improved node classification network provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0024] Technical term explanation: 1. Linear computational complexity: refers to the linear relationship between the amount of model calculation and the growth of data; 2. Information aggregation: refers to the process in which the central node aggregates the information of neighboring nodes in the Graph Neural Network (GNN).

[0025] Embodiment 1 like Figure 1 As shown, this embodiment discloses a linear complexity node classification method based on Graph Transformer, comprising the following steps: S1: Obtain a graph to be classified and its adjacency matrix; the graph to be classified is a recommended relationship graph of commodities to be classified; S2: Input the graph to be classified and its adjacency matrix into the improved graph classification network, extract the classification features, and obtain the classification results based on the classification features.

[0026] Next, combine Figure 1, a linear complexity node classification method based on Graph Transformer disclosed in this embodiment is described in detail.

[0027] The FuseFormer method provided in this embodiment, as a Graph Transformer model, can classify nodes in large graphs with linear computational complexity. Graph Transformer uses a parallel architecture to cleverly integrate the topological characteristics of graph structured data and the modeling advantages of the Transformer architecture, providing a powerful solution for complex graph data relationship processing.

[0028] Graph Transformer maintains the local connection information and structural features between nodes through GNN, and with the powerful feature extraction capability of Transformer, it can effectively capture the implicit long-distance dependencies in the graph. This parallel architecture coincides with the architectural concept of StartNet (StartNet achieves feature fusion by mapping low-dimensional inputs to two independent high-dimensional feature spaces and then multiplying them element by element). Inspired by this, this embodiment applies the feature fusion idea of ​​StartNet to Graph Transformer, and uses this architecture to elegantly integrate the local structural features from GNN and the global dependency representation of Transformer.

[0029] In order to enhance the interaction between the two parallel modules of GNN and Transformer, the cross-attention mechanism is introduced as a bridge for feature interaction. This method is not only highly compatible with the self-attention mechanism, but also does not increase the amount of additional calculations. From the perspective of architecture design, Figure 1 The network structure shown actually acts as a connection through the cross-attention mechanism, thus combining two classic architectures in Graph Transformer to a certain extent: parallel architecture and serial architecture.

[0030] As a specific implementation method, first, a graph to be classified and its adjacency matrix are obtained, and the graph to be classified and its adjacency matrix are input into an improved graph classification network to extract classification features, and obtain classification results based on the classification features.

[0031] The improved graph classification network includes a Transformer branch and a GNN branch, wherein the GNN branch is used to extract the graph structure features of the graph to be classified based on the graph to be classified and its adjacency matrix. After GNN branching (this embodiment mainly uses GCN, and can also be expanded to other excellent GNN methods such as GAT), and combined with the adjacency matrix , and obtain the graph structure features processed by the GNN branch .

[0032] This process is to integrate the input data into the structural information of the graph, because GNN can effectively capture the topological relationship and feature propagation between nodes in the graph, thereby generating a feature representation that is more graph-structure aware. .

[0033] The Transformer branch includes a linear transformation module and a linear complexity module. Through the cross-attention mechanism, Enter the Transformer branch and input the data and In the Transformer branch, linear processing is performed based on the linear transformation module: After linear transformation, we get the query vector query. After linear transformation, we get the key vector key and the value vector value. Specifically:

[0034] in, is the GNN algorithm, is a sparse adjacency matrix, N is the number of nodes, Represent the original node input features and the output feature matrix after GNN conversion, There are three sets of learnable parameters.

[0035] From formula (2), we can see that the process of generating the key vector key is: After and learnable parameters Multiply and pass through the ReLU activation function and then perform the chunk operation to get and ; At the same time, the generation process of the value vector value is and Multiply to get ; and the query vector query ( and ) is the original input After After multiplication and ReLU activation function, the chunk operation is performed. By combining the original input information and the features of the graph structure information after GNN processing, the attention weight is calculated and the information is weightedly aggregated, which can more accurately capture the node relationship and structural features in the graph, thereby better processing graph data.

[0036] In addition, based on the Differential Transformer, in the linear complexity module, FNG is used as the kernel function for mapping the query vector query and the key vector key, and the ReLU activation function is introduced to ensure the non-negativity of the attention score, so as to better approximate the distribution of the softmax function and realize the linear Differential Transformer calculation.

[0037] It should be understood that in recent years, a lot of research has been carried out around the quadratic computational complexity of Transformer, such as: Linformer, Sparse Attention, Efficient Attention, Linear Transformer, FLatten Transformer. Among them, a more advanced method is to map the query and key by designing a kernel function, split the softmax function, and use the associative law of matrix multiplication to adjust the calculation order of query, key and value, thereby avoiding explicit calculation of attention scores. This method can significantly reduce the computational complexity from Reduce to , and is therefore considered to be an efficient and practical method. The core is to design a suitable kernel function , common kernel functions include: relu, elu, focused function, frobenius norm. However, these kernel functions have significant limitations in practical applications: on the one hand, overly simple kernel functions cannot impose sufficient constraints on the attention score; on the other hand, complex kernel functions lead to a large amount of additional calculations. In addition, these general kernel functions lack specific designs for graph structured data.

[0038] Therefore, in order to solve the problems that the existing kernel functions are too simple to impose sufficient constraints on the attention score, and that complex kernel functions will lead to a large amount of additional calculations, and that general kernel functions lack specific designs for graph structured data, this embodiment proposes a specific kernel function on the graph (Frobenius Norm on Graph, FNG). FNG effectively captures the essential characteristics of graph data by extending the Frobenius norm theory to the field of graph structured data.

[0039] Unlike the traditional Frobenius norm, which uses a global normalization strategy, FNG introduces a dynamic normalization mechanism based on local neighborhoods, which adds a large number of nonlinear transformations. Specifically, the global normalization factor used by the traditional Frobenius norm is determined by all nodes in the entire graph. This global perspective scaling strategy ignores local structural features, while FNG constrains the normalization calculation within multi-hop neighbors. Therefore, the general calculation of FNG can be defined as: (3) in, In this embodiment, the node The set of first-order neighbor nodes of Representation Node The characteristic vector of represents the order of the norm. In this embodiment, we set , calculated using the Frobenius norm, can of course also be extended to higher orders. On the other hand, when When , the Frobenius norm can be regarded as Expanded from neighborhood nodes to global nodes.

[0040] The FNG provided in this embodiment introduces a dynamic normalization mechanism based on local neighborhoods, constraining the normalization calculation within the multi-order node neighborhood. The set of first-order neighbor nodes Internal Node The FNG operation of the feature vector can achieve differentiated scaling of nodes on neighborhood nodes due to the different neighborhood conditions of different nodes, providing a more fine-grained feature representation for graph learning.

[0041] FNG not only inherits the stability advantage of Frobenius norm in numerical calculation, but also cleverly combines the graph structure information to achieve the transformation from traditional global unified linear transformation to local adaptive nonlinear transformation, thereby ensuring the stability of numerical calculation while improving the model's ability to express local features. The details are as follows:

[0042] in, is the Frobenius Norm on Graph proposed in this embodiment. Formula (4) represents , , and Normalize them separately by dividing them by their respective Norm. Normalization can make data of different dimensions or scales comparable, which helps stabilize the training process and improve the convergence speed of the model. The norm is more suitable for the characteristics of graph data and can better reflect the information related to the graph structure, so it is used Normalizing the norm can make subsequent calculations based on these normalized vectors (such as attention weight calculation, etc.) more reasonable and effective.

[0043] Formula (5) is the formula for calculating attention output. and is the weighted sum of attention calculated based on query and key, is a constant additional term related to the number of nodes or other factors, and finally we get and By calculating the attention output in this way, the query, key and value information are combined to achieve weighted aggregation of the input data. Information can be filtered and integrated according to the importance of different parts, thereby extracting more valuable feature representations, which helps to improve the model's understanding and processing capabilities of graph data.

[0044] Formula (6) is similar to formula (4). and Do it again Normalizing the norm, we get and Once again based on Normalization by norm can further standardize the output results, making them more numerically stable and comparable, so as to facilitate subsequent processing or better interaction and combination with other parts.

[0045] In this embodiment, the formula (5) is used The softmax function is split and the calculation order is exchanged to achieve Computational complexity. Although have , but compared with the quadratic complexity, the computational efficiency is greatly improved. Compared with the previously used Frobenius Norm method, a large number of applications in equations (4) and (6) Normalization is performed and the relu function is introduced to ensure the non-negativity of attention, which greatly enhances the nonlinear transformation ability of the model and effectively improves the expression ability of the model. In addition, the softmax function is removed in formula (5), thereby omitting the exp function. Correspondingly, The exp function in should also be removed and redefined for:

[0046] in, , , , and represents a learnable parameter.

[0047] In formula (5), the softmax function is removed, and accordingly, The exp function in should also be removed, and the relevant parameters or variables should be redefined and adjusted to adapt to the new calculation process and model structure to ensure the normal operation and performance of the model.

[0048] The overall computational complexity of FuseFormer is , the core of which is The computational complexity of the Linear Differential Transformer is , E is the number of edges, N is the number of nodes, is the edge set. Since we use FNG in equations (4) and (5) to reduce the computational complexity of Transformer from Reduce to In addition, this embodiment sets the batch size for large graphs for batch training, and directly uses the whole graph training for small graphs. Therefore, FuseFormer can efficiently perform training and reasoning on datasets of different sizes with linear computational complexity.

[0049] The core implementation of Differential Transformer is and By doing pixel-by-pixel subtraction, the noise topology can be removed. Although the original graph topology has noise structure, most of the structure still has a positive effect on GNN. Therefore, the graph structure information is still introduced in the denoising process to avoid excessive denoising. To this end, the denoising global attention definition combined with the graph structure is as follows:

[0050] Where A represents a sparse adjacency matrix and V represents a value vector.

[0051] In order to ensure the numerical stability of the training process, RMSNorm is used for normalization. Therefore, formula (8) can be written as:

[0052] in is a learnable parameter matrix. Finally, according to the StartNet architecture design, pixel-by-pixel multiplication is used for fusion. Therefore, the input Output after conversion by GNN and Transformer The following formula:

[0053] in is the activation function, The Linear Differential Transformer uses a lot of different normalization methods to greatly improve the generalization ability of the model and avoid overfitting to a certain extent (especially in small graphs).

[0054] The graph node classification method provided in this embodiment can be applied to product recommendation, protein molecule prediction, social network processing, paper attribute prediction, etc. In order to verify the effectiveness of this embodiment, an implementation method applied to product recommendation is given.

[0055] On homogeneous graph datasets such as Cora, Citeseer, Pubmed, Amazon-Computer, Amazon-Photo, Coauthor-CS, and Coauthor-Physics, FuseFormer has shown better performance than SGFormer. For example, on the Cora dataset, FuseFormer's accuracy reaches 85.46 ± 0.72, which is about 1.38 percentage points higher than SGFormer's 84.08 ± 0.41; on the Amazon-Computer dataset, FuseFormer's accuracy is 93.68 ± 0.21, significantly higher than SGFormer's 91.46 ± 0.66. These results show that FuseFormer can better capture the global dependencies between nodes on homogeneous graph data, especially on datasets with slightly more nodes, and is more stable.

[0056] FuseFormer also shows significant advantages on heterogeneous graph datasets such as Film, Squirrel, Chameleon, Wikics, Amazon-Ratings, and Tolokers. For example, on the Squirrel dataset, FuseFormer's accuracy is 44.25 ± 1.88, significantly higher than SGFormer's 39.52 ± 2.56; on the Tolokers dataset, FuseFormer's accuracy reaches 86.33 ± 0.51, further verifying its strong modeling ability on heterogeneous graph data. These datasets usually have complex graph structures and noise, and FuseFormer's performance shows that it has stronger robustness and adaptability when processing heterogeneous graph data.

[0057] FuseFormer outperforms SGFormer on both homogeneous and heterogeneous graph datasets. This is mainly due to its ability to capture global dependencies through the denoising self-attention mechanism and its design that effectively integrates graph structure information. On homogeneous graphs, FuseFormer can better model the similarities and associations between nodes; on heterogeneous graphs, FuseFormer can handle complex node types and edge relationships, showing stronger generalization capabilities.

[0058] Inspired by StartNet and Differential Transformer, this specific embodiment proposes FuseFormer, which is a Graph transformer method and an improved model on SGFormer, named Fuseformer. Fuseformer combines the advantages of both Transformer and GNN, and has linear computational complexity, and can be used to process various types of node classification tasks including homogeneous graphs, heterogeneous graphs and large graphs. Specifically, first, the input X (which can be graph data composed of user behavior data, product attribute data, etc.) is mapped to a high-dimensional nonlinear feature space through GNN and Linear Differential Transformer respectively, so that the features of products and users can be deeply mined and abstractly expressed; in order to strengthen the connection between the two parallel modules, the GNN information is fused in the Transformer module by cross-attention, so that the long-range dependency relationship between users and products and local neighborhood information can be fully captured, and the potential needs of users and the key attributes of products can be accurately grasped; finally, the two modules are multiplied pixel by pixel to realize the feature fusion between the two modules, providing a richer and more accurate feature basis for graph node classification.

[0059] The linear calculation of Differential Transformer is realized thanks to the FNG operator (Frobenius Normon Graph). FNG realizes the calculation of Frobenius norm on the graph, thereby achieving local normalization. In addition, FNG is used as the kernel function to split the replacement function to realize the linear calculation of Differential Transformer. This not only greatly improves the computing efficiency, but also ensures that when processing large-scale product recommendation graph data, the node classification task can be completed quickly and accurately, thereby improving the accuracy of product recommendations, improving the user experience in the product recommendation system, and increasing the success rate of matching users and products, bringing more commercial value to e-commerce platforms.

[0060] Embodiment 2 This embodiment provides a linear complexity node classification system based on Graph Transformer, including: A data acquisition unit, used to acquire a graph to be classified and its adjacency matrix; the graph to be classified is a recommended relationship graph of commodities to be classified; A node classification unit is used to input the graph to be classified and its adjacency matrix into the improved graph classification network, extract classification features, and obtain classification results based on the classification features; The improved graph classification network includes a Transformer branch and a GNN branch; the GNN branch is used to extract the graph structure features of the graph to be classified based on the graph to be classified and its adjacency matrix; the Transformer branch includes a linear transformation module and a linear complexity module; the linear transformation module is used to linearly transform the graph to be classified into a query vector, and linearly transform the graph structure features into a key vector and a value vector; the linear complexity module is used to establish a mapping relationship between the query vector and the key vector based on the FNG kernel function, locally normalize the query vector, and obtain a normalized feature vector; the normalized feature vectors are fused through a star operation to obtain an aggregate vector, and a node classification result is predicted based on the aggregate vector.

[0061] Embodiment 3 This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the linear complexity node classification method based on Graph Transformer as described in the first embodiment above are implemented.

[0062] Embodiment 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the linear complexity node classification method based on Graph Transformer as described in the first embodiment are implemented.

[0063] The steps or modules involved in the above embodiments 2 to 4 correspond to those in embodiment 1. For the specific implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0064] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A linear complexity node classification method based on Graph Transformer, characterized in that: include: Get the graph to be classified and its adjacency matrix; The to-be-classified graph is a recommended relationship graph of commodities to be classified; The graph to be classified and its adjacency matrix are input into the improved graph classification network to extract classification features and obtain classification results based on the classification features; The improved graph classification network includes a Transformer branch and a GNN branch; the GNN branch is used to extract the graph structure features of the graph to be classified based on the graph to be classified and its adjacency matrix; the Transformer branch includes a linear transformation module and a linear complexity module; the linear transformation module is used to linearly transform the graph to be classified into a query vector, and linearly transform the graph structure features into a key vector and a value vector; the linear complexity module is used to establish a mapping relationship between the query vector and the key vector based on the FNG kernel function, locally normalize the query vector, and obtain a normalized feature vector; the normalized feature vectors are fused through a star operation to obtain an aggregate vector, and a node classification result is predicted based on the aggregate vector.

2. A linear complexity node classification method based on Graph Transformer as claimed in claim 1, characterized in that: The Transformer branch and the GNN branch are connected based on a cross attention layer, which is used to input the graph structure features obtained by the GNN branch into the Transformer branch.

3. A linear complexity node classification method based on Graph Transformer as claimed in claim 1, characterized in that: The linear transformation module is used to linearly transform the to-be-classified graph into a query vector, and linearly transform the graph structure features into a key vector and a value vector; specifically, it includes: in, represents the query vector, represents the key vector, represents a value vector; Represent the original node input features and the output feature matrix after GNN conversion respectively; is the GNN algorithm, is a sparse adjacency matrix, N is the number of nodes; represents the ReLU activation operation, Indicates chunk operation; There are three sets of learnable parameters.

4. A linear complexity node classification method based on Graph Transformer as claimed in claim 1, characterized in that: The FNG kernel function is specifically: in, Representation Node The set of first-order neighbor nodes of Representation Node The characteristic vector of Indicates the order of the norm, set ; For Node The neighbor nodes of When FNG is calculated, it is regarded as Expanded from neighborhood nodes to global nodes.

5. A linear complexity node classification method based on Graph Transformer as claimed in claim 1, characterized in that: The linear complexity module is used to establish a mapping relationship between the query vector and the key vector based on the FNG kernel function, and locally normalize the query vector to obtain a normalized feature vector; specifically including: Calculate the FNG norm of the query vector and the key vector based on the FNG kernel function; Divide the query vector and key vector by their respective FNG norms for preliminary local normalization; Based on the value vector and the preliminary locally normalized query vector and key vector, attention weighted fusion is performed to obtain a fused feature vector; The fused feature vector is divided again by its FNG norm to obtain the final normalized feature vector.

6. A linear complexity node classification method based on Graph Transformer as claimed in claim 1, characterized in that: Before fusing the normalized feature vectors through the star operation, the method further includes: denoising the normalized feature vectors, avoiding excessive denoising by introducing an adjacency matrix, and achieving a high degree of fusion of graph structure information.

7. A linear complexity node classification method based on Graph Transformer as claimed in claim 1, characterized in that: The normalized feature vectors are fused through a star operation, and the fusion method is element-by-element multiplication.

8. A linear complexity graph classification system based on Graph Transformer, characterized in that: include: A data acquisition unit, used for acquiring the graph to be classified and its adjacency matrix; The to-be-classified graph is a recommended relationship graph of commodities to be classified; A node classification unit is used to input the graph to be classified and its adjacency matrix into the improved graph classification network, extract classification features, and obtain classification results based on the classification features; The improved graph classification network includes a Transformer branch and a GNN branch; the GNN branch is used to extract the graph structure features of the graph to be classified based on the graph to be classified and its adjacency matrix; the Transformer branch includes a linear transformation module and a linear complexity module; the linear transformation module is used to linearly transform the graph to be classified into a query vector, and linearly transform the graph structure features into a key vector and a value vector; the linear complexity module is used to establish a mapping relationship between the query vector and the key vector based on the FNG kernel function, locally normalize the query vector, and obtain a normalized feature vector; the normalized feature vectors are fused through a star operation to obtain an aggregate vector, and a node classification result is predicted based on the aggregate vector.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a linear complexity node classification method based on Graph Transformer as described in any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the linear complexity node classification method based on Graph Transformer as described in any one of claims 1-7 are implemented.