Multi-head attention divergence regularization-based graph neural network graph structure identification method and device, storage medium and equipment

By introducing a multi-head attention divergence regularization mechanism into graph attention networks, the problems of attention redundancy and insufficient expressive power are solved, improving the accuracy and stability of graph structure recognition. It is applicable to complex graph data scenarios such as social networks, molecular structure recognition, and transportation networks.

CN120932062APending Publication Date: 2025-11-11WUHAN VOCATIONAL COLLEGE OF SOFTWARE & ENG (WUHAN OPEN UNIV)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510951782.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing graph attention networks suffer from serious attention redundancy, insufficient expressive power, and limited recognition accuracy in graph structure recognition tasks. They are particularly difficult to accurately extract structural features in complex graph data scenarios, leading to a decline in the model's generalization ability.

Method used

A multi-head attention divergence regularization mechanism is introduced, which constrains the weight differences between attention heads through cosine similarity and introduces Euclidean distance at the node level to constrain the differences in output features. A joint loss function is constructed for model training to enhance the diversity and stability of the model.

Benefits of technology

It significantly improves the structure recognition ability and training stability of graph neural networks, increases node classification accuracy, enhances the generalization performance and anti-overfitting ability of the model, and is suitable for complex graph data scenarios such as social networks, molecular structure recognition and transportation networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932062A_ABST
    Figure CN120932062A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-head attention divergence regularization-based graph neural network graph structure identification method and device, a storage medium and equipment, and belongs to the field of application of a graph neural network in a structure identification task. According to the method, a multi-head attention divergence regularization mechanism is introduced, so that redundancy between attention channels is effectively inhibited, the independence and expression diversity of information channels are improved, and the discrimination capability of a graph neural network on a complex structure mode is remarkably enhanced; meanwhile, the node feature discrimination degree and the model structure recognition precision are improved through the constraint of the node embedding expression difference. The method has good training stability and generalization ability, adapts to various heterogeneous graph structure recognition tasks, has high modularization and pluggable performance, can be flexibly integrated into an existing graph neural network framework, and is suitable for efficient deployment in resource-constrained environments such as edge computing and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of graph neural network applications in structural recognition tasks, and specifically relates to a graph neural network multi-head attention optimization method, apparatus, storage medium and device for graph structural recognition tasks. Background Technology

[0002] Graph structure recognition tasks are widely used in various practical applications such as social network analysis, molecular structure modeling, traffic network planning, and intelligent security monitoring. The core of this task lies in extracting effective node representations or subgraph patterns from graph data with complex topologies for operations such as node classification, edge relationship identification, and structural anomaly detection. Examples include identifying potentially influential user nodes in social networks, identifying specific functional clusters in molecular graphs, or identifying congestion-sensitive areas in traffic maps.

[0003] Graph Neural Networks (GNNs) have become the mainstream method for graph structure recognition tasks in recent years due to their powerful graph structure modeling capabilities. GNNs effectively capture complex topological dependencies in graph structures by aggregating and embedding information about nodes and their adjacencies, thereby generating structural representations that can be used for classification or prediction. Common applications include node classification, graph classification, and link prediction.

[0004] Graph Attention Network (GAT) is an improved model developed based on Generative Neural Networks (GNNs). It innovatively introduces an attention mechanism, allowing each node to adaptively allocate weights based on the importance of its neighbors during information aggregation. This overcomes the information loss problem caused by the average weight strategy used in traditional GNNs when processing structurally heterogeneous graphs. To further enhance the model's ability to characterize structural details, GAT typically employs a multi-head attention mechanism, learning neighbor information from multiple perspectives through multiple parallel attention channels, aiming to improve the stability and diversity of representations.

[0005] However, in complex tasks such as graph structure recognition, existing GAT models still have the following technical bottlenecks:

[0006] Attention redundancy is severe: During training, multi-head attention channels often converge to similar or repetitive attention weight distributions, resulting in a lack of effective complementarity between multiple channels and weakening the model's ability to discriminate in different graph structure regions.

[0007] Decreased expressive power: Due to the lack of sufficient differences between attention heads, the model's overall nonlinear feature modeling ability is limited, making it difficult to accurately identify complex local structural patterns in heterogeneous graphs, thus affecting the accuracy of graph structure recognition.

[0008] Overfitting and convergence instability: Redundancy among multiple heads may also lead to excessive concentration of gradient propagation paths during the model training phase, resulting in local optima, convergence oscillations, and other phenomena. Especially when facing sparse or complex graphs, the model's generalization ability decreases, ultimately causing problems such as low graph structure recognition accuracy and low node classification accuracy.

[0009] Therefore, while existing graph attention networks have shown some performance in structure recognition tasks, further optimization is needed in areas such as information independence, diverse representation, and training stability of multi-head mechanisms to improve the performance and robustness of graph structure recognition. Summary of the Invention

[0010] This invention provides a graph structure recognition method, apparatus, storage medium, and device based on multi-head attention divergence regularization, aiming to solve key problems of existing GAT models in graph structure recognition tasks, such as severe attention channel redundancy, insufficient node representation ability, and limited recognition accuracy. Especially in complex graph data scenarios, such as social network pattern analysis, chemical molecule structure recognition, or path anomaly detection in intelligent traffic maps, traditional graph neural network models struggle to accurately extract structural features from multiple perspectives, resulting in weak structural discrimination and decreased generalization ability.

[0011] A graph structure recognition method based on multi-head attention divergence regularization using graph neural networks, with the following specific steps:

[0012] Step 1: For different scenarios, preprocess the input graph to construct the node feature matrix and the self-loop enhanced adjacency matrix, and input the feature vector into the GNN network;

[0013] The scenarios include social network scenarios, molecular structure recognition, and transportation network scenarios.

[0014] The preprocessing involves initializing the input graph G=(V,E) structurally. Graph G represents the entity relationship structure in a specific application, and the nodes... ,side Each node Corresponding to an initial feature vector This feature vector consists of the node's original attributes or observed data; the node feature matrix... Input the neural network. Here, N represents the number of nodes, and F represents the feature dimension of each node.

[0015] The construction of the self-loop enhanced adjacency matrix is ​​as follows: Based on the graph topology, an adjacency matrix A is constructed, and an identity matrix I is introduced to form a self-loop enhanced adjacency matrix ildeA=A+I, in order to increase node self-connection.

[0016] Step 2: The information of each node and its neighbors is aggregated in parallel through the multi-head attention mechanism of the GNN network to obtain normalized attention weights, which are then stacked globally as a weight distribution matrix, and the final feature representation of the node is obtained.

[0017] In a multi-head attention mechanism, each attention head has independent parameters. Let represent the linear mapping weight matrix of the k-th attention head.

[0018] First, for any node i and its neighboring nodes j, the k-th attention head projects the input features onto the feature subspace through a linear mapping, obtaining the mapped feature representation:

[0019]

[0020] Then, the feature projections of node i and its neighboring node j are concatenated, and the k-th attention head is used with the LeakyReLU activation function to calculate its original attention score:

[0021]

[0022] in, The attention scoring weight vector for the k-th attention head is used to learn the importance of node relationships;

[0023] Finally, the raw attention scores calculated for each attention head are normalized to obtain a weighted representation of the multiple attention heads:

[0024]

[0025] in, This represents the normalized attention weight of node i to its neighboring node j under the k-th attention head; N(i) represents the set of neighboring nodes of node i.

[0026] Integrating the feature representations of all K attention heads at node i, the final feature of node i is represented by the average concatenation:

[0027]

[0028] Step 3: Construct an attention divergence regularization term based on the cosine similarity of the weight distribution matrices among multiple attention heads to constrain the weight distribution differences between different attention heads and capture different adjacency patterns.

[0029] The attention divergence regularization term is represented as:

[0030]

[0031] in, This represents the weight distribution matrix of the i-th attention head. This represents the cosine phase velocity function.

[0032] Step 4: Calculate the differences between the final feature representations of nodes under different attention heads, construct an output divergence regularization term to keep the representations of nodes differentiated under different attention head outputs, and enhance the representation diversity of the GNN model.

[0033] The node-level output branching regularization term is represented as:

[0034]

[0035] .

[0036] Step 5: Combine the main task loss function, attention divergence regularization term, and output divergence regularization term to form the final optimization objective function, and then train the model.

[0037] The final optimized loss function is expressed as:

[0038]

[0039] in, . It is the main task loss function, used to guide core tasks such as node classification.

[0040] During model training, the Adam optimizer is used to update the parameters. The model training is complete when the final optimized loss function converges.

[0041] Step 6: After the model training is completed, apply it to the actual graph structure recognition task and output the recognition results;

[0042] In practical applications, after inputting graph data such as social networks, traffic networks, or molecular structure diagrams into the GAT model, the model outputs the category prediction, structure attribution, or anomaly identification results for each node through forward propagation. It is widely applicable to practical technical scenarios such as user profile analysis, key structure identification, and path optimization suggestions.

[0043] A graph neural network multi-head attention optimization device incorporating a divergence regularization mechanism includes:

[0044] ① Graph representation generation module, used to model node features and adjacency structure of the original graph data;

[0045] ② Multi-head attention computation module, used to build and run multiple attention heads and output the node representation of each head;

[0046] ③ Attention divergence regularization module, used to calculate the similarity between multiple attention matrices and output attention divergence regularization terms;

[0047] ④ Node representation divergence regularization module, used to calculate the difference between different attention head representations of each node and output the embedding divergence regularization term;

[0048] ⑤ Optimizer module, used to integrate multiple regularization terms with the main task loss to jointly train the graph neural network model.

[0049] The multi-head attention calculation module includes multiple attention calculation units with independent parameters, each of which performs weighted aggregation of adjacent nodes.

[0050] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.

[0051] A computing device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method described above.

[0052] The advantages and beneficial effects of this invention are as follows:

[0053] (1) This invention constructs an attention divergence regularization term based on cosine similarity to measure the difference in attention weights among multiple attention heads. This mechanism can significantly suppress the problem of convergence of weights among heads in multi-head attention, allowing each head to focus on learning different adjacency patterns and structural features. This improves the information diversity and representational ability of the multi-head attention mechanism, and enhances the model's adaptability and generalization performance to complex graph structures.

[0054] (2) Following the attention mechanism, this invention further designs a divergence regularization term for node representations, using Euclidean distance to constrain the similarity of output results from different attention heads, thereby enhancing feature differences at the output level. This improves the discriminativeness and separation of node representations, making the model's performance more stable in tasks such as classification and clustering, and reducing the possibility of redundant representations interfering with downstream tasks.

[0055] (3) This invention introduces a joint loss function, which integrates two regularization terms—attention divergence and node divergence—on the basis of the main task loss, and sets adjustable hyperparameters to control their influence. The training framework has good flexibility and scalability, making it easy to dynamically adapt and optimize the regularization strength in different tasks or data scenarios.

[0056] (4) The graph neural network multi-head attention optimization device of the present invention adopts a modular design, is compatible with mainstream graph neural network structures, does not depend on a specific model architecture, and is compatible with mainstream graph neural network models such as GAT, GATv2, and Graph Transformer, and has the flexibility of modular plug-and-play. While keeping the original model structure and computation graph unchanged, it can achieve low-cost functional enhancement, which is especially suitable for the lightweight deployment needs of embedded AI systems and edge computing environments.

[0057] (5) The regularization terms constructed in this invention are based on cosine distance and Euclidean distance, respectively, and have clear geometric meaning and statistical properties, which can be used for visualization analysis. This enhances the interpretability of the model training process, helps to analyze attention learning behavior and embedding change trends, and provides an intuitive basis for model debugging and algorithm optimization. Attached Figure Description

[0058] Figure 1 This is a flowchart illustrating the overall processing flow of a multi-head attention optimization method for graph neural networks according to the present invention.

[0059] Figure 2 This is a schematic diagram of the computational structure of the multi-head attention mechanism in this invention;

[0060] Figure 3 This is a design logic diagram of the attention divergence regularization term in this invention;

[0061] Figure 4 The logic diagram for designing the node embedding representation of the divergence regularization term in this invention is shown below;

[0062] Figure 5 The structure diagram of the joint loss function after introducing the double-divergence regularization term;

[0063] Figure 6 This is a comparison chart of the convergence curves of the loss function during training between the method of the present invention and the comparative method in this embodiment of the invention;

[0064] Figure 7 A schematic diagram of a computer storage medium designed to implement the method of the present invention;

[0065] Figure 8 A schematic diagram of a computer device designed to implement the method of the present invention. Detailed Implementation

[0066] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0067] This invention designs a dual regularization optimization mechanism based on the standard multi-head GAT framework: first, it introduces an attention divergence regularization term based on cosine similarity to suppress weight convergence between different attention heads; second, it introduces a node embedding divergence regularization term based on Euclidean distance to enhance the diversity of output representations across different channels. These two regularization terms are integrated into the loss function for end-to-end training, significantly improving the model's structure recognition ability and training stability without increasing model structural complexity.

[0068] This invention effectively suppresses redundancy between attention channels by introducing a multi-head attention divergence regularization mechanism, thereby improving the independence and expressive diversity of information channels and significantly enhancing the ability of graph neural networks to discriminate complex structural patterns. Simultaneously, by constraining the differences in node embedding representations, it improves the discriminative power of node features and the model's structural recognition accuracy. This method exhibits good training stability and generalization ability, adapts to various heterogeneous graph structure recognition tasks, and possesses high modularity and pluggability, enabling flexible integration into existing graph neural network frameworks and suitable for efficient deployment in resource-constrained environments such as edge computing.

[0069] This invention provides a multi-head attention optimization method and apparatus that introduces a divergence regularization mechanism in graph neural networks. By employing two techniques—attention mechanism regularization and node representation regularization—it effectively solves the channel redundancy and representation convergence problems in multi-head attention mechanisms, thereby improving the performance of graph learning tasks such as node classification.

[0070] In this embodiment, standard graph datasets such as Cora, Citeseer, and PubMed are used as examples. The number of attention heads is set to 8, and the Adam optimizer is used for training. The joint loss function has λ1 = 0.1 and λ2 = 0.01. The results show that this method has higher accuracy and stronger generalization ability than the standard GAT in node classification tasks.

[0071] The present invention is implemented in the following environment:

[0072] Software platform: Python 3.8, PyTorch 1.12, PyTorch-Geometric 2.1

[0073] Hardware platform: NVIDIA A100 GPU, CUDA 11.6, 2.0GHz clock speed, 40GB RAM

[0074] Experimental datasets: Cora (2708 nodes), Citeseer (3327 nodes), Pubmed (19717 nodes)

[0075] Parameter settings: Number of attention heads K=8, embedding dimension Optimizer: Adam, learning rate 0.005, weight decay 5e-4; regularization coefficients: λ1 = 0.1 (attention regularization), λ2 = 0.01 (embedding regularization).

[0076] Based on the aforementioned graph dataset and parameters, this invention proposes a multi-head attention optimization method in graph neural networks that incorporates a divergence regularization mechanism, such as... Figure 1 As shown, the specific steps are as follows:

[0077] Step 1: Graph structure initialization

[0078] In graph structure recognition tasks, the input graph G=(V,E) first needs to undergo structure initialization processing. Graph G represents the entity relationship structure in a specific application, including:

[0079] ① In social network scenarios, nodes Indicates social users, side This indicates that there is an interactive relationship between the users;

[0080] ② In the molecular structure recognition task, nodes represent atoms, and edges represent chemical bonds between atoms;

[0081] ③ In traffic network scenarios, nodes represent intersections or traffic sensing devices, and edges represent road connections or sensor communication paths.

[0082] Each node Corresponding to an initial feature vector This feature vector consists of the node's original attributes or observed data, for example:

[0083] ① In social networks, features may include user activity, number of posts, and encoding of interest categories;

[0084] ② In a molecular diagram, features can be composed of the types of atoms, electronic structure, hybrid orbital types, etc.;

[0085] ③ In a traffic map, node characteristics may include congestion index, traffic flow, average vehicle speed, traffic light status, etc.

[0086] Based on the graph's topology, an adjacency matrix A is constructed, and an identity matrix is ​​introduced to form ildeA=A+I to increase node self-connections; node features Input neural network.

[0087] This step sets up the standard GNN input layer by constructing a self-loop enhanced adjacency matrix, which helps nodes fully capture their own semantics and improves the stability of the initial feature representation.

[0088] Step 2: Calculation of features of multi-head attention mechanism

[0089] like Figure 2 As shown, the multi-head attention mechanism in this invention performs node feature representation and attention normalization of the adjacency structure through the operational structure of linear transformation, attention scoring, softmax normalization and aggregated output in each attention head. By integrating the perspectives of multiple attention channels, the robustness and diversity of feature representation are improved.

[0090] Assume each attention head has independent parameters. Let represent the linear mapping weight matrix of the k-th attention head; for any node i and its neighboring nodes j, the k-th attention head projects the input features onto the feature subspace through linear mapping, obtaining the mapped feature representation:

[0091]

[0092] Integrating the feature representations of all K attention heads for node i, the final feature representation of node i is:

[0093]

[0094] For each pair of adjacent nodes i,j, the k-th attention head calculates its original attention score:

[0095]

[0096] in, The attention scoring weight vector for the k-th attention head is used to learn the importance of node relationships;

[0097] The raw attention scores calculated for each attention head are normalized to obtain a weighted representation of the multiple attention heads:

[0098]

[0099] This represents the normalized attention weight of node i to its neighboring node j under the k-th attention head, where N(i) represents the set of neighboring nodes of node i.

[0100] Step 3: Attention Divergence Regularization Mechanism

[0101] like Figure 3 As shown, this mechanism demonstrates the similarity calculation process between attention matrices of different heads, reflecting the construction method of cosine similarity.

[0102] To reduce the similarity between attention heads, the following regularization term is introduced:

[0103]

[0104] in, Let represent the weight distribution matrix of the i-th attention head.

[0105] This regularization term is used to penalize situations where the weight distributions of multiple attention heads are too uniform, while encouraging differences in attention distributions and encouraging them to capture different adjacency patterns, thereby improving the network's discriminative ability.

[0106] Step 4: Embedding the divergence regularization mechanism

[0107] This mechanism demonstrates how Euclidean distance is calculated between the node representations output by each head to construct dissimilarity constraints.

[0108] Further introduce node-level output divergence regularization terms:

[0109]

[0110] This mechanism encourages different expressions of node features from different attention heads, captures heterogeneous information channels, avoids redundancy in output from multiple head pairs, and improves the expressive diversity of GNN models.

[0111] Step 5: Joint Loss and Training Optimization

[0112] like Figure 5 As shown, the final loss function is formed by combining the main task loss and the two regularization terms:

[0113]

[0114] The main task uses cross-entropy loss:

[0115]

[0116] The Adam optimizer is used to update parameters, and the training cycle is set to 200 epochs. This loss function structure, while preserving the original GAT performance, addresses its redundant channels and training instability issues, thus improving generalization performance. During training, a smaller loss function value generally indicates a stronger fit of the model to the current task, meaning that the model parameters are better under the current optimization objective.

[0117] Step six: After the model training is complete, it can be applied to actual graph structure recognition tasks:

[0118] After inputting graph data (such as social networks, traffic networks, or molecular structure diagrams), the model outputs the category prediction, structure attribution, or anomaly identification results for each node through forward propagation. It is widely applicable to practical technical scenarios such as user profiling analysis, key structure identification, and path optimization suggestions.

[0119] Example 1

[0120] On the Cora dataset, the DRGAT proposed in this invention improves accuracy by 1.3% compared to GAT, by 1.8% on Citeseer, and by 1.2% on PubMed;

[0121] The experimental data are shown in Table 1. The attention weights trained on the Cora dataset are visualized and analyzed, and the adjacency attention graphs corresponding to the 8 attention heads in the first layer are extracted.

[0122] In the GAT model, most attention heads focus on nodes of the same type with similar adjacent structures and features, and the distribution maps of multiple heads in the figure almost overlap.

[0123] In the DRGAT model, attention weights are clearly differentiated, with some heads focusing on edge nodes, minority class nodes, or structurally anomalous regions, demonstrating good heterogeneity.

[0124] Table 1. Experimental data for Example 1

[0125] Dataset GAT model DRGAT model Improvement rate Cora 81.5 82.8 1.3% Citeseer 70.3 72.1 1.8% Pubmed 79.0 80.2 1.2%

[0126] Comparison of convergence curves of loss functions, such as Figure 6 As shown:

[0127] The X-axis represents the epoch, and the Y-axis represents the training loss. GAT exhibits loss oscillations before the 50th epoch, with significant curve fluctuations. The DRGAT model's loss curve decreases steadily and converges within 100 epochs, demonstrating superior convergence speed and lower volatility. Figure 6 As can be seen, the DRGAT method proposed in this invention converges faster and the results are more stable during the training process. The attention head distribution is more dispersed, demonstrating stronger expressive power and anti-overfitting ability.

[0128] Summary of technical advantages:

[0129] Enhanced expressive power: By introducing attention divergence regularization, the attention regions of multiple attention heads complement each other, thereby improving the node representation ability;

[0130] Enhanced generalization performance: It demonstrates higher accuracy than GAT on multiple datasets, especially with a 1.8% improvement on complex Citeseer structures;

[0131] Training stability optimization: The joint regularization term makes the training process more controllable, accelerates the convergence speed by about 20%, and enhances the robustness of the model;

[0132] Easy to integrate and extend: This mechanism can be plugged into any GAT variant and is suitable for large-scale heterogeneous graph structure tasks.

[0133] Example 2

[0134] Please see the appendix Figure 7 As shown, based on the same inventive concept, this application also provides a computer-readable storage medium 300, on which a computer program 311 is stored, which, when executed, implements the method described in Embodiment 1.

[0135] Since the computer-readable storage medium described in Embodiment 3 of this invention is the same as that used in implementing the graph neural network graph structure recognition method based on multi-head attention divergence regularization of this invention, those skilled in the art can understand the specific structure and variations of this computer-readable storage medium based on the method described in Embodiment 1 of this invention, and therefore will not be repeated here. All computer-readable storage media used in the method of Embodiment 1 of this invention fall within the scope of protection of this invention.

[0136] Example 3

[0137] Based on the same inventive concept, this application also provides a computer device, such as... Figure 8 As shown, it includes a storage 401, a processor 402, and a computer program 403 stored in the storage and executable on the processor. When the processor 402 executes the above program, it implements the method in Embodiment 1.

[0138] Since the computer device described in Embodiment 4 of this invention is the same computer device used to implement the graph neural network graph structure recognition method based on multi-head attention divergence regularization in Embodiment 1 of this invention, those skilled in the art can understand the specific structure and variations of this computer device based on the method described in Embodiment 1 of this invention, and therefore will not be repeated here. All computer devices used in the method of Embodiment 1 of this invention fall within the scope of protection of this invention.

Claims

1. A graph structure recognition method based on multi-head attention divergence regularization in graph neural networks, characterized in that, The specific steps are as follows: Step 1: For different application scenarios, preprocess the input graph to construct the node feature matrix and the self-loop enhanced adjacency matrix, and input the feature vector into the GNN network; Step 2: The information of each node and its neighbors is aggregated in parallel through the multi-head attention mechanism of the GNN network to obtain normalized attention weights, which are then stacked globally as a weight distribution matrix, and the final feature representation of the node is obtained. In a multi-head attention mechanism, each attention head has independent parameters. This represents the linear mapping weight matrix for the k-th attention head; First, for any node i and its neighboring nodes j, the k-th attention head projects the input features onto the feature subspace through a linear mapping, obtaining the mapped feature representation: Then, the feature projections of node i and its neighboring node j are concatenated, and the k-th attention head is used with the LeakyReLU activation function to calculate its original attention score: in, The attention scoring weight vector for the k-th attention head is used to learn the importance of node relationships; Finally, the raw attention scores calculated for each attention head are normalized to obtain a weighted representation of the multiple attention heads: in, This represents the normalized attention weight of node i to its neighboring node j under the k-th attention head; N(i) represents the set of neighboring nodes of node i. Integrating the feature representations of all K attention heads at node i, the final feature of node i is represented by the average concatenation: Step 3: Construct an attention divergence regularization term based on the cosine similarity of the weight distribution matrices among multiple attention heads to constrain the weight distribution differences between different attention heads and capture different adjacency patterns. The attention divergence regularization term is represented as: in, This represents the weight distribution matrix of the i-th attention head. Represents the cosine phase velocity function; Step 4: Calculate the differences between the final feature representations of nodes under different attention heads, construct an output divergence regularization term to keep the representations of nodes differentiated under different attention head outputs, and enhance the representation diversity of the GNN model; The node-level output branching regularization term is represented as: ; Step 5: Combine the main task loss function, attention divergence regularization term, and output divergence regularization term to form the final optimization objective function, and then train the model. The final optimized loss function is expressed as: in, ; It serves as the main task loss function, guiding core tasks such as node classification. Step 6: After the model training is completed, apply it to the actual graph structure recognition task and output the recognition results; In practical applications, after inputting graph data such as social networks, traffic networks, or molecular structure diagrams into the GAT model, the model outputs the category prediction, structure attribution, or anomaly identification results for each node through forward propagation.

2. The graph structure recognition method based on multi-head attention divergence regularization of graph neural networks according to claim 1, characterized in that, The application scenarios include social network scenarios, molecular structure recognition, and transportation network scenarios. In social network scenarios, user profiles are analyzed; in molecular structure recognition scenarios, key structures are identified; and in transportation network scenarios, path optimization suggestions are provided.

3. The graph structure recognition method based on multi-head attention divergence regularization of graph neural networks according to claim 1, characterized in that, The preprocessing involves initializing the input graph G=(V,E) structurally. Graph G represents the entity relationship structure in a specific application, and the nodes... ,side Each node Corresponding to an initial feature vector This feature vector consists of the node's original attributes or observed data; the node feature matrix... Input the neural network, where N represents the number of nodes and F represents the feature dimension of each node.

4. The graph structure recognition method based on multi-head attention divergence regularization of graph neural networks according to claim 1, characterized in that, The construction of the self-loop enhanced adjacency matrix is ​​specifically as follows: based on the graph topology, an adjacency matrix A is constructed, and an identity matrix I is introduced to form a self-loop enhanced adjacency matrix ildeA=A+I, in order to increase node self-connection.

5. The graph structure recognition method based on multi-head attention divergence regularization of graph neural networks according to claim 1, characterized in that, The model training process involves updating parameters using the Adam optimizer. The model training is complete when the final optimized loss function converges.

6. A graph neural network multi-head attention optimization device for implementing the method of claim 1 by introducing a divergence regularization mechanism, characterized in that, include: ① Graph representation generation module, used to model node features and adjacency structure of the original graph data; ② Multi-head attention computation module, used to build and run multiple attention heads and output the node representation of each head; ③ Attention divergence regularization module, used to calculate the similarity between multiple attention matrices and output attention divergence regularization terms; ④ Node representation divergence regularization module, used to calculate the difference between different attention head representations of each node and output the embedding divergence regularization term; ⑤ Optimizer module, used to integrate multiple regularization terms with the main task loss to jointly train the graph neural network model.

7. The apparatus according to claim 6, characterized in that, The multi-head attention calculation module includes multiple attention calculation units with independent parameters, each of which performs weighted aggregation of adjacent nodes.

8. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 5.

9. A computing device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the method according to any one of claims 1 to 5.