Transferable graph hint attack method based on hierarchical sub-graph enhancement

Through the hierarchical subgraph enhancement method, the vulnerable nodes are selected and the perturbation is iteratively optimized by combining local and global information, which solves the problems of graph prompt attack adaptability and perturbation balance, and achieves efficient and stable attack effects.

CN120671124APending Publication Date: 2025-09-19SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510746348.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The figure suggests that the attack needs to adapt to the diversity of various downstream data sets, and it is difficult to strike a balance between attack concealment and disturbance effect. The conflict between the internal and external optimization goals makes the attack design complex.

Method used

A hierarchical subgraph enhancement method is adopted to extract local and global information through Ego network and Cut subgraph, combined with random walk return probability coding, VoteRank and K-Shell strategies are used to select vulnerable nodes, and hybrid perturbations are generated through iterative alternating optimization to attack the pre-trained graph model.

Benefits of technology

It significantly reduces the accuracy of downstream tasks of pre-trained models on multiple graph-structured datasets, demonstrating excellent attack performance, low budget and high efficiency, and strong transferability. The modules work together to enhance the overall attack capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671124A_ABST
    Figure CN120671124A_ABST
Patent Text Reader

Abstract

The invention discloses a transferable graph hint attack method based on hierarchical sub-graph enhancement, which aims at fragile graph hint attack in a pre-training graph model, strives to solve the transferability and stability problems which are less concerned by a traditional graph hint attack method, and mainly realizes a reliable scheme aiming at graph hint from the following three aspects. The method comprises the following steps: firstly, proposing a hierarchical sub-graph information extraction method, and obtaining global information and local information in data on which graph prompts depend by executing sub-graph extraction strategies of different hierarchies; secondly, executing a hierarchical node selection strategy according to the obtained hierarchical sub-graph information so as to determine the most suitable attack node; and finally, designing an alternate iteration disturbance generation optimization method, generating malicious disturbance for the selected target node, adding the malicious disturbance to the node, completing malicious attack on graph prompt dependence data, and further generating a poisoned graph prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a transferable graph hint attack method based on layered subgraph enhancement, belonging to the technical field of graph hint engineering. Background Art

[0002] Inspired by the application of pre-trained models in language and vision, graph pre-training [1] has become a powerful paradigm for learning intrinsic graph properties, significantly alleviating the challenges brought by the complexity and high cost of graph data labeling. It requires further steps (such as fine-tuning) to update the weights of the pre-trained model in order to transfer knowledge to downstream tasks. When the pre-trained model is too large, the fine-tuning process becomes very complicated. To alleviate this challenge, hint learning has emerged as an alternative to fine-tuning. Sun et al. [2] designed a feasible graph hint tuning method that uses link prediction as a pre-training task and uses learnable graph hints in the downstream node classification task. Tan et al. [3] introduced a virtual node as a graph hint to learn task-specific information, thereby improving the performance of the pre-trained model in downstream tasks. However, both methods are only applicable to node classification tasks and lack adaptability to other downstream tasks. Liu et al. [3] proposed a novel method to alleviate this limitation, which unifies the pre-training task and downstream task into a subgraph similarity task and then introduces learnable hints to guide the downstream task. Despite this, designing and computing a general template for graph tasks is a complex task. Fang et al. introduced a learnable vector as a graph hint, which was then embedded into the node features of downstream graph structure data, demonstrating the feasibility of using only graph hints to achieve the same goal. In addition, Yu et al. [6] successfully extended hints to heterogeneous graph learning by adopting a dual template design and dual hints, unifying pre-training and downstream tasks, and effectively bridging the gap caused by heterogeneous differences. Yu et al. [7] attempted to improve the performance of hint learning by introducing pre-labeling in pre-training and proposing a dual hint design combined with open hints. At present, research in the field of graph hint learning is still limited, and research on graph hint attacks is even rarer. So far, only a few studies have explored graph hint attacks. Song et al. [8] proposed a graph hint backdoor that effectively converts hints into triggers through a custom-designed trigger generation method. Lin et al.

[12] designed a backdoor attack framework for graph hint learning, using a feature-aware trigger generator and a fine-tuning-resistant graph hint poisoning method to implement the graph hint backdoor. Lyu et al. [9] implemented a cross-context backdoor attack on graph hint learning by formulating the backdoor attack as an optimization problem for feature collision in the pre-training stage. Zhu et al.

[10] successfully conducted attribute inference attacks and link inference attacks on graph-cue learning, demonstrating the existence of privacy vulnerabilities. Most current research on graph-cue threats focuses on backdoor attacks and privacy risks in graph-cue learning. However, little attention has been paid to studying the impact of cue attacks on the accuracy of pre-trained graph models. Motivated by this, this paper proposes an attack scheme targeting cue based on the fact that cue generation depends on downstream datasets.Once an attacker obtains the downstream dataset used for prompt learning and poisons it, the prompts generated using this data become poisoned prompts. Since most existing defenses

[10] focus on the pre-trained model itself, attacks via these poisoned graph prompts are difficult to detect.

[0003] References:

[0004] [1]Xia, J., Zhu, Y., Du, Y., Li, SZ, 2022. A-survey of pretraining ongraphs: Taxonomy, methods, and applications. arXiv preprint arXiv:2202.07893.

[0005] [2]Sun, M., Zhou, K., He, X., Wang, Y., Wang,

[0006] [3] Tan, Z., Guo, R., Ding, K., Liu, H., 2023. Virtual node tuning for few-shotnode classification, in: Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pp. 2177-2188.

[0007] [4] Liu, Z., Yu, X., Fang, Y., Zhang,

[0008] [5]Fang,T.,Zhang,Y.,Yang,Y.,Wang,C.,Chen,L.,2024.Universal prompttuning for graph neural networks.Advances in Neural Information ProcessingSystems 36.

[0009] [6]Yu,X.,Fang,Y.,Liu,Z.,Zhang,X.,2024b.Hgprompt:Bridging homogeneousand heterogeneous graphs for few-shot prompt learning,in:Proceedings of theAAAI Conference on Artificial Intelligence,pp.16578-16586.

[0010] [7]Yu,X.,Zhou,C.,Fang,Y.,Zhang,X.,2024c.Multigprompt for multitaskpre-training and prompting on graphs,in:Proceedings of the ACM on WebConference 2024,pp.515-526.

[0011] [8]Song,Y.,Singh,R.,Palanisamy,B.,2024.Krait:A backdoor attackagainst graph prompt tuning.arXiv preprint arXiv:2407.13068.

[0012] [9]Lyu,X.,Han,Y.,Wang,W.,Qian,H.,Tsang,I.,Zhang,X.,2024.Cross-contextbackdoor attacks against graph prompt learning,in:Proceedings of the 30th ACMSIGKDD Conference on Knowledge Discovery and Data Mining,pp.2094-2105.

[0013]

[10] Zhu, J., Lin,

[0014]

[11] Zhou, C., Li, G., Weng, H., Xiang, Y., 2025. Training-free graph anomalydetection: A simple approach via singular value decomposition, in: THE WEBCONFERENCE 2025.

[0015]

[12] Lin, M., Zhang, Z., Dai, E., Wu, Z., Wang, Y., Zhang, X., Wang, S., 2024. Trojanprompt attacks on graph neural networks. arXiv preprint arXiv:2410.13974. Summary of the Invention

[0016] Technical issues:

[0017] 1) Due to the diversity of downstream tasks, graph hint attacks must be adapted to various downstream datasets. Each dataset has different internal structures and properties, so a general attack strategy is usually ineffective.

[0018] 2) Since graph cues indirectly affect the performance of pre-trained models, sufficient perturbations are required to achieve a satisfactory attack. Doing so while meeting the requirement of attack stealth is a huge challenge.

[0019] 3) Given that graph-cue learning inherently involves optimization, the attack perturbation generation module becomes an optimization problem within its inner layer. However, the conflict between the inner and outer optimization objectives makes designing a reasonable optimization process a daunting challenge.

[0020] Technical solution:

[0021] The present invention provides a transferable graph hint attack method based on layered subgraph enhancement, comprising:

[0022] 1) Hierarchical Subgraph Information Extraction: Utilizing the node-based Ego network strategy and the graph-based Cut subgraph strategy, we extract subgraph information at different levels from the original graph. By using random walk return probabilistic coding, we combine subgraph information with backbone graph information to obtain hierarchical subgraph aggregation features that contain rich local and global information.

[0023] 2) Multi-perspective vulnerable node selection: For local information, the VoteRank strategy is adopted to select key nodes by initializing node voting information, calculating voting scores, weakening neighbor influence, and iteratively updating. For global information, node removal strategies or the K-Shell method are used to calculate node influence scores for different task data to identify vulnerable nodes.

[0024] 3) Iteratively alternating perturbation optimization: First, key parameters such as the perturbation learning rate and the hint learning rate, as well as the graph hint and perturbation, are initialized. Then, while keeping the perturbation fixed, the hint is optimized using a weighted cross-entropy loss function. Next, while keeping the hint fixed, methods such as projected gradient descent are used to optimize the node and edge perturbations, respectively. By iteratively alternating these two steps, the perturbation is gradually adjusted to generate adversarial examples.

[0025] The specific steps include:

[0026] Step 1: Select the downstream task data used for graph cue learning as the target dataset;

[0027] Step 2: extract global and local information from the target dataset through the hierarchical subgraph information extraction module;

[0028] Step 3: Use a hierarchical vulnerable node selection strategy to identify target vulnerable nodes based on global information and local information respectively;

[0029] Step 4: By alternately iterating the optimization module, a mixed perturbation of node features and topological structures is generated to generate poisoned samples to mislead graph prompt learning;

[0030] Step 5: Use poisoned samples to generate malicious graph prompts to attack the downstream task reasoning process of the pre-trained graph model and reduce the model prediction accuracy.

[0031] Furthermore, the specific implementation of step 1 is as follows:

[0032] For some downstream task datasets used in graph hint learning, the downstream node classification task dataset, graph classification task dataset and link prediction task dataset are defined as Set node , Set graph and Set edge .

[0033] Furthermore, the specific implementation of step 2 is as follows:

[0034] (a) Extracting local subgraphs using a node-based strategy: extracting subgraphs from a single node, taking local information about the graph rather than information about the entire graph, using an Ego network, and the number of subgraphs generated is equal to the number of nodes; specifically: for each node v in the graph, select the nodes and edges within k hops to construct an Ego subgraph, denoted as Ego(v) k , which contains node v and its k-hop neighbor nodes;

[0035] (b) A graph-based strategy is used to extract global subgraphs: Global information about the graph is required to obtain subgraphs. The number of generated subgraphs has no direct relationship with the number of nodes in the original graph. The improved Cut subgraph is selected as the extraction target of the global subgraph. Specifically, the betweenness centrality of the edges in the graph is calculated, and the edges with high betweenness centrality are deleted in sequence until the graph is divided into v blocks. The generated Cut subgraph is denoted as Cut(v). k ;

[0036] (c) Perform random walk return probability encoding on the Ego subgraph, Cut subgraph, and backbone graph to generate aggregate node features and And through the linear layer, the subgraph information of different levels and the backbone graph information are fused to obtain the aggregated features of node v corresponding to the Ego subgraph and the Cut subgraph. and The random walk return probability encoding formula is:

[0037]

[0038] in Represents subgraph G sub The return probability of a random walk starting from node v is, Represents subgraph G sub The probability of returning after a random walk of S steps starting from node v.

[0039] Furthermore, the specific implementation of step 3 is as follows:

[0040] (a) Local information selection strategy: Aggregate features of the Ego layer subgraph containing rich local information obtained in step 2 The improved VoteRank algorithm is used to make full use of these local information to iteratively select vulnerable nodes; specifically: each node will have a voting weight and voting ability And initialize it by defining:

[0041]

[0042] Among them, d irepresents the degree of node i, d j The degree of node j is obtained by counting the number of neighbor nodes including itself. In practice, the voting ability of different nodes should be distinguished from multiple perspectives, including position, importance and role in the graph. The obtained aggregate features It can be used to consider node importance information; therefore, the following improvements are made:

[0043]

[0044] in, is the aggregate feature of node v corresponding to the Ego subgraph is the maximum aggregate feature of the Ego subgraph;

[0045] Afterwards, in each round of voting, the node receives voting scores from its neighbors and votes for its neighbors. In each round of voting, the node with the highest score is determined as the most critical node. In practice, the strength of the relationship between different nodes is usually different and is determined by the edges between them. Therefore, voting weights are assigned to the edges between different nodes. In addition, to ensure the representativeness of the selected node, the voting power of the selected node's high-order neighbors is weakened in the calculation, thereby reducing the voting power of the selected node's two-hop neighbors. Finally, after calculating the voting scores of all nodes, the node with the highest score is selected.

[0046] In each round, only the information of two-hop neighbor nodes is used to update the selected nodes, and the steps of calculating voting scores and weakening influence are repeated until a sufficient number of nodes are selected;

[0047] (b) Global information selection strategy: Aggregate features of the Cut layer subgraph containing rich global information obtained in step 2 This information is used to efficiently calculate the influence of all nodes through a node removal strategy. This strategy is divided into three parts, each corresponding to a specific effect of node removal. For graph classification tasks, it may not be possible to directly use the node removal strategy to calculate the node influence score due to the lack of certain features in these data sets. Therefore, an alternative global information node selection strategy called K-Shell is adopted.

[0048] Initially, the calculation of node influence scores is explained in a general framework:

[0049]

[0050] in, Represents node v r The impact score, g θ (G) i Represents the graph neural network g θ For a node v in graph Gi The prediction results;

[0051] Due to the presence of the l1 norm in the summation term, the node influence cannot be directly calculated using the first-order derivative as in Equation (5). Therefore, the calculation process of the influence score is decomposed. If deleting a node will continuously change the class distribution of other nodes, the effect expressed in Equation (5) is redefined as:

[0052]

[0053] where δh cut This is due to the deletion of v r The resulting L-th layer representation h cut Further analysis shows that the impact can be decomposed into three parts: T1, T2 and T3; where T1 measures the vanishing information of the potential representation contributed by a node to its neighboring nodes; T2 captures the change of the normalized term associated with the neighboring nodes; T3 evaluates the change of the potential representation of the neighboring nodes; Subsequently, each part is approximated independently; If each node in the graph is structurally and functionally equivalent, the term T1 is approximated as:

[0054]

[0055] in, Indicates the ratio of the missing information of the potential representation contributed by the neighboring nodes to the total representation information.

[0056] If every node in the graph is structurally and functionally equivalent, T2 is approximately:

[0057]

[0058] in, Indicates the change of the adjacency matrix after removing the node.

[0059] Similarly, the term T3 is approximated as:

[0060]

[0061] Through the above analysis, the calculation of node impact factor is decomposed into three parts:

[0062] δf r ≈-T1+T2+T3(10)

[0063] Afterwards, vulnerable nodes are selected based on the calculated node influence scores. Further, the specific implementation of step 4 is as follows: the nodes determined by the local and global node selection strategies are denoted as V r1 , V r2 ,…,V rn; Then, this priority is used to add perturbations to obtain adversarial samples; and a perturbation generation method based on alternating iterative optimization is proposed for node attacks and topology attacks, where the ratios of the two perturbations are quantified by parameters ɑ and 1-α, respectively;

[0064] First, the basic parameters of the optimization process are initialized. A black-box attack setting is adopted, with the perturbation learning rate set to 0.02 and the prompt learning rate defined by the user. The number of iterations of the outer and inner optimization loops are denoted as numOuter and numInner, respectively. Afterwards, the graph prompt prompt and the perturbation ptb are initialized. The initialization of the graph prompt is determined by the prompt learning strategy, while the perturbation ptb is divided into node perturbation and edge perturbation. Specifically, the node perturbation is initialized by adding a random vector f to ptb×a nodes. For the initialization of the edge perturbation, the altered edge dataset is constructed by randomly sampling ptb×(1-α) node pairs from the obtained vulnerable nodes.

[0065] Secondly, the fixed perturbation is optimized for the prompt, and the optimization objective is defined as follows:

[0066]

[0067] ΔD=ΔX+ΔA (12)

[0068] Where A(M,P,D+ΔD) represents the accuracy of the pre-trained model M when using P as the image prompt and D+ΔD as the perturbed training data; due to the diversity of downstream tasks, the weighted cross entropy loss function is used to calculate the loss:

[0069]

[0070] In the i-th optimization step, a more ideal graph hint P is obtained by minimizing the loss function L(M,P,D+ΔD) i ;The prompt update strategy is determined by the specific prompt learning algorithm;

[0071] Then, the fixed prompt is optimized for adding disturbances, and the optimization objective is as follows:

[0072]

[0073] ΔD=Δ node +Δ edge (15)

[0074] ||Δ||2≤ε (16)

[0075] The corresponding loss is calculated using weighted cross entropy as follows:

[0076]

[0077] Since the perturbation is divided into two parts, whose ratio is controlled by the parameter α, the loss is also divided proportionally to update the node perturbation and the topology perturbation separately;

[0078] For node perturbations, we first calculate the gradient to guide the update process; select the target node for attack based on the priority ranking; and add the perturbation ΔX∈R n×d To perturb the node feature matrix X∈R n×d ; The node perturbation loss based on proportional allocation is L node ; Therefore, the corresponding gradient is calculated as follows:

[0079]

[0080] in represents the gradient information about the node perturbation; then, the perturbation is updated according to the gradient as follows:

[0081]

[0082] Update the perturbation according to the calculated gradient to obtain the perturbation ΔX after the jth round of optimization j ;This perturbation is then used in the subsequent optimization process;

[0083] For edge perturbation, projected gradient descent is used to solve the discrete edge perturbation optimization problem. Specifically, for the adjacency matrix A, the corresponding perturbation matrix ΔA is constructed. First, a set of candidate edges E is determined based on the vulnerable nodes. set , from which the edge perturbations are obtained; subsequently, the edge gradients are calculated using the derived loss As shown in formula (22):

[0084]

[0085] Use the obtained edge gradient to update the edge perturbation to get the j-th round edge perturbation ΔA j ; This is achieved by flipping the edge with the highest gradient score:

[0086]

[0087] Among them, e * represents the edge after perturbation;

[0088] This edge perturbation is then added to the perturbation matrix:

[0089] A j =Union(A j-1 ,e * ) (twenty two)

[0090] The above optimization process is repeated alternately until the set optimization target or optimization times are reached; finally, the optimal perturbation for the vulnerable node is obtained through iterative optimization: and The generated poisoning map prompt is P poison .

[0091] Furthermore, the specific implementation of step 5 is as follows: the poisoned sample P poison When used for downstream classification tasks, the accuracy of the graph pre-training model decreases.

[0092] Beneficial effects:

[0093] 1) Excellent Attack Performance: Experiments conducted on multiple different graph datasets demonstrate that TGPA can significantly reduce the accuracy of pre-trained graph models in downstream tasks. On a node classification dataset, the accuracy of the attacked pre-trained model decreased by 25% compared to the baseline; on a graph classification dataset, the accuracy decreased by 30%, significantly outperforming other compared methods.

[0094] 2) Low-Budget, Efficient Attack: Even with a low attack budget, TGPA can make pre-trained models perform worse than the baseline. Through an effective vulnerable node selection strategy, it reduces the budget required to achieve the desired attack effect, making it more feasible in practical applications.

[0095] 3) Strong Transferability: TGPA demonstrates excellent transferability across different datasets for the same task, as well as datasets for different tasks. Its attack performance exhibits minimal fluctuation across datasets, with a standard deviation significantly lower than that of most comparable methods, enabling stable attacks against pre-trained models.

[0096] 4) Module synergy advantage: The diverse sample generation module provides rich information for vulnerable node selection, improving node identification accuracy; the vulnerable node selection module effectively utilizes global and local information to accurately locate attack targets; the iterative alternating optimization perturbation module optimizes the attack effect through the optimization process. The modules work together to enhance the overall attack capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0097] Figure 1 It is a flow chart of the present invention.

[0098] Figure 2 Schematic diagram of the implementation process of the present invention. DETAILED DESCRIPTION

[0099] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0100] The goal of this paper is to attack the hints required by pre-trained graph models, thereby affecting the model's performance in downstream tasks. To achieve this goal more effectively and improve the success rate of the attack, the present invention focuses the budget on nodes that are easily perturbed, achieving the goal of ensuring both attack performance and stealth.

[0101] The overall framework of the proposed scheme is as follows Figure 1 As shown in the figure, it consists of three main components: (a) subgraph-based diverse sample generation; (b) vulnerable node selection based on multi-level information; and (c) hybrid perturbations generated through iterative alternating optimization. Initially, a subgraph-based expansion is used on the input graph data to capture different levels of information by splitting the subgraph and self-subgraph. Subsequently, the subgraph is integrated with the backbone graph to encode and derive initial features rich in additional samples. Then, the properties of the subgraph are utilized to identify vulnerable nodes in local and global node selection strategies. Finally, by applying iterative alternating optimization, hybrid perturbations are successfully introduced to the obtained target nodes. Using these perturbed graph structure data, poisoning hints can be generated to successfully attack the inference process of the pre-trained model.

[0102] 1. Diverse Sample Generation Based on Subgraphs

[0103] 1.1, Subgraph Extraction Strategy

[0104] The idea behind the node-based strategy is to extract a subgraph from a single node. This strategy captures local information about the graph, rather than information about the entire graph. The strategy employed is the Ego network, which generates a number of subgraphs equal to the number of nodes. Specifically, for each node v in the graph, we select nodes and edges within k hops of it to construct an Ego subgraph, denoted as Ego(v). k , which contains node v and its k-hop neighbor nodes; for the global subgraph, a graph-based strategy is used to extract it. The graph-based strategy requires global information about the graph to obtain the subgraph. The number of subgraphs generated by this strategy has no direct relationship with the number of nodes in the original graph. The improved Cut subgraph is selected as the extraction target of the global subgraph. Specifically, the betweenness centrality of the edges in the graph data is calculated, and the edges with high betweenness centrality are deleted in sequence until the graph is divided into v blocks. The generated Cut subgraph is represented as Cut(v) k .

[0105] 1.2, Acquisition of hierarchical subgraph aggregation features

[0106] In order to obtain richer aggregate feature information, the extracted subgraph must be integrated with the backbone graph, which requires an efficient graph encoding method. Among the existing methods, random walk return probability coding achieves the best balance between complexity and expressiveness. Formally, in the subgraph G sub The random walk return probability encoding of node v is defined as:

[0107]

[0108] in Represents the subgraph G sub The return probability of a random walk of s steps starting from node v in .

[0109] By applying random walk return probability coding to the Ego subgraph and the Cut subgraph, the corresponding return probabilities are obtained respectively. and Using only subgraph probabilistic coding is not enough, because subgraph information needs to be combined with backbone graph information to obtain effective hierarchical aggregation information. First, the same method is applied to derive backbone graph probabilistic coding. Subsequently, a linear layer is used to encode the subgraph and backbone graph probabilities to obtain the hierarchical subgraph aggregation information. The aggregation information of node v corresponding to the Ego subgraph and the Cut subgraph is expressed as and here, Contains rich local information about the original graph, and Contains rich global information.

[0110] 2. Vulnerable Node Selection Based on Multi-level Information

[0111] The next step is to select a series of vulnerable nodes to be attacked. For the hierarchical subgraph aggregation features containing local and global information obtained from the Ego subgraph and the Cut subgraph, different vulnerable node selection strategies based on local and global information are used to obtain the vulnerable nodes to be attacked in the next step.

[0112] 2.1, Vulnerable Node Selection Based on Local Information

[0113] Already obtained It contains rich local information about the original graph. We are now ready to leverage this local information using a node-based strategy. VoteRank is an important node selection strategy inspired by real-world voting mechanisms. Specifically, when person A votes for person B, A's support for the other person is typically weaker. Similarly, key nodes can be identified by iteratively calculating the voting scores of their neighboring nodes in the graph.

[0114] The improved VoteRank algorithm will be used to make full use of this local information to iteratively select vulnerable nodes. Specifically: Each node will have a voting weight and voting ability They are initialized by the following definitions:

[0115]

[0116] D i represents the degree of node i, which is obtained by counting the number of neighboring nodes including itself. The above calculation only considers the degree information between adjacent nodes. In practice, the voting ability of different nodes should be distinguished from multiple perspectives, such as position, importance, and role in the graph. The representation vector obtained in the previous section It can be used to consider node importance information. Therefore, the following improvements are made:

[0117]

[0118] Afterwards, in each voting round, a node receives voting scores from its neighbors and votes for them. In each voting round, the node with the highest score is determined as the most critical node. In practice, the strength of relationships between different nodes often varies, determined by the edges between them. To address this, voting weights are assigned to the edges between different nodes. Furthermore, to ensure the representativeness of the selected node, the voting power of the selected node's higher-order neighbors is weakened during the calculation, reducing the voting power of the selected node's two-hop neighbors. Finally, after calculating the voting scores of all nodes, the node with the highest score is selected.

[0119] The steps of calculating voting scores and weakening influence are repeated until a sufficient number of nodes are selected. It is important to emphasize that, as mentioned earlier, each round only uses information from two-hop neighbors to update the selected node, rather than information from all nodes. This operation not only helps avoid high computational costs but also embodies the principle of leveraging local information.

[0120] 2.2, Vulnerable Node Selection Based on Global Information

[0121] Then, the global information is used to identify vulnerable nodes, and the aggregated features of the Cut layer subgraph containing rich global information are obtained. This information is exploited through a node removal strategy to efficiently compute the influence of all nodes. The strategy is divided into three parts, each corresponding to a specific effect of node removal. For data from graph classification tasks, it may not be possible to directly compute node influence scores using the node removal strategy due to the absence of certain features in these data sets. Therefore, an alternative global information node selection strategy, called K-Shell, is employed. This approach is based on iteratively deleting nodes based on their degree centrality. Each node is given an index K, nodes with a degree less than or equal to K are deleted, and the degree value of the network is recalculated. This process is iterated until no more nodes can be deleted based on the threshold K. Nodes with higher K values ​​are more important in the network.

[0122] Initially, the calculation of node influence scores is explained in a general framework:

[0123]

[0124] Due to the presence of the l1 norm in the summation term, we cannot directly use the first-order derivative to calculate the node influence as above. To solve this problem, we decompose the calculation process of the influence score. If deleting a node will continuously change the class distribution of other nodes, then the effect expressed in equation (5) can be redefined as:

[0125]

[0126] where δh cut This is due to the deletion of v r Resulting δh cut Further analysis shows that the impact can be decomposed into three components: T1, T2, and T3. T1 measures the vanishing information of the latent representation that a node contributes to its neighboring nodes. T2 captures the change in the normalization term associated with neighboring nodes. T3 evaluates the change in the latent representation of neighboring nodes. Subsequently, each component is approximated independently.

[0127] If every node in the graph is structurally and functionally equivalent, the term T1 can be approximated as:

[0128]

[0129] If every node in the graph is structurally and functionally equivalent, T2 can be approximated as:

[0130]

[0131] Similarly, the term T3 can be approximated as:

[0132]

[0133] Through the above analysis, the calculation of node impact factor is successfully decomposed into three parts:

[0134] δf r ≈-T1+T2+T3(10)

[0135] Afterwards, vulnerable nodes can be selected based on the calculated node influence scores.

[0136] 3. Hybrid perturbations generated by iterative alternating optimization

[0137] The last step of the present invention is to add perturbations according to the node importance obtained in the previous section. The nodes determined by the local and global node selection strategies are prioritized and are denoted as V r1 , V r2 ,…,V rnAfterwards, perturbations are added using this priority to obtain adversarial samples. Here, a perturbation generation method based on alternating iterative optimization is proposed for node attacks and topology attacks, where the ratios of the two perturbations are quantified by parameters α and 1-α, respectively.

[0138] 3.1 Initialization

[0139] First, the basic parameters of the optimization process must be initialized. Given the black-box attack setting, the perturbation learning rate is set to 0.02, while the prompt learning rate is defined by the user. The number of iterations of the outer and inner optimization loops are denoted as numOuter and numInner, respectively. Afterwards, the graph prompt prompt and perturbation ptb are initialized. The initialization of the graph prompt is determined by the prompt learning strategy, while the perturbation ptb is divided into node perturbation and edge perturbation. Specifically, the initialization of the node perturbation is achieved by adding a random vector f to ptb×a nodes. For the initialization of the edge perturbation, the altered edge dataset is constructed by randomly sampling ptb×(1-α) node pairs from the obtained vulnerable nodes.

[0140] 3.2, Prompt Optimization of Fixed Perturbation

[0141] During the perturbation generation process, the proposed solution is transformed into a dual optimization problem. The outer optimization, called prompt optimization, is the core step of prompt learning. It is important to note that the perturbation remains unchanged during the prompt optimization process. The optimization objective is defined as follows:

[0142]

[0143] ΔD=ΔX+ΔA (12)

[0144] Where A(M,P,D+ΔD) represents the accuracy of the pre-trained model M when using P as the image prompt and D+ΔD as the perturbed training data. Due to the diversity of downstream tasks, the weighted cross entropy loss function is used to calculate the loss:

[0145]

[0146] In the i-th optimization step, a more ideal graph hint P is obtained by minimizing the loss function L(M,P,D+ΔD) i The cue updating strategy is determined by the specific cue learning algorithm.

[0147] 3.3, Perturbation Optimization of Fixed Prompts

[0148] Internal perturbation optimization is the opposite of the external prompt optimization. In this process, the fixed prompt P obtained in the previous step is used i The optimization objectives are as follows:

[0149]

[0150] ΔD=Δ node +Δ edge (15)

[0151] ||Δ||2≤ε (16)

[0152] The corresponding loss is calculated using weighted cross entropy as follows:

[0153]

[0154] Since the perturbation is divided into two parts with the ratio controlled by the parameter α, the loss is also divided proportionally to update the node perturbation and the topology perturbation separately.

[0155] For node perturbations, the gradient is first calculated to guide the update process. The target node of the attack is selected according to the priority ranking. By adding the perturbation ΔX∈R n×d To perturb the node feature matrix X∈R n×d The node perturbation loss based on proportional allocation is L node . Therefore, the corresponding gradient can be calculated as follows:

[0156]

[0157] in Represents the gradient information about the node perturbation. Afterwards, the perturbation is updated according to the gradient as follows:

[0158]

[0159] Update the perturbation according to the calculated gradient to obtain the perturbation ΔX after the jth round of optimization j This perturbation is then used in the subsequent optimization process.

[0160] For edge perturbation, Projected Gradient Descent (PGD) is used to solve the discrete edge perturbation optimization problem. Specifically, for the adjacency matrix A, the corresponding perturbation matrix ΔA is constructed. First, a set of candidate edges E are determined based on the vulnerable nodes. set , from which the edge perturbation is obtained. Subsequently, the edge gradient is calculated using the derived loss as shown in Equation (22):

[0161]

[0162] Use the obtained edge gradient to update the edge perturbation to get the j-th round edge perturbation ΔA j This is done by flipping the edge with the highest gradient score:

[0163]

[0164] This edge perturbation is then added to the perturbation matrix:

[0165] A j =Union(A j-1 ,e * ) (twenty two)

[0166] 3.4, Iterative Alternation Optimization Strategy

[0167] To solve a bi-level optimization problem with conflicting optimization objectives, an iterative alternating optimization strategy is proposed. Specifically, the bi-level optimization is decomposed into two interdependent steps: an outer layer and an inner layer, which are iteratively alternating. The outer layer optimization depends on the perturbations generated by the inner layer optimization, while the inner layer optimization depends on the hints obtained from the outer layer.

Claims

1. A transferable graph hint attack method based on hierarchical subgraph enhancement, characterized in that: The steps include: Step 1: Select the downstream task data used for graph cue learning as the target dataset; Step 2: extract global and local information from the target dataset through the hierarchical subgraph information extraction module; Step 3: Use a hierarchical vulnerable node selection strategy to identify target vulnerable nodes based on global information and local information respectively; Step 4: By alternately iterating the optimization module, a mixed perturbation of node features and topological structures is generated to generate poisoned samples to mislead graph prompt learning; Step 5: Use poisoned samples to generate malicious graph prompts to attack the downstream task reasoning process of the pre-trained graph model and reduce the model prediction accuracy.

2. The method for migratable graph hint attack based on layered subgraph enhancement as claimed in claim 1, characterized in that: The specific implementation of step 1 is as follows: For some downstream task datasets used in graph hint learning, the downstream node classification task dataset, graph classification task dataset and link prediction task dataset are defined as Set node , Set graph and Set edge .

3. The transferable graph hint attack method based on layered subgraph enhancement as claimed in claim 1, characterized in that: The specific implementation of step 2 is as follows: (a) Extracting local subgraphs using a node-based strategy: extracting subgraphs from a single node, taking local information about the graph rather than information about the entire graph, using an Ego network, and the number of subgraphs generated is equal to the number of nodes; specifically: for each node v in the graph, select the nodes and edges within k hops to construct an Ego subgraph, denoted as Ego(v) k , which contains node v and its k-hop neighbor nodes; (b) Using a graph-based strategy to extract global subgraphs: Global information about the graph is required to obtain subgraphs. The number of generated subgraphs has no direct relationship with the number of nodes in the original graph. The improved Cut subgraph is selected as the extraction target of the global subgraph. Specifically, the betweenness centrality of the edges in the graph is calculated, and the edges with high betweenness centrality are deleted in sequence until the graph is divided into v blocks. The generated Cut subgraph is denoted as Cut(v). k ; (c) Perform random walk return probability encoding on the Ego subgraph, Cut subgraph, and backbone graph to generate aggregate node features and And through the linear layer, the subgraph information of different levels and the backbone graph information are fused to obtain the aggregated features of node v corresponding to the Ego subgraph and the Cut subgraph. and The random walk return probability encoding formula is: in Represents subgraph G sub The return probability of a random walk starting from node v is, Represents subgraph G sub The probability of returning after a random walk of S steps starting from node v.

4. The method for migratable graph hint attack based on layered subgraph enhancement as claimed in claim 3, characterized in that: The specific implementation of step 3 is as follows: (a) Local information selection strategy: Aggregate features of the Ego layer subgraph containing rich local information obtained in step 2 The improved VoteRank algorithm is used to make full use of these local information to iteratively select vulnerable nodes; specifically: each node will have a voting weight and voting ability And initialize it by defining: Among them, d i represents the degree of node i, d j The degree of node j is obtained by counting the number of neighbor nodes including itself. In practice, the voting ability of different nodes should be distinguished from multiple perspectives, including position, importance and role in the graph. The obtained aggregate features It can be used to consider node importance information; therefore, the following improvements are made: in, is the aggregate feature of node v corresponding to the Ego subgraph is the maximum aggregate feature of the Ego subgraph; Afterwards, in each round of voting, the node receives voting scores from its neighbors and votes for its neighbors. In each round of voting, the node with the highest score is determined as the most critical node. In practice, the strength of the relationship between different nodes is usually different and is determined by the edges between them. Therefore, voting weights are assigned to the edges between different nodes. In addition, to ensure the representativeness of the selected node, the voting power of the selected node's high-order neighbors is weakened in the calculation, thereby reducing the voting power of the selected node's two-hop neighbors. Finally, after calculating the voting scores of all nodes, the node with the highest score is selected. In each round, only the information of two-hop neighbor nodes is used to update the selected nodes, and the steps of calculating voting scores and weakening influence are repeated until a sufficient number of nodes are selected; (b) Global information selection strategy: Aggregate features of the Cut layer subgraph containing rich global information obtained in step 2 This information is used to efficiently calculate the influence of all nodes through a node removal strategy. This strategy is divided into three parts, each corresponding to a specific effect of node removal. For graph classification tasks, it may not be possible to directly use the node removal strategy to calculate the node influence score due to the lack of certain features in these data sets. Therefore, an alternative global information node selection strategy called K-Shell is adopted. Initially, the calculation of node influence scores is explained in a general framework: Among them, F gθ (v r ) represents node v r The impact score, g θ (G) i Represents the graph neural network g θ For a node v in graph G i The prediction results; Due to the presence of the l1 norm in the summation term, the node influence cannot be directly calculated using the first-order derivative as in Equation (5). Therefore, the calculation process of the influence score is decomposed. If deleting a node will continuously change the class distribution of other nodes, the effect expressed in Equation (5) is redefined as: where δh cut This is due to the deletion of v r The resulting L-th layer representation h cut Further analysis shows that the impact can be decomposed into three parts: T1, T2 and T3; where T1 measures the vanishing information of the potential representation contributed by a node to its neighboring nodes; T2 captures the change of the normalized term associated with the neighboring nodes; T3 evaluates the change of the potential representation of the neighboring nodes; Subsequently, each part is approximated independently; If each node in the graph is structurally and functionally equivalent, the term T1 is approximated as: in, Indicates the ratio of the missing information of the potential representation contributed by the neighboring nodes to the total representation information. If every node in the graph is structurally and functionally equivalent, T2 is approximately: in, Indicates the change of the adjacency matrix after removing the node. Similarly, the term T3 is approximated as: Through the above analysis, the calculation of node impact factor is decomposed into three parts: δf r ≈-T1+T2+T3(10)After that, the vulnerable nodes are selected based on the calculated node influence scores.

5. The method for migratable graph hint attack based on layered subgraph enhancement as claimed in claim 4, characterized in that: The specific implementation of step 4 is as follows: the nodes determined by the local and global node selection strategies are represented as V r1 , V r2 ,…,V rn ; Then, this priority is used to add perturbations to obtain adversarial samples; and a perturbation generation method based on alternating iterative optimization is proposed for node attacks and topology attacks, where the ratios of the two perturbations are quantified by parameters α and 1-α, respectively; First, the basic parameters of the optimization process are initialized. A black-box attack setting is adopted, with the perturbation learning rate set to 0.02 and the prompt learning rate defined by the user. The number of iterations of the outer and inner optimization loops are denoted as numOuter and numInner, respectively. Afterwards, the graph prompt prompt and the perturbation ptb are initialized. The initialization of the graph prompt is determined by the prompt learning strategy, while the perturbation ptb is divided into node perturbation and edge perturbation. Specifically, the node perturbation is initialized by adding a random vector f to ptb×a nodes. For the initialization of the edge perturbation, the altered edge dataset is constructed by randomly sampling ptb×(1-α) node pairs from the obtained vulnerable nodes. Secondly, the fixed perturbation is optimized for the prompt, and the optimization objective is defined as follows: ΔD=ΔX+ΔA (12) Where A(M,P,D+ΔD) represents the accuracy of the pre-trained model M when using P as the image prompt and D+ΔD as the perturbed training data; due to the diversity of downstream tasks, the weighted cross entropy loss function is used to calculate the loss: In the i-th optimization step, a more ideal graph hint P is obtained by minimizing the loss function L(M, P, D+ΔD) i ;The prompt update strategy is determined by the specific prompt learning algorithm; Then, the fixed prompt is optimized for adding disturbances, and the optimization objective is as follows: ΔD=Δ node +D edge (15) ||Δ||2≤ε (16) The corresponding loss is calculated using weighted cross entropy as follows: Since the perturbation is divided into two parts, whose ratio is controlled by the parameter α, the loss is also divided proportionally to update the node perturbation and the topology perturbation separately; For node perturbations, gradients are first calculated to guide the update process; Select the target node of the attack according to the priority ranking; by adding perturbation ΔX∈R n×d To perturb the node feature matrix X∈R n×d ; The node perturbation loss based on proportional allocation is L node ; Therefore, the corresponding gradient is calculated as follows: in represents the gradient information about the node perturbation; then, the perturbation is updated according to the gradient as follows: Update the perturbation according to the calculated gradient to obtain the perturbation ΔX after the jth round of optimization j ; This perturbation is then used in the subsequent optimization process; For edge perturbations, projected gradient descent is used to solve the discrete edge perturbation optimization problem; Specifically: For the adjacency matrix A, construct the corresponding perturbation matrix ΔA; first, determine a set of candidate edges E based on the vulnerable nodes set , from which the edge perturbation is obtained; The derived loss is then used to calculate the edge gradients As shown in formula (22): Use the obtained edge gradient to update the edge perturbation to get the j-th round edge perturbation ΔA j ; This is achieved by flipping the edge with the highest gradient score: Among them, e * represents the edge after perturbation; This edge perturbation is then added to the perturbation matrix: A j =Union(A j-1 ,have been * ) (22) The above optimization process is repeated alternately until the set optimization target or optimization times are reached; finally, the optimal perturbation for the vulnerable node is obtained through iterative optimization: and The generated poisoning map prompt is P poison .

6. The method for migratable graph hint attack based on layered subgraph enhancement as claimed in claim 5, characterized in that: The specific implementation of step 5 is as follows: poison When used for downstream classification tasks, the accuracy of the graph pre-training model decreases.