Graph backdoor encoder defense method and device, and readable storage medium

By generating unlabeled augmented datasets and adversarial contrastive learning, combined with attention distillation mechanism, the concealment problem of graph backdoor attacks in self-supervised scenarios is solved, achieving efficient backdoor defense, improving the robustness and defense capability of the model, and making it suitable for low-computing scenarios such as edge devices.

CN120913046APending Publication Date: 2025-11-07NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511077647.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing graph backdoor encoder defense methods are difficult to effectively identify and defend against backdoor attacks in self-supervised scenarios without label information guidance. Furthermore, existing methods rely heavily on label information during supervised training, which leads to a decrease in defense accuracy when label guidance is lacking.

Method used

By generating at least two unlabeled augmented datasets, adversarial contrastive learning and attention distillation mechanisms are used to train the teacher encoder, calculate the distribution shift of the layer attention map, update the parameters of the encoder to be processed, and construct multiple loss functions to clean up the backdoor encoder, including the first loss function, the second loss function and the third loss function. Combining adversarial learning and image contrastive learning, the robustness and defense capability of the model are improved.

Benefits of technology

It effectively reduces the success rate of self-supervised backdoor attacks, improves the system's backdoor defense capabilities, adapts to low-computing-power scenarios such as edge devices, maintains high accuracy of downstream tasks, achieves lightweight deployment, and enhances overall defense capabilities without sacrificing the core functionality of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913046A_ABST
    Figure CN120913046A_ABST
Patent Text Reader

Abstract

The invention discloses a graph backdoor encoder defense method and device, and a readable storage medium. The method comprises the steps that at least two label-free enhanced data sets are generated based on a preset training data set, a to-be-processed encoder is trained based on the enhanced data sets to obtain a teacher encoder, and the training data set is a downstream data set or a subset of the downstream data set; defining identical attention operators in corresponding layers of the teacher encoder and the to-be-processed encoder, wherein the attention operators are used for outputting layer attention maps of corresponding encoder layers; calculating the distribution offset of the layer attention map of the teacher encoder and the layer attention map of the encoder to be processed in the same encoder layer; and fixing the parameters of the teacher encoder, and updating the parameters of the to-be-processed encoder based on the distribution offset. Compared with the prior art, through lightweight deployment, on the premise that the model precision is not reduced, the attack success rate of backdoor attack on graph self-supervised learning is effectively reduced, and the backdoor defense capability of the system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of graph self-supervised learning, and particularly relates to a graph backdoor encoder defense method, equipment and a readable storage medium. BACKGROUND

[0002] On the basis of the success of GNN, graph self-supervised learning (GSSL) as a new graph learning paradigm enables the encoder to learn high-quality feature representations from unlabeled data. However, the process of obtaining a well-trained encoder through GSSL is time-consuming and expensive, leading many practitioners to rely on third-party pre-trained encoders provided on public platforms. However, these third-party pre-trained encoders may be backdoor encoders injected with triggers by malicious attackers, posing a significant threat to the security and integrity of the GNN system.

[0003] As a typical type of attack, backdoor attacks in GNN pose a serious threat to the security of graph models. Graph backdoor attacks have attracted increasing attention and have been extensively studied by some researchers, including supervised and self-supervised scenarios. Among them, in the inference stage of the self-supervised scenario, the backdoor encoder incorrectly classifies the downstream test samples injected with triggers into the target class while correctly identifying normal samples. Therefore, the intrinsic need for label modification may arouse suspicion in supervised attacks, but it does not exist at all in the self-supervised scenario, making graph self-supervised backdoor attacks significantly more covert. Existing graph backdoor encoder defense methods rely on label information in the supervised training setting, and cannot accurately defend against self-supervised backdoor attacks without label information guidance.

[0004] Therefore, in view of the above technical problems, it is necessary to provide a graph backdoor encoder defense method, equipment and a readable storage medium.

[0005] The information disclosed in this BACKGROUND section is only intended to increase an understanding of the general context in which the present application can be practiced. It is not admitted that any of the information provided in this BACKGROUND section constitutes prior art to the present application. SUMMARY

[0006] The purpose of the present application is to provide a graph backdoor encoder defense method, equipment and a readable storage medium, which can effectively deal with backdoor attacks in the graph self-supervised learning scenario and break through the concealment barrier of self-supervised backdoor attacks.

[0007] In order to achieve the above-mentioned purpose, the technical scheme provided by an embodiment of the present application is as follows:

[0008] In a first aspect, the present application provides a graph backdoor encoder defense method applied to graph self-supervised learning, which comprises:

[0009] generating at least two unlabeled augmented data sets based on a preset training data set, and training a to-be-processed encoder based on the augmented data sets to obtain a teacher encoder, wherein the training data set is a downstream data set or a subset of the downstream data set;

[0010] defining the same attention operator in the corresponding layers of the teacher encoder and the to-be-processed encoder, the attention operator being used to output a layer attention map of a corresponding encoder layer;

[0011] calculating a distribution offset of the layer attention map of the teacher encoder and the layer attention map of the to-be-processed encoder in the same encoder layer;

[0012] fixing the parameters of the teacher encoder, updating the parameters of the to-be-processed encoder based on the distribution offset.

[0013] In one or more embodiments of the present application, generating at least two unlabeled augmented data sets and training a to-be-processed encoder based on the augmented data sets to obtain a teacher encoder comprises:

[0014] obtaining an original image in the training data set, and modifying the topological structure of the original image to generate at least two corresponding augmented images for each original image;

[0015] calculating an embedding vector of each node in the augmented image based on the to-be-processed encoder;

[0016] constructing a first loss function based on the similarity of the embedding vectors of the same node in different augmented images corresponding to the same original image, and / or the node similarity of the augmented images corresponding to different original images, and / or the similarity of the embedding vectors of different nodes in different augmented images corresponding to the same original image;

[0017] updating the parameters of the to-be-processed encoder based on the first loss function until a preset condition is reached.

[0018] In one or more embodiments of the present application, the method further comprises:

[0019] adding perturbations to the original image to generate a corresponding adversarial image; wherein the sum of the absolute values of the perturbations added to each node in all feature latitudes in the adversarial image is less than a preset first threshold value, and the total sum of the number of changes in the connection state between nodes in the adversarial image is less than a preset second threshold value, compared with the original image;

[0020] The second loss function is constructed based on the similarity of the same node embedding vectors in the enhanced image and the adversarial image corresponding to the same original image, and / or the node similarity of the enhanced image and the adversarial image corresponding to different original images, and / or the similarity of different node embedding vectors in different enhanced images and adversarial images corresponding to the same original image

[0021] The parameters of the to-be-processed encoder are updated based on the second loss function until a preset condition is reached.

[0022] In one or more embodiments of the present application, the first loss function and the second loss function are respectively:

[0023]

[0024]

[0025]

[0026] wherein, is an enhanced data set; represents the loss of each contrast instance ; represents the cosine similarity of embedding and ; is a temperature parameter; is a weight coefficient for balancing the two losses; is the first loss function; is the second loss function.

[0027] In one or more embodiments of the present application, the method further comprises updating the related features of the adversarial image in each training round:

[0028]

[0029]

[0030]

[0031]

[0032]

[0033]

[0034] wherein, represents the node features of the adversarial image ; represents the adjacency matrix of the adversarial image ; representing the original image node features of the original graph; representing the original image adjacency matrix of the original graph; representing the perturbation to be added to the node feature matrix; is a symmetric matrix for modifying edges in the graph, and the element value 1 represents modification and 0 represents no change; represents an element-wise multiplication operation for determining whether to add or delete an edge; is an element-wise sign function; and are hyperparameters affecting the generation of the t-th round of adversarial samples; representing ; and represent projection operations on and respectively; is a complementary graph; is a matrix of all 1s, is an identity matrix, and A represents the original adjacency matrix; only in the original graph without connection and not self-loop is 1.

[0035] In one or more embodiments of the present application, the distribution of the layer attention map of the teacher encoder is offset from the layer attention map of the to-be-processed encoder in the same encoder layer, including:

[0036] normalizing the layer attention map output by the layer attention operator of each layer of the teacher encoder to obtain the corresponding layer attention weight;

[0037] normalizing the layer attention map output by the layer attention operator of each layer of the to-be-processed encoder to obtain the corresponding layer attention weight;

[0038] obtaining the difference between the layer attention weights of the teacher encoder and the to-be-processed encoder in the same encoder layer, and calculating the distribution difference between the layer attention weights of the teacher encoder and the to-be-processed encoder based on the L2 norm.

[0039] In one or more embodiments of the present application, the to-be-processed encoder parameters are updated based on the distribution offset, including:

[0040] inputting the original image to calculate the cross-correlation matrix of the output features of the teacher encoder and the to-be-processed encoder;

[0041] constructing a third loss function based on the cross-correlation matrix and the distribution difference between the layer attention weights;

[0042] update the to-be-processed encoder parameters based on the third loss function.

[0043] In one or more embodiments of the present application, the third loss function comprises:

[0044]

[0045] wherein, represents the cross-correlation matrix between the mth feature of the teacher encoder and the nth feature of the to-be-processed encoder; is a weight for controlling the consistency of the diagonal elements and the independence of the non-diagonal elements; is the third loss function.

[0046] In a second aspect, the present application provides a computer device, comprising a memory and a processor, which are in communication connection with each other, and the memory stores computer instructions, and the processor executes the computer instructions to perform the graph backdoor encoder defense method.

[0047] In a third aspect, the present application provides a computer readable storage medium, which stores computer instructions for causing a computer to execute the graph backdoor encoder defense method.

[0048] Compared with the prior art, the graph backdoor encoder defense method provided by the present application can break through the concealment barrier of self-supervised backdoor attacks, effectively reduce the attack success rate of backdoor attacks on graph self-supervised learning, and improve the backdoor defense capability of the system. Secondly, the method provided by the present application discards the complex architecture of traditional dependence on clean encoder contrast, is closer to the extreme harsh scene under actual application, and can complete the training of the teacher model only with a small amount of downstream data set, so that the resource consumption can be reduced, the method is suitable for low-power scenarios such as edge devices, and lightweight deployment is realized. In addition, the BT loss constraint and the adversarial regularization mechanism based on self-supervised distillation eliminate the backdoor neurons while maintaining the integrity of the feature space, and still have a high downstream task accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0050] Figure 1 It is a scene diagram of the graph backdoor encoder defense in an embodiment of the present application.

[0051] Figure 2 Figure 1 is a flow chart of a backdoor encoder defense method in an embodiment of the present application;

[0052] Figure 3 Figure 1 is a structural block diagram of a backdoor encoder defense system in an embodiment of the present application;

[0053] Figure 4 Figure 1 is a structural block diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0055] Unless otherwise explicitly stated, throughout the specification and claims, the term "comprise" or variations such as "comprises" or "comprising" will be understood to imply the inclusion of a stated element or component, but not the exclusion of any other element or component.

[0056] The complex topology inherent in graph data provides a unique concealment advantage for backdoor attacks, making it a significant challenge to implement effective defense in graph learning scenarios. Current mainstream graph backdoor defense mechanisms mainly focus on attack response in supervised learning environments. Although some of these techniques exhibit certain defense effects, there is a fundamental limitation common to these methods: they rely heavily on accurate label information provided during supervised training as a defense basis. Once such explicit guidance is lacking, existing methods are unable to effectively identify and resist more concealed self-supervised backdoor attacks, and their defense accuracy will decrease significantly.

[0057] The inventors of the present application have found the main shortcomings of the prior art and, based on the shortcomings of the prior art, have proposed a new technical implementation idea: generating a teacher model based on adversarial contrastive learning of the encoder to be processed, and implementing backdoor defense in cooperation with an attention distillation mechanism. Specifically, this method first uses a small amount of downstream task data to perform adversarial contrastive learning on the teacher encoder, effectively overcoming the problem of model performance degradation caused by multiple rounds of training iterations in a label-free environment, thereby significantly improving the representation ability and robustness of the teacher model. On this basis, the attention distillation technology is innovatively deployed in the hidden layer of the model, guiding the student model to learn the key feature focusing pattern of the teacher model, with the aim of fundamentally identifying and eliminating the hidden backdoor mechanism.

[0058] Compared with the prior art, the backdoor defense method provided by the application can effectively penetrate the concealment barrier of the self-supervised backdoor attack. The method can significantly reduce the success rate of the backdoor attack on the graph self-supervised learning model while maintaining a high accuracy of the downstream task, thereby systematically improving the overall defense capability of the graph learning model against the backdoor threat without sacrificing the core function of the model, and providing a strong guarantee for safe and reliable graph self-supervised learning.

[0059] Please refer to Figure 1 which shows an application scene schematic diagram of the graph backdoor encoder defense method provided by the application in an embodiment, and the scene specifically includes a model enhancement unit 101, a backdoor purification unit 102 and a user terminal 103.

[0060] It should be noted that the model enhancement unit 101, the backdoor purification unit 102 and the user terminal 103 are all provided with communication connection, and the communication network extended by the above communication connection can include various connection types, including but not limited to wired connection, wireless connection or optical cable connection and the like. At the same time, the communication network can be a local area network, a metropolitan area network, a wide area network or any combination of the three.

[0061] The model enhancement unit 101 is configured to form a corresponding enhanced data set and an adversarial data set based on a preset training data set. The standard contrast loss and the adversarial loss are combined to train the to-be-processed encoder, and a teacher encoder for purifying the to-be-processed encoder is formed.

[0062] The backdoor purification unit 102 is configured to configure an attention operator at each encoder layer, quantify the importance of neurons based on a layer attention graph, and compare each layer attention of the teacher encoder and the to-be-processed encoder. The attention distribution difference is used as a loss function, and the attention distribution of the to-be-processed encoder is forced to imitate the attention distribution of the teacher encoder by back propagation, so as to suppress the backdoor activation.

[0063] The user terminal 103 is configured to configure the parameters, algorithms, downstream data sets, suspicious backdoor encoders or encoders to be processed necessary for the execution of the graph backdoor encoder defense method provided by the application. It should be noted that the user terminal 103 is installed with a computer software program matched with the graph backdoor encoder defense method provided by the application; the user terminal 103 can include but is not limited to a desktop computer (PC terminal), a desktop computer, a smart phone, a handheld computer, a tablet computer, a personal digital assistant (PDA) and the like portable electronic devices or wearable electronic devices, and the embodiments of the application do not limit the above content.

[0064] It should be noted that the graph backdoor encoder defense method of the embodiment of the present application can be applied to the graph backdoor encoder defense system of the embodiment of the present application. The graph backdoor encoder defense system can be configured in a terminal. The terminal can include, but is not limited to, a PC (Personal Computer), a PDA (tablet computer), a smart phone, a smart wearable device, and the like.

[0065] Please refer to Figure 2 FIG. 1 shows a flowchart of the graph backdoor encoder defense in an embodiment of the present application. The graph backdoor encoder defense method specifically includes the following steps:

[0066] S201: generating at least two unlabeled enhanced data sets based on a preset training data set, and training a to-be-processed encoder based on the enhanced data sets to obtain a teacher encoder;

[0067] It should be noted that, considering the actual application environment, the to-be-processed encoder is mostly a backdoor encoder injected with a trigger by a malicious attacker from a third party. Such an encoder not only poses a major threat to the security and integrity of the GNN system, but also means that only downstream data sets can be used for defense. In order to simulate this kind of extreme and severe condition, the present application also trains the teacher encoder for backdoor defense only with the downstream data sets or a subset of the downstream data sets. That is, the training data set defined in the present application is the downstream data set or a subset of the downstream data set.

[0068] It can be understood that in a self-supervised backdoor attack, the pre-training phase establishes a weak association between the trigger feature and the target class semantic feature in the feature space through data poisoning, and forces the samples containing the trigger to be clustered into the target class region. At this time, when the downstream user fine-tunes the classification head on the clean data set, the feature space structure of the encoder remains fixed. At this time, the normal samples of the target class will strengthen the label mapping of the class feature region, and passively inherit the association between the trigger and the target class in the pre-training, and formally include the entire trigger feature region into the decision domain of the target class. The fine-tuning process solidifies and amplifies the backdoor binding using the clean data provided by the user, which converts the originally ambiguous association into an explicit mapping from the trigger to the target label, and instead improves the success rate of the backdoor attack.

[0069] Based on this, in order to obtain the teacher encoder through training to purify the to-be-processed encoder in subsequent operations, and at the same time not to amplify the backdoor binding process due to the data labels of the downstream data, the present application provides the following embodiments to decouple the training of the teacher encoder from the data labels of the downstream data, that is, to realize the unlabeled encoder training:

[0070] In an exemplary embodiment, at least two unlabeled enhanced data sets are generated, and a to-be-processed encoder is trained based on the enhanced data sets to obtain a teacher encoder, including: obtaining original images in the training data set, and modifying the topology of the original images to generate at least two corresponding enhanced images for each original image; based on the to-be-processed encoder, calculating the embedding vectors of each node in the enhanced images; based on the similarity of the embedding vectors of the same node in different enhanced images corresponding to the same original image, and / or the node similarity of the enhanced images corresponding to different original images, and / or the similarity of the embedding vectors of different nodes in different enhanced images corresponding to the same original image, a first loss function is constructed; based on the first loss function, the parameters of the to-be-processed encoder are updated until a preset condition is reached. The preset condition can be that the training round reaches a preset threshold, or the value of the loss function reaches a target threshold, and the embodiments of the present application do not limit this.

[0071] The topology of the original image, that is, the connection mode between the nodes in the original image, can be represented by an adjacency matrix. The way to modify the topology can include, but is not limited to, one or a combination of the following ways: adding edges that do not exist on the original image, deleting edges that exist on the original image, masking part of the feature vectors of the nodes, subgraph sampling, etc.

[0072] Further, based on the modified topology to obtain the enhanced image, what should be changed is the non-essential aspect of the original image, such as the specific connection details, part of the noise features, etc. The semantic structure information of the original image such as the overall topology, community structure, and node function role should be preserved, so that in graph contrast learning, the enhanced image and the original image can be regarded as different "perspectives" of the same thing, and have similar representations. Each of the enhanced data sets formed based on the enhanced image corresponds to one enhanced image of each original image, and the enhanced images corresponding to the same original image in different enhanced data sets are not exactly the same.

[0073] The similarity of the embedding vectors of the same node in different enhanced images corresponding to the same original image is included in the loss function, which can maximize the similarity of positive sample pairs; the node similarity of the enhanced images corresponding to different original images and the similarity of the embedding vectors of different nodes in different enhanced images corresponding to the same original image are included in the loss function, which can minimize the similarity of negative sample pairs. The positive sample pair, that is, the embedding of two enhanced images obtained by different enhancements of the same original image, and the negative sample pair is the combination of other enhanced images except the positive sample pair. Through this learning method, the model can understand "which graphs or nodes are similar (positive sample pair) and which are not (negative sample pair),

[0074] Specifically, in an embodiment, the first loss function can be defined as:

[0075]

[0076]

[0077] wherein, is an enhanced dataset; represents a loss of each contrastive instance . represents a cosine similarity of embedding and . is a temperature parameter; is a first loss function.

[0078] It should be noted that in the multi-round contrastive learning iteration, the cosine similarity between the original image and the enhanced image will decrease, resulting in a decrease in the performance of the teacher encoder. Therefore, in another embodiment of the present application, on the basis of the above learning mode, an adversarial learning can be introduced to enhance the performance of the teacher encoder, thereby laying a foundation for improving the performance of the distilled backdoor encoder.

[0079] In an exemplary embodiment, in the aforementioned image contrastive learning, the adversarial learning specifically includes: adding a disturbance to the original image to generate a corresponding adversarial image; constructing a second loss function based on the similarity of the same node embedding vectors in the same original image corresponding to the enhanced image and the adversarial image, and / or the node similarity of the enhanced image and the adversarial image corresponding to different original images, and / or the similarity of different node embedding vectors in different enhanced images and adversarial images corresponding to the same original image; updating the parameters of the to-be-processed encoder based on the second loss function until a preset condition is reached.

[0080] It should be noted that the adversarial image is a malicious sample generated by adding a small disturbance designed by man to the original input, which aims to induce the model to produce an error output. Unlike the enhanced image, the enhanced image is a sample generated under the premise of semantic invariance, which aims to improve the generalization ability of the model.

[0081] In an embodiment of the present application, considering the algorithm performance and time complexity, the node features of the adversarial image and the adjacency matrix of the adversarial image can be defined by adding a disturbance, specifically including:

[0082]

[0083]

[0084]

[0085]

[0086] wherein, denotes the node feature of the adversarial image ; denotes the adjacency matrix of the adversarial image ; denotes the node feature of the original image ; denotes the adjacency matrix of the original image ; is a complementary graph; is a matrix full of 1s, is an identity matrix, and A denotes an original adjacency matrix; is 1 only when there is no connection in the original graph and it is not a self-loop.

[0087] In another embodiment, the constraint condition can be further optimized on this basis and to control the adversarial noise added in the training process. The further optimized constraint condition varies with the iteration round and can be expressed as:

[0088]

[0089]

[0090]

[0091] wherein, denotes the perturbation to be added to the node feature matrix; is a symmetric matrix for modifying the edges in the graph, and the element value of 1 indicates modification and 0 indicates no change; denotes an element-wise multiplication operation for determining whether to add or delete an edge; is an element-wise sign function; and are hyperparameters affecting the generation of the t-th round of adversarial samples; and respectively denote projection operations on and .

[0092] On this basis, in combination with adversarial learning and the aforementioned image contrast learning, the second loss function can be defined as:

[0093]

[0094] wherein, is an enhanced data set; is a weight coefficient for balancing the two losses, which can be artificially adjusted based on the actual use scenario of the present application; is the first loss function; This is the second loss function.

[0095] Regarding the adversarial image, it should also be noted that, in order to ensure that the perturbations are not easily detected and to guarantee the effectiveness of the adversarial sample attack, in one embodiment of the present invention, the total magnitude of the modification of node features when generating the adversarial sample is further limited, so as to prevent the defender from detecting the attack through feature statistical anomalies. Specifically, the sum of the absolute values ​​of the perturbations added to each node in the adversarial image across all feature dimensions is less than a preset first threshold, and the sum of the number of connection state changes between each node in the adversarial image is less than a preset second threshold. The first threshold and the second threshold can be dynamically adjusted by the experimenter based on the specific experimental environment, and the embodiments of the present invention do not impose such limitations.

[0096] S202: Define the same attention operator in each corresponding layer of the teacher encoder and the encoder to be processed, and the attention operator is used to output the layer attention map of the corresponding encoder layer;

[0097] In one exemplary embodiment, for a graph self-supervised encoder The first The activation graph of the layer is represented as ,in Indicates the number of nodes. Indicates the first The dimension of the layer output channels. To accurately represent the contribution of each element to the classification result, a function is defined. This is used to map the activation graph to the attention representation. Based on this, a degree attribute is added to each node. This fully considers the inherent topological structure information in the graph data. In summary, the first... Layer attention operator Defined as:

[0098]

[0099] in Degree attribute representing all nodes The maximum value, Used to characterize the relationship between the degree attribute of each node and the maximum value of the degree attribute in the node; The activation graph of each channel is represented as follows: ; It is an absolute value function, ensuring Non-negative; This is used to amplify the distance between benign and malicious neurons, and is typically set in experiments. .

[0100] It should be noted that the degree attribute value of a node refers to the number of other nodes connected to that node in the topology. That is, if a node's degree attribute value is 3, then there are 3 other nodes connected to that node in the topology. Introducing the node degree attribute into the operator ensures that the semantic features of high-order nodes are enhanced and that the backdoor trigger region is highlighted due to topological inconsistencies.

[0101] The activation map is a spatial visualization of neuron responses in the hidden layers of a GNN encoder, used to quantify the response intensity of node features in the channel dimension. High activation regions represent neurons sensitive to specific features, while backdoor attacks can cause abnormal activation of trigger regions. In this invention, the activation map is the raw input for generating the attention map, which in turn serves as the alignment target for distillation.

[0102] S203: Calculate the distribution offset between the layer attention map of the teacher encoder and the layer attention map of the encoder to be processed in the same encoder layer;

[0103] In an exemplary embodiment, calculating the distribution offset of the layer attention map of the teacher encoder and the layer attention map of the encoder to be processed in the same encoder layer includes: normalizing the layer attention map output by each layer attention operator of the teacher encoder to obtain the corresponding layer attention weights; normalizing the layer attention map output by each layer attention operator of the encoder to be processed to obtain the corresponding layer attention weights; obtaining the difference between the layer attention weights of the teacher encoder and the encoder to be processed in the same encoder layer, and calculating the layer attention weight distribution difference between the teacher encoder and the encoder to be processed based on the L2 paradigm.

[0104] Normalizing the attention representations of each layer ensures that the attention representations of each layer remain consistent. This normalization can be expressed as:

[0105]

[0106] Based on the L2 paradigm, the difference in layer attention weight distribution between the teacher encoder and the encoder to be processed can be specifically expressed as follows:

[0107]

[0108] in and They represent the first Attention maps of the teacher encoder and the encoder to be processed in the layer. Through The distributional differences between them are calculated using the paradigm.

[0109] S204: Fix the parameters of the teacher encoder, update the parameters of the to-be-processed encoder based on the distribution offset.

[0110] On the basis of the above scheme, the teacher encoder parameters are fixed, the backdoor neurons in the student model are aligned with the clean neurons in the teacher model, and the backdoor purification can be realized. Specifically, the greater the difference between the layer attention weight distribution of the above-mentioned teacher encoder and the to-be-processed encoder, the more significant the abnormal activation of the backdoor neuron. Through the loss function (distribution difference of layer attention map) The student model (to-be-processed encoder) is forced to The attention distribution of the layer approximate the teacher encoder Since the teacher encoder does not respond to the backdoor trigger, after the teacher encoder with fixed parameters and the to-be-processed encoder attention are aligned, the to-be-processed encoder will also be forced to reduce the activation intensity of the trigger area. With the to-be-processed encoder enhancing the attention to semantic features (such as high nodes, feature clusters), the backdoor logic will be gradually covered, ensuring that the backdoor neurons are suppressed.

[0111] On this basis, an embodiment of the present application further introduces a BT loss, improves the performance of the purified encoder by calculating the cross-correlation matrix of the attention representation graph of the teacher encoder and the to-be-processed encoder. This embodiment can ensure the consistency of the diagonal elements and the independence of the non-diagonal elements of the features, so as to learn more useful information under different encoders. Further, based on the cross-correlation matrix and the layer attention weight distribution difference, a third loss function is constructed; the parameters of the to-be-processed encoder are updated based on the third loss function. Formally, it can be represented as:

[0112]

[0113] Among them, represents the cross-correlation matrix between the mth feature of the teacher encoder and the nth feature of the to-be-processed encoder; is the weight between controlling the consistency of the diagonal elements and the independence of the non-diagonal elements; is the third loss function.

[0114] Finally, the total loss of the backdoor removal algorithm can be defined as:

[0115]

[0116] Among them and are two hyperparameters for balancing the performance of the purified backdoor encoder, represents the number of layers of the encoder.

[0117] In order to more intuitively show the graph backdoor encoder defense method of the present application, a set of comparative data in a real experimental scene is also provided. The performance of the graph backdoor encoder defense method of the present application in performing node classification on different data sets is comprehensively evaluated

[0118] In the comparative experiment, four real-world data sets are used to evaluate the defense method of the present application in the node classification task, and the data sets include Cora, CiteSeer, Pubmed and Flickr. The specific information of the above data sets is shown in Table I.

[0119] Table I: Analysis of node classification data sets

[0120]

[0121] In the experiment, five most advanced graph backdoor attack methods in the self-supervised scene are selected, including GRBA, GCBA, CGBR, ECGBA and CLGA. At the same time, four SOTA defense methods are compared, including BloGBaD, E-SAGE, SimGuard and DShield. It is worth noting that since there is no defense method specially designed for graph backdoor attacks in the self-supervised scene, the above graph backdoor defense techniques that can be transferred from the supervised scene to the self-supervised scene are selected and evaluated. The evaluation index is defined as attack success rate (ASR) and accuracy (ACC). ASR is an attack evaluation index, which is used to evaluate the proportion of poisoned samples that are successfully misclassified as target labels by the attack method. In the comparative experiment, the effectiveness of the defense attack is measured by this index. ACC is a measure used to evaluate the performance of the model. Since backdoor attacks or defenses may affect the performance of the model. Therefore, ACC is used to measure the performance of the model after defense.

[0122] In the comparative experiment, the graph backdoor encoder defense method provided by the present application is implemented based on the Python language using the PyTorch framework, and the experimental environment is configured as: Intel(R) Xeon(R) Silver 4310 CPU 377G (CPU), NVIDIA RTXA6000 (GPU), 75GiB memory and Ubuntu 20.04 operating system. DGI is used as the pre-training model. In the training stage, 20% of the data are randomly selected as the training set, and 5% of the data are injected into the trigger as backdoor samples for training. The remaining 80% of the data are used as the test set, of which only 5% of the data are selected for defense. The defense performance of the graph backdoor encoder defense method of the present application and the SOTA defense method on the SOTA self-supervised graph backdoor attack on different data sets is shown in Table II.

[0123] Table II: Comparison of defense performance of GDetox and SOTA methods

[0124]

[0125] The experimental results show that the present application achieves lower ASR than all SOTA defenses, indicating its advantage in mitigating various backdoor attacks in the graph self-supervised scenario. At the same time, the model accuracy (ACC) of the present application also exceeds all SOTA defenses, proving that the present application maintains strong model performance while providing effective defense. The accuracy of the present application is within 2% of the difference between the Before models, and some of the After-defense models actually perform better than the attacked models.

[0126] Referring to FIG. 3, Figure 3 Based on the same inventive concept as the aforementioned graph backdoor encoder defense method, an embodiment of the present application provides a graph backdoor encoder defense system 300, which includes a training module 301, a definition module 302, a calculation module 303, and an update module 304.

[0127] Specifically, the training module 301 is configured to generate at least two unlabeled augmented data sets based on a preset training data set, and train a to-be-processed encoder based on the augmented data sets to obtain a teacher encoder, wherein the training data set is a downstream data set or a subset of the downstream data set; the definition module 302 is configured to define the same attention operator in the corresponding layers of the teacher encoder and the to-be-processed encoder, and the attention operator is used to output the layer attention map of the corresponding encoder layer; the calculation module 303 is configured to calculate the distribution offset of the layer attention map of the teacher encoder and the layer attention map of the to-be-processed encoder in the same encoder layer; and the update module 304 is configured to fix the parameters of the teacher encoder, and update the parameters of the to-be-processed encoder based on the distribution offset.

[0128] Referring to FIG. 4, Figure 4 The present application further provides an electronic device 400, which includes at least one processor 401, a memory 402 (for example, a non-volatile memory), a memory 403, and a communication interface 404, and the at least one processor 401, the memory 402, the memory 403, and the communication interface 404 are connected together via an internal bus 405. The at least one processor 401 is configured to invoke at least one program instruction stored or encoded in the memory 402, so as to enable the at least one processor 401 to perform various operations and functions of the graph backdoor encoder defense method described in various embodiments of the present specification.

[0129] In embodiments of the present specification, the electronic device 400 can include, but not limited to, a personal computer, a server computer, a workstation, a desktop computer, a laptop computer, a notebook computer, a mobile electronic device, a smart phone, a tablet computer, a cellular phone, a personal digital assistant (PDA), a handheld device, a messaging device, a wearable electronic device, a consumer electronic device, and the like.

[0130] The present application also provides a computer readable medium, which carries computer execution instructions, and the computer execution instructions can be used to implement various operations and functions of the post-GAN encoder defense method described in various embodiments of the present specification when executed by a processor.

[0131] The computer readable medium in the present application can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium may, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any combination of the above. More specific examples of the computer readable storage medium can include, but not limited to, an electrical connection with one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or apparatus.

[0132] In the present application, the computer readable signal medium can include a data signal propagating in a baseband or as a carrier wave in a propagated data signal, which carries computer readable program code. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer readable signal medium can also be any computer readable medium that can send, propagate or transfer program for use by or in connection with an instruction execution system, device or apparatus, other than the computer readable storage medium. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, or the like, or any suitable combination thereof.

[0133] Those skilled in the art will appreciate that embodiments of the present application can be devised for a variety of applications. It is intended that the present application covers all such applications of the embodiments disclosed herein, whether or not the specific data is disclosed. It is intended that the present application encompasses all such embodiments falling within the scope of the appended claims and their equivalents. It will be apparent to those skilled in the art that substantial variations can be made in accordance with specific detailed embodiments described herein without departing from the spirit or scope of the application. Therefore, the scope of the present application is not intended to be limited to the particular embodiments described herein, but is only limited by the scope of the appended claims.

[0134] The present application is described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus, systems, and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer, special purpose computer, embedded processing system, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in the flowchart illustrations and / or block diagrams. Figure 1 one or more functions specified in one or more of the flowchart illustrations and / or block diagrams. Figure 1 an apparatus to perform the functions specified in one or more of the flowchart illustrations and / or block diagrams.

[0135] The foregoing description of specific exemplary embodiments of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed, and various modifications and variations are possible in light of the above teachings. It is intended that the embodiments be limited only by the claims as interpreted in accordance with the principles of patent law, including the principles of equivalents. It is intended that the scope of the application encompass all technical equivalents.

[0136] It will be apparent to those skilled in the art that the present application is not limited to the specific embodiments described herein, which are intended as illustrative only, and that numerous modifications can be made therein without departing from the spirit or scope of the application. It is intended that the scope of the application encompass all such modifications and alterations as fall within the scope of the appended claims and their equivalents. No limitation is intended to the details of construction or design herein shown, other than as described in the claims below, and their equivalents. No limitation is intended to the diagrams of the claims, other than as described in the claims below.

[0137] Furthermore, it should be understood that although the specification is described in terms of embodiments, not every embodiment includes every feature described. The specification can include implicit combinations of explicitly mentioned features and / or explicit combinations of implicitely mentioned features. Each embodiment depends on the explicit combinations of features and / or the implicit combinations of features made specifically within that embodiment, and each such embodiment can be combined with every other such embodiment to create further embodiments.

Claims

1. A method for defending a graph backdoor encoder, applied to graph self-supervised learning, characterized in that, The method comprises: generating at least two unlabeled enhanced data sets based on a preset training data set, and training a to-be-processed encoder based on the enhanced data sets to obtain a teacher encoder, wherein the training data set is a downstream data set or a subset of the downstream data set; defining the same attention operator in the corresponding layers of the teacher encoder and the to-be-processed encoder, the attention operator being used to output a layer attention map of a corresponding encoder layer; calculating the distribution offset of the layer attention map of the teacher encoder and the layer attention map of the to-be-processed encoder in the same encoder layer; fixing the parameters of the teacher encoder, updating the parameters of the to-be-processed encoder based on the distribution offset.

2. The postgate encoder defense method of claim 1, wherein, Generating at least two unlabeled enhanced data sets and training a to-be-processed encoder based on the enhanced data sets to obtain a teacher encoder comprises: obtaining original images in the training data set, and modifying the topological structure of the original images to generate at least two corresponding enhanced images for each original image; calculating the embedding vector of each node in the enhanced image based on the to-be-processed encoder; constructing a first loss function based on the similarity of the embedding vectors of the same node in different enhanced images corresponding to the same original image, and / or the node similarity of the enhanced images corresponding to different original images, and / or the similarity of the embedding vectors of different nodes in different enhanced images corresponding to the same original image; updating the parameters of the to-be-processed encoder based on the first loss function until a preset condition is reached.

3. The postgate encoder defense method of claim 2, wherein, The method further comprises: adding perturbations to the original images to generate corresponding adversarial images; wherein the sum of the absolute values of the perturbations added to each node in all feature latitudes in the adversarial image is less than a preset first threshold, and the total sum of the number of changes in the connection state between nodes in the adversarial image is less than a preset second threshold; constructing a second loss function based on the similarity of the embedding vectors of the same node in the enhanced images and the adversarial images corresponding to the same original image, and / or the node similarity of the enhanced images and the adversarial images corresponding to different original images, and / or the similarity of the embedding vectors of different nodes in different enhanced images and the adversarial images corresponding to the same original image updating the parameters of the to-be-processed encoder based on the second loss function until a preset condition is reached.

4. The postgate encoder defense method of claim 3, wherein, The first loss function and the second loss function are respectively: wherein, is a loss for augmented dataset; is a loss for augmented dataset; is a loss for each contrastive instance is a cosine similarity of embedding and ; is a temperature parameter; is a weight coefficient for balancing two losses; is a first loss function; is a second loss function.

5. The postgate encoder defense method of claim 3, wherein, The method further comprises updating the relevant features of the adversarial image in each training round: where, represents the node features of the adversarial image ; represents the adjacency matrix of the adversarial image ; represents the node features of the original image ; represents the adjacency matrix of the original image ; represents the perturbation to be added to the node feature matrix; is a symmetric matrix for modifying the edges in the graph, and the element value of 1 represents modification and 0 represents no change; represents an element-wise multiplication operation for determining whether to add or delete an edge; is an element-wise sign function; and are hyperparameters affecting the generation of the t-th round of adversarial samples; and respectively represent the projection operations on and ; is a complementary graph; is a matrix of all 1s, is an identity matrix, and A represents the original adjacency matrix; is 1 only when there is no connection in the original graph and it is not a self-loop.

6. The postgate encoder defense method of claim 1, wherein, The calculation of the distribution offset of the layer attention map of the teacher encoder and the layer attention map of the to-be-processed encoder in the same encoder layer comprises: normalizing the layer attention map output by each layer attention operator of the teacher encoder to obtain corresponding layer attention weights; normalizing the layer attention map output by each layer attention operator of the to-be-processed encoder to obtain corresponding layer attention weights; obtaining the difference between the layer attention weights of the teacher encoder and the to-be-processed encoder in the same encoder layer, and calculating the distribution difference between the layer attention weights of the teacher encoder and the to-be-processed encoder based on the L2 norm.

7. The postgate encoder defense method of claim 6, wherein, The updating the to-be-processed encoder parameter based on the distribution offset comprises: inputting an original image, calculating a cross-correlation matrix of the teacher encoder and the to-be-processed encoder output feature; constructing a third loss function based on the cross-correlation matrix and the layer attention weight distribution difference; updating the to-be-processed encoder parameter based on the third loss function.

8. The postgate encoder defense method of claim 7, wherein, The third loss function comprises: wherein, represents the cross-correlation matrix between the mth feature of the teacher encoder and the nth feature of the encoder to be processed; is the weight between the consistency of the diagonal elements and the independence of the off-diagonal elements; is the third loss function.

9. A computer device, comprising: comprises: a memory and a processor, which are connected in communication with each other, the memory stores computer instructions, and the processor executes the computer instructions to perform the image backdoor encoder defense method of any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing a computer to execute the image backdoor encoder defense method of any one of claims 1-7.

Citation Information

Cited By

  • Back door defense method and device, image classification method and device, equipment and medium

    CN122313237A

  • A backdoor defense method, an image classification method, an apparatus, a device and a medium

    CN122313237B