A method, system, medium, and apparatus for detecting poisoning in a data sample

CN122693014APending Publication Date: 2026-09-04ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004427.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-04

AI Technical Summary

Technical Problem

该类方法的局限性在于:其一,多依赖攻击先验知识(如触发器形态、投毒比例)或人工设计的判别规则,面对未知新型隐蔽攻击时适应性较差;其二,大多需要额外获取独立于训练集的干净验证数据集,而实际工业场景中往往难以满足该条件;其三,多采用单层激活特征或单一维度信息进行判别,无法充分挖掘中毒样本在不同隐藏层的微弱异常响应,跨层特征融合能力不足,在低投毒率、特征高度重叠场景下检测精度显著下降

Benefits of technology

[0022] The beneficial effects of this invention are as follows: Compared with the prior art, this invention extracts the activation maps of multiple hidden layers of a deep neural network and performs cross-layer analysis to extract the activation maps of poisoned samples under the condition of no prior attack; secondly, by normalizing the activation maps of different hidden layers to a unified numerical scale, the internal covariate shift and numerical dimension difference between different layers are eliminated; furthermore, by extracting extreme values ​​along the spatial dimension as abnormal features, the masked abnormal activations are accurately captured in scenarios where the features of poisoned samples and benign samples highly overlap; further still, by introducing a particle swarm optimization algorithm and using the micro-cluster stability of HDBSCAN clustering as the fitness function to inversely optimize the layer attention weights, the optimal weight distribution for different attack modes is adaptively learned under the condition of no additional clean validation set; finally, by re-weighting with the optimal weights and combining threshold and stability dual conditions to screen abnormal micro-clusters, the extremely small proportion of poisoned samples can be accurately separated from the large benign main clusters in scenarios with low poisoning rates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122693014A_ABST
    Figure CN122693014A_ABST
Patent Text Reader

Abstract

The application discloses a data sample poisoning detection method, system, medium and equipment, and relates to the technical field of artificial intelligence security.The application normalizes and aligns a multi-hidden layer activation map of a deep neural network, extracts channel extreme values along a spatial dimension, and constructs a sample cross-layer abnormal feature vector; a particle swarm optimization algorithm is introduced, attention weights of each hidden layer are modeled as particle positions in a continuous search space, and an abnormal micro-cluster stability score output by HDBSCAN after low-dimensional mapping is used as a fitness index; under the condition of no attack prior and no clean validation set, optimal layer attention weight distribution is adaptively searched through population cooperative iteration, cross-layer abnormal consistency of poisoned samples is accurately amplified, abnormal micro-clusters are finally locked by combining scale threshold and stability double constraints, accurate separation of poisoned samples and benign main clusters is realized under a low poisoning rate, and training data purification precision and model security performance are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence security technology, and in particular to a method, system, medium, and device for detecting data sample poisoning. Background Technology

[0002] In recent years, deep learning technology has played a core supporting role in various intelligent systems due to its powerful feature extraction and nonlinear modeling capabilities. However, the performance of deep learning models is highly dependent on large-scale, high-quality training data. In practical applications, training data is often obtained through public datasets, third-party platforms, and multi-party collaborative collection, resulting in complex data flow links and difficulty in fully guaranteeing the reliability of the source. Attackers can inject a small number of carefully crafted poisoned samples with triggers into the training set, inducing the model to learn the hidden correlation between triggers and attack target labels during training: the model can maintain high classification accuracy under normal input, but will output incorrect results preset by the attacker once a sample with triggers is input. This type of data poisoning backdoor attack is characterized by strong concealment, high success rate, and high detection difficulty, and has become one of the core threats to the secure implementation of deep learning.

[0003] To address the aforementioned issues, the industry has proposed various backdoor defense solutions, which can be categorized into three types based on their operational phase: The first category is dataset-level defense methods, which are applied before model training. They achieve data purification by detecting and removing poisoned samples from the training set. Typical solutions include spectral feature detection methods based on singular value decomposition, activation clustering (AC) methods, and frequency domain analysis methods based on discrete cosine transform. The limitations of this type of method are: first, they often rely on prior knowledge of the attack (such as trigger morphology and poisoning ratio) or manually designed discrimination rules, making them less adaptable to unknown and novel covert attacks; second, most require additional clean validation datasets independent of the training set, which is often difficult to meet in real-world industrial scenarios; and third, they often use single-layer activation features or single-dimensional information for discrimination, failing to fully exploit the weak anomalous responses of poisoned samples in different hidden layers, exhibiting insufficient cross-layer feature fusion capabilities, and resulting in a significant decrease in detection accuracy in scenarios with low poisoning rates and highly overlapping features.

[0004] The second category is input-level defense methods, which operate during the model inference stage. They block backdoor activation by detecting and filtering input samples carrying triggers. Typical solutions include randomness detection methods based on image overlay, SentiNet methods based on Grad-CAM, and defense methods based on geometric spatial transformations. The limitations of this type of method are: it can only handle scenarios with visible triggers; its effectiveness against triggers without obvious trigger features or dynamically hidden triggers is limited; and it cannot fundamentally purify the training data, leaving the risk of backdoors being implanted into the model.

[0005] The third category is model-level defense methods, which are implemented after model training. Backdoor removal is achieved by pruning abnormal neurons, adjusting the model structure, or constructing a meta-classifier. Typical solutions include Fine-Pruning pruning methods targeting backdoor-related neurons and meta-classifier methods based on universal litmus detection. The limitations of this type of method are: complex deployment process, high computational cost, and the potential to disrupt the original parameter structure of the model, leading to decreased normal classification performance and difficulty in adapting to complex and ever-changing attack scenarios.

[0006] Therefore, under the conditions of no attack prior, no additional clean validation set, low poisoning rate, and high overlap between the characteristics of poisoned samples and benign samples, how to eliminate the internal covariate shifts and numerical dimension differences between different layers, accurately capture the masked abnormal activations, and adaptively learn the optimal weight distribution of different attack modes, so as to accurately separate the very small number of poisoned samples from the huge benign main cluster, has become a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0007] Therefore, it is necessary to provide a method, system, medium, and device for detecting poisoning in data samples to address the aforementioned technical problems.

[0008] The following technical solution is adopted in this specification: This specification provides a method for detecting poisoning in data samples, specifically including: Obtain an untrusted training dataset and a deep neural network trained on the untrusted training dataset; extract the activation maps of multiple hidden layers of the deep neural network for each sample in the untrusted training dataset, and normalize the activation maps of different hidden layers to a uniform numerical scale to obtain the aligned feature map of each hidden layer.

[0009] For each hidden layer's aligned feature map, extract extreme values ​​along the spatial dimension by channel, and use these extreme values ​​as the abnormal features of that channel; concatenate the abnormal features of each channel in the current activated map to obtain the hidden layer channel abnormal concatenation features; concatenate the concatenation features of all hidden layers to generate the cross-layer extreme value feature vector for each sample.

[0010] A particle swarm optimization algorithm is employed, using the number of hidden layers as the search space dimension. The attention weights of all hidden layers are mapped to particle positions within the search space. For each particle, the attention weights corresponding to that particle are used to weight and concatenate the cross-layer extreme feature vectors of all samples, resulting in a weighted feature matrix. After dimensionality reduction of the weighted feature matrix, density-based hierarchical clustering is performed to filter out anomalous micro-clusters whose sample proportion in the training set is less than a preset threshold. The stability value with the highest stability score in the anomalous micro-cluster set is used as the fitness value of that particle. The particle positions are iteratively updated until convergence, yielding the globally optimal particle positions as the optimal layer attention weight distribution.

[0011] The optimal layer attention weight distribution is used to reweight and concatenate the cross-layer extreme feature vectors of all samples to obtain the final weighted feature matrix. After dimensionality reduction of the final weighted feature matrix, density-based hierarchical clustering is performed to obtain a secondary screening abnormal micro-cluster set. The sample with the highest stability score in the secondary screening abnormal micro-cluster set, whose proportion of the training set is less than a preset threshold, is identified as a poisoned sample and output.

[0012] Furthermore, the method for normalizing the activation maps of different hidden layers to a uniform numerical scale to obtain the aligned feature map of each hidden layer is as follows: Regarding the first Hidden layer Channel 1 The original activation values ​​of the activation map of the hidden layer at the location, and the moving average of the batch-normalized values ​​of that hidden layer. and variance Normalization is performed to obtain the normalized activation response value. ; ; in, For the first Layer Channel 1 The original activation value of the location. To prevent division by zero smoothing constants.

[0013] The normalized activation response values ​​of all channels of the hidden layer are combined into an aligned feature map.

[0014] Furthermore, for each hidden layer's aligned feature map, along the spatial dimension, extreme values ​​are extracted by channel as the abnormal features of that channel. Specifically, for each channel of the activation map of each hidden layer, the global minimum value of the response values ​​at all spatial locations on the feature map of that channel is extracted as the abnormal feature of that channel.

[0015] Furthermore, the particle swarm optimization algorithm employs a linearly decreasing inertia weight strategy, specifically as follows: ; in, For the first The inertia weight of the next iteration, and These are the maximum inertia weight and the minimum inertia weight, respectively. This represents the maximum number of iterations.

[0016] Furthermore, the step of using the stability value with the highest stability score in the abnormal micro-cluster set as the fitness value of the particle specifically means: for any particle in each iteration, if its corresponding abnormal micro-cluster set is not empty, then the score corresponding to the micro-cluster with the highest stability evaluation in the set is selected as the fitness value of the particle; if the abnormal micro-cluster set is empty, then the fitness value of the particle is set to zero.

[0017] Furthermore, for each particle, the attention weight corresponding to that particle is used to weight and concatenate the cross-layer extreme feature vectors of all samples to obtain a weighted feature matrix. Specifically, for each particle's current attention weight, the attention weight is multiplied sequentially by the cross-layer extreme feature vector of the corresponding layer and then concatenated to form a weighted feature matrix of all samples. The multiple hidden layers are selected from at least two convolutional layers of different depths in the deep neural network.

[0018] Furthermore, the untrusted training dataset is an autonomous driving domain dataset or a medical imaging dataset containing at least one attack mode; the attack mode includes at least All-to-One, A2O, All-to-All, A2A, or Untargeted, UT; the poisoned sample is a sample with a trigger injected into the untrusted training set; the label of the poisoned sample is tampered with to an attack target label associated with the trigger, so that the deep neural network establishes a hidden mapping relationship between the trigger features and the attack target label during training; when the untrusted training set is an autonomous driving domain dataset, the backdoor trigger of the poisoned sample is an incorrect traffic signal instruction; when the untrusted training set is a medical imaging dataset, the backdoor trigger of the poisoned sample is an incorrect lesion classification result.

[0019] Furthermore, the preset threshold ranges from greater than 0.05 to less than 0; when screening micro-clusters, only clusters containing a number of samples less than the product of the preset threshold and the total number of samples in the training set are identified as micro-clusters that meet the size condition, thereby excluding the main cluster composed of a large number of benign samples.

[0020] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described above.

[0021] This specification provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described above.

[0022] The beneficial effects of this invention are as follows: Compared with the prior art, this invention extracts the activation maps of multiple hidden layers of a deep neural network and performs cross-layer analysis to extract the activation maps of poisoned samples under the condition of no prior attack; secondly, by normalizing the activation maps of different hidden layers to a unified numerical scale, the internal covariate shift and numerical dimension difference between different layers are eliminated; furthermore, by extracting extreme values ​​along the spatial dimension as abnormal features, the masked abnormal activations are accurately captured in scenarios where the features of poisoned samples and benign samples highly overlap; further still, by introducing a particle swarm optimization algorithm and using the micro-cluster stability of HDBSCAN clustering as the fitness function to inversely optimize the layer attention weights, the optimal weight distribution for different attack modes is adaptively learned under the condition of no additional clean validation set; finally, by re-weighting with the optimal weights and combining threshold and stability dual conditions to screen abnormal micro-clusters, the extremely small proportion of poisoned samples can be accurately separated from the large benign main clusters in scenarios with low poisoning rates. Attached Figure Description

[0023] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0024] Figure 1 This is a schematic diagram of the overall framework of an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall process of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort are within the scope of protection of this invention.

[0026] This invention provides a method for detecting data sample poisoning, comprising four steps: acquiring and normalizing the activation maps of hidden layers in a trained deep neural network for each sample in an untrusted dataset; extracting extreme values ​​along the spatial dimension as anomalous features and generating cross-layer extreme value feature vectors; optimizing attention weights using a particle swarm optimization algorithm; and re-weighting the data using the optimal weights and combining threshold and stability conditions to screen anomalous micro-clusters, thereby identifying poisoned samples and outputting the results. The method involves extracting activation maps from multiple hidden layers of the deep neural network and performing cross-layer analysis to extract activation maps of poisoned samples under conditions without prior attack. Furthermore, by normalizing the activation maps of different hidden layers to a uniform numerical scale, the method eliminates inter-layer interference. The study examines the internal covariate shifts and numerical dimension differences. Then, by extracting extreme values ​​along the spatial dimension as anomalous features, it accurately captures masked anomalous activations in scenarios where the features of poisoned and benign samples highly overlap. Furthermore, by introducing a particle swarm optimization algorithm, using the micro-cluster stability of HDBSCAN clustering as the fitness function to optimize the attention weights of the inverse optimization layer, it adaptively learns the optimal weight distribution for different attack modes without an additional clean validation set. Finally, by reweighting with the optimal weights and combining threshold and stability conditions to screen anomalous micro-clusters, it can accurately separate a very small percentage of poisoned samples from the large benign main cluster in scenarios with low poisoning rates.

[0027] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0028] A method for detecting poisoning in data samples, such as Figure 2 The specific steps shown are as follows: Step 1: Obtain the untrusted training dataset and the deep neural network trained on the untrusted training dataset.

[0029] Step 2: Extract the activation maps of multiple hidden layers of the deep neural network for each sample in the untrusted training dataset, and normalize the activation maps of different hidden layers to a uniform numerical scale to obtain the aligned feature map of each hidden layer.

[0030] Specifically, the method for normalizing the activation maps of different hidden layers to a uniform numerical scale to obtain the aligned feature map of each hidden layer is as follows: Regarding the first Hidden layer Channel 1 The original activation values ​​of the activation map of the hidden layer at the location, and the moving average of the batch-normalized values ​​of that hidden layer. and variance Normalization is performed to obtain the normalized activation response value. ; ; in, For the first Layer Channel 1 The original activation value of the location. To prevent division by zero smoothing constants.

[0031] The normalized activation response values ​​of all channels of the hidden layer are combined into an aligned feature map.

[0032] Step 3: For the aligned feature map of each hidden layer, extract the extreme values ​​along the spatial dimension by channel as the abnormal features of that channel; concatenate the abnormal features of each channel of the current activation map to obtain the hidden layer channel abnormal concatenation features; concatenate the concatenation features of all hidden layers to generate the cross-layer extreme value feature vector of each sample.

[0033] Furthermore, for each hidden layer's aligned feature map, along the spatial dimension, extreme values ​​are extracted by channel as the abnormal features of that channel. Specifically, for each channel of the activation map of each hidden layer, the global minimum value of the response values ​​at all spatial locations on the feature map of that channel is extracted as the abnormal feature of that channel.

[0034] Step 4: Employ the Particle Swarm Optimization (PSO) algorithm, using the number of hidden layers as the search space dimension, and map the attention weights of all hidden layers to particle positions within the search space. For each particle, use its corresponding attention weight to weighted concatenate the cross-layer extreme feature vectors of all samples to obtain a weighted feature matrix. After dimensionality reduction of the weighted feature matrix, perform density-based hierarchical clustering to filter out anomalous micro-clusters whose sample proportion in the training set is less than a preset threshold. Use the stability value with the highest stability score in the anomalous micro-cluster set as the fitness value of the particle. Iteratively update the particle positions until convergence, obtaining the globally optimal particle positions as the optimal layer attention weight distribution.

[0035] Preferably, the particle swarm optimization algorithm employs a linearly decreasing inertia weight strategy, specifically: ; in, For the first The inertia weight of the next iteration, and These are the maximum inertia weight and the minimum inertia weight, respectively. This represents the maximum number of iterations.

[0036] Specifically, the fitness value of a particle is determined by the highest stability score among the abnormal micro-clusters: for any particle in each iteration, if its corresponding abnormal micro-cluster set is not empty, the score corresponding to the micro-cluster with the highest stability evaluation in that set is selected as the fitness value of the particle; if the abnormal micro-cluster set is empty, the fitness value of the particle is set to zero.

[0037] Furthermore, for each particle, the attention weight corresponding to that particle is used to weight and concatenate the cross-layer extreme feature vectors of all samples to obtain a weighted feature matrix. Specifically, for the current attention weight of each particle, the attention weight is multiplied by the cross-layer extreme feature vector of the corresponding layer in turn and then concatenated to form the weighted feature matrix of all samples. Multiple hidden layers are selected from at least two convolutional layers of different depths in the deep neural network.

[0038] Step 5: Reweight and concatenate the cross-layer extreme feature vectors of all samples using the optimal layer attention weight distribution to obtain the final weighted feature matrix; after dimensionality reduction of the final weighted feature matrix, perform density-based hierarchical clustering to obtain a secondary screening set of abnormal micro-clusters; identify the samples in the secondary screening set of abnormal micro-clusters that account for less than the preset threshold and have the highest stability score as poisoned samples and output them.

[0039] Furthermore, the untrusted training dataset is an autonomous driving domain dataset or a medical image dataset containing at least one attack mode; the attack modes include at least All-to-One, A2O, All-to-All, A2A, or Untargeted, UT; the poisoned sample is a sample with triggers injected into the untrusted training set; the label of the poisoned sample is tampered with to the attack target label associated with the trigger, so that the deep neural network establishes a hidden mapping relationship between the trigger features and the attack target label during the training process; when the untrusted training set is an autonomous driving domain dataset, the backdoor trigger of the poisoned sample is an incorrect traffic signal instruction; when the untrusted training set is a medical image dataset, the backdoor trigger of the poisoned sample is an incorrect lesion classification result.

[0040] Preferably, the preset threshold ranges from greater than 0.05 to less than 0; when screening micro-clusters, only clusters containing a number of samples less than the product of the preset threshold and the total number of samples in the training set are identified as micro-clusters that meet the size condition, thereby excluding the main cluster composed of a large number of benign samples.

[0041] Example The overall framework of this embodiment is as follows: Figure 1 As shown, the specific steps are as follows: S1: Establish a threat model Attacker Model: This model assumes an attacker, acting as a potential data provider or data contaminant, who can access and modify a portion of the training data, injecting a certain proportion of poisoned samples into the original training set. The attacker can freely design triggers based on the attack target, including visible triggers, hidden triggers, or dynamic triggers, and can also tamper with the labels of poisoned samples, thus constructing different types of backdoor attacks. However, the attacker cannot interfere with the defender's subsequent model training process and cannot directly modify the model structure, training algorithm, optimizer settings, or data cleansing process. To cover common backdoor attack scenarios, this paper considers three typical attack paradigms: All-to-One (A2O), All-to-All (A2A), and Untargeted (UT).

[0042] Defender Model: This model assumes the defender is the trainer of the model and can only obtain potentially contaminated, untrusted training datasets. The defender has no prior knowledge of the existence of backdoors, the specific form of triggers, the attack target labels, or the locations of poisoned samples, and does not rely on additional clean validation sets. The defender can train the model according to the normal process and use the feature responses of training samples in the model's hidden layers to detect, filter, and clean the dataset, but does not pre-determine the attack type or the proportion of poisoned samples.

[0043] S2: Extreme Value Feature Extraction In deep neural networks, hidden features at different layers exhibit significantly different response patterns to benign semantics and backdoor triggers. To effectively suppress redundant background noise and accurately amplify weak anomalous activations induced by backdoors, a cross-layer latent representation extraction mechanism is designed. Addressing the issues of inconsistent numerical scales between layers and high-dimensional feature redundancy and noise interference, two steps are performed sequentially: feature distribution alignment and outlier extraction. In the alignment step, the feature maps are normalized to a uniform scale based on the statistical data of the Batch Normalization (BN) layer. Then, an outlier is extracted from each feature map as its representative, and these are aggregated across all hidden layers to form the final latent representation. Specific details are as follows:

[0044] S21: Feature Alignment To eliminate internal covariate shifts between activations of different layers and achieve effective fusion of cross-layer features at a uniform scale, batch normalization (BN) is used to align the outputs of each convolutional layer. Let the training dataset contain... One sample, Indicates the first There are training samples. Let the deep network select... Several candidate hidden layers participate in cross-layer feature fusion, among which Indicates the hidden layer number. Overall deep neural network. Represented in cascade form:

[0045] ; in, Indicates the final classification layer. Indicates the first A hidden layer typically consists of a convolutional layer, a batch normalization (BN) layer, and an activation function. For the input sample Let it be in the first The original output activation map of each convolutional layer is as follows: ,in For the number of channels, and These represent the height and width of each feature map, respectively.

[0046] Because the activation values ​​of different layers exhibit severe internal covariate shift, the global moving average of the corresponding BN layer is used. and variance Align the activation maps to a uniform scale to obtain normalized activation response values. : ; in, For the first Layer Channel 1 The original activation value of the location. To prevent the smoothing constant from being divided by zero, the above transformation aligns the features of different hidden layers to a space with approximately zero mean and unit variance on a numerical scale, thus providing a unified scale basis for subsequent cross-layer feature comparison and fusion.

[0047] S22: Extracting abnormal features: Backdoor triggers typically induce anomalous responses in specific neurons or channels, often appearing as extreme values ​​deviating from the normal semantic distribution in the standardized feature map. To capture these weak but discriminative backdoor signals, channel extrema are extracted along the spatial dimension of the aligned feature map as a representation of the anomalous response in that channel. Specifically, the focus is on the anomalous inhibitory responses induced by the triggers in certain channels. Observations show that backdoor triggers often elicit significantly lower inhibitory responses than normal samples (i.e., negative anomalous responses) in certain channels; therefore, the global minimum value for each channel is selected as the anomalous feature. If the trigger might induce positive anomalous responses, the maximum value can be used symmetrically.

[0048] For the The sample at the th Feature maps aligned with hidden layers , its first Representative anomaly characteristics of each channel: ; in, Indicates the first The sample at the th The first hidden layer, the... Extreme value anomaly features extracted from each channel This represents the activation response value after BN alignment. and These represent the spatial location indices in the feature map of that channel.

[0049] The first The extreme value features of all channels in the layer are concatenated in channel order to obtain the sample. In the Anomaly feature vectors on hidden layers: ; in, Indicates the first The sample at the th Abnormal extreme value feature vectors on hidden layers Indicates the first The feature dimensions obtained after extreme value extraction from each hidden layer For the first The number of channels in each hidden layer. Due to the feature map (spatial dimension) for each channel... Extracting the global minimum along the spatial dimension, that is, retaining only one scalar value for each channel, thus the feature dimension of this layer is... Equal to the number of channels in this layer .

[0050] Then, the sample In all The abnormal feature vectors from the candidate hidden layers are concatenated to obtain an unweighted cross-layer latent representation: ; in, Indicates the first Unweighted cross-layer latent representations of individual samples This represents a vector concatenation operation. The total feature dimension after cross-layer concatenation is represented by the following formula: ; in, for The feature dimensions extracted from each hidden layer This represents the total number of hidden layers.

[0051] Therefore, each training sample can be represented as a latent feature vector composed of multiple layers of anomalous extreme responses. This representation preserves anomalous activation information in different hidden layers that may be related to backdoor triggers, and provides a feature basis for subsequent layer attention weight search based on particle swarm optimization.

[0052] S3: Layer Attention Weight Optimization Layer attention weight optimization is essentially a gradient-free search problem in a continuous space, making it difficult to optimize directly using explicit supervised labels or differentiable objective functions. Particle Swarm Optimization (PSO), which does not rely on gradient information, can collaboratively search through individual and collective experience among particles, gradually converging to a better weight configuration while maintaining global exploration capabilities. This makes it suitable for adaptive layer weight optimization tasks under conditions of no clean validation set and no attack priors. Therefore, we further introduce a gradient-free layer attention optimization mechanism based on PSO, modeling the attention weights of different hidden layers as particle coordinates in a continuous search space. Through cluster stability feedback, we adaptively find the optimal layer weights that best amplify the abnormal response of poisoned samples.

[0053] S31: Particle Coding Initialization The attention weights of each hidden layer are modeled as follows: The coordinates of particles in a continuous search space. Let the particle swarm size be... , No. The particle in the first The position vector in the next iteration is represented as a set of candidate layer attention weight configurations:

[0054] ; in, Indicates the first The particle in the first In the nth iteration, it is assigned to the ... The attention weights of each hidden layer. These weights are layer weights shared across all samples.

[0055] In the initial stage of the algorithm (i.e.) To ensure sufficient global exploration diversity for the particle swarm throughout the high-dimensional attention space, a uniform random distribution is used to unbiasedly initialize the initial positions of all particles. This set of initialized particle coordinates forms the first set of initial weights for the layer attention mechanism. Specifically, firstly, the initial weights of the first layer are... The first particle The dimensional weights are independently and randomly sampled, and their calculation formula is as follows:

[0056] ; Then it was normalized: ; in, For the first The first particle The initial sampled values ​​of the dimensional weights, where ~ indicates that they follow a certain probability distribution. Representing an interval The standard uniform distribution on the surface, and the initial velocity It is also initialized to 0. Since the initial position of the particles has been obtained through uniform random sampling, the particle swarm can still maintain good search diversity in the initial stage even if the initial velocity is zero.

[0057] S32: Weighted Representation Construction Based on the The particle in the first Position coordinates in the next iteration , for the The outlier features extracted from each hidden layer of a sample are dynamically weighted and fused to obtain a weighted high-dimensional latent representation of the sample under the current particle weight configuration. : ; For the full training samples, the second step can be further constructed. The particle in the first The weighted feature matrix corresponding to the next iteration : ; in, The total number of samples.

[0058] S33: Weight Iterative Optimization In each iteration, the particle swarm updates the historical best position and the global best position of each individual particle based on fitness feedback. For the ... The particle, up to the [number]th particle. The best historical position of an individual in the next iteration is denoted as... :

[0059] ; up to the In the next iteration, the global optimal position among all particles is denoted as: ; in, For the first The first particle The weight vector of the next iteration For the first The first particle The fitness value of the next iteration. For the first The first particle The fitness value of the next iteration.

[0060] Subsequently, each particle updates its velocity and position based on its individual historical best position and the population's global best position: ; ; in, Indicates the first The particle in the first The velocity vector at the next iteration; Inertial weight; and These are individual learning factors and social learning factors, respectively. The variables are random. The particle swarm optimization algorithm in this embodiment adopts a linearly decreasing inertia weight strategy:

[0061] ; in, For the first The inertia weight of the next iteration, and These are the maximum inertia weight and the minimum inertia weight, respectively. This represents the maximum number of iterations.

[0062] Since the layer attention weights need to satisfy the constraints of nonnegativity and summing to 1, they are nonnegated and normalized after each particle position update: ; in, For the first The first particle Normalized values ​​after dimension weight update The updated unconstrained weight values. For the hidden layer to be calculated, To perform the summation of all hidden layers during normalization, To prevent division by zero smoothing constants.

[0063] After a preset maximum number of iterations Afterward, the particle swarm optimization process converges. At this point, the final globally optimal particle position is used as the optimal layer attention weight distribution. :

[0064] ; in, Indicates the final assignment to the first The optimal attention weights for each hidden layer. These weights reflect the degree to which different hidden layers contribute to the backdoor anomaly response; the larger the weight, the more important the hidden layer is in amplifying the anomalous consistency of the poisoned samples.

[0065] S4: Abnormal microcluster detection S41: Low-dimensional density modeling To provide a clear optimization direction for PSO, dimensionality reduction and density clustering results are embedded into the fitness calculation process. For each particle's corresponding weighted feature matrix... First, UMAP is used to map it to a low-dimensional manifold space:

[0066] ; in, The dimension representing the low-dimensional embedding space is usually taken as... Subsequently, in the low-dimensional embedding space HDBSCAN density clustering is performed to capture anomalous microclusters with local high density and high stability.

[0067] HDBSCAN first constructs the density relationship between samples based on the distance between them. For sample points and Their mutual distance for:

[0068] ; in, Represents sample points To its first The core distance of the nearest neighbor, Represents sample points and The Euclidean distance between them is used. Based on the reach distance, HDBSCAN constructs a minimum spanning tree and generates a hierarchical clustering structure. Furthermore, it calculates the stability score of each candidate cluster through the density change process. Higher cluster stability indicates that the cluster is less likely to be split during density hierarchy changes and has stronger topological cohesion.

[0069] Suppose HDBSCAN is in The cluster set obtained above: ; in, Indicates the first The particle in the first The first iteration obtained There are several clusters.

[0070] S42: Fitness Calculation To eliminate the influence of large-scale benign main clusters, an abnormal micro-cluster size constraint is introduced, and samples with a proportion less than a threshold are selected. The set of candidate anomalous microclusters: ; in, This represents the prior threshold for the size of abnormal microclusters. This represents the total number of training samples. Based on this, the... The particle in the first The fitness at the next iteration is defined as:

[0071] ; in, Indicates the output of HDBSCAN. The stability score of each candidate microcluster. If empty, set the particle's fitness to 0. Fitness The larger the value, the more the layer weight corresponding to the particle can amplify the abnormal consistency of the backdoor sample, making it form a more stable topologically collapsed micro-cluster in the low-dimensional space.

[0072] S43: Poisoning Sample Detection Obtain the optimal layer attention weights Then, it is applied to the cross-layer extreme value features of the entire training sample to obtain the final globally weighted latent representation. For the th Each sample has a final weighted feature vector. for:

[0073] ; Furthermore, the final weighted feature matrix corresponding to all training samples is: ; Finally, Enter UMAP and HDBSCAN again to locate the anomalous microcluster that meets the size constraint and has the highest stability, and identify the samples in the microcluster as poisoned samples.

[0074] Comparison of experimental results To comprehensively evaluate the proposed method's ability to detect backdoor poisoning samples and the safety performance of the cleaned-up model, this paper evaluates it from two aspects: sample detection effectiveness and model defense effectiveness. Sample detection effectiveness is mainly measured using the True Positive Rate (TPR) and False Positive Rate (FPR).

[0075] The true positive rate (TPR) represents the proportion of correctly identified poisoned samples out of all truly poisoned samples. A higher TPR indicates a stronger ability of the detection method to identify poisoned samples and a lower false negative rate. The formula for calculating TPR is:

[0076] ; The false positive rate (FPR) represents the proportion of benign samples that are mistakenly identified as poisoned samples out of the total number of truly benign samples. A lower FPR indicates fewer false positives for clean samples and more reliable results. The formula for calculating FPR is:

[0077] ; On the CIFAR-10 and Tiny-ImageNet benchmark datasets, the performance of the proposed method was compared with that of four advanced defense methods, AC, SCALE-UP, AIBD and CCD, under three attack modes of four typical backdoor attacks: BadNets, Blend, Trojan and WaNet. Experimental results are shown in Tables 1 and 2. On the CIFAR-10 dataset, this invention achieves a TPR close to 100% and an FPR below 1% across all attack types and modes. In contrast, the AC method almost completely fails in A2A and UT modes, with TPR generally close to 0. SCALE-UP's detection performance drops sharply under dynamically triggered attacks such as WaNet, and its FPR remains above 20% for a long time. AIBD maintains a high detection rate, but its FPR is significantly higher than PED. CCD performs well under simple attacks, but its detection capability is limited under complex modes such as A2A and UT. On the more complex Tiny-ImageNet dataset, all the comparison methods show varying degrees of performance degradation, while this invention still maintains a stable performance of high TPR and low FPR, fully demonstrating its strong adaptability to complex data distributions and high-dimensional feature spaces. It shows significant advantages, especially in scenarios where traditional defenses are difficult to handle, such as A2A and UT, effectively solving the problems of existing methods being sensitive to attack modes and having poor generalization.

[0078] Table 1. Detection performance (%) of the ResNet-18 model on the CIFAR-10 dataset. Table 2. Detection performance (%) of the ResNet-18 model on the Tiny-ImageNet dataset. Unlike traditional backdoor defense methods that rely on attack priors, additional clean validation sets, or single-layer feature discrimination, the method proposed in this invention starts from the perspective of training data purification. It achieves effective identification of poisoned samples in data poisoning backdoor attacks through a three-stage collaborative mechanism: "cross-layer extreme value feature extraction - particle swarm layer weight optimization - density clustering anomaly micro-cluster detection." This method can fully exploit the weak anomalous responses and cross-layer consistency of poisoned samples in different hidden layers without attack priors or additional clean data, thereby improving the feature separability between poisoned and benign samples. This approach breaks through the dependence of existing methods on clean data and manual priors, providing a new theoretical path for dataset-level backdoor defense and theoretical support for poisoned sample detection and training data purification in complex backdoor attack scenarios.

[0079] The method proposed in this invention does not rely on additional clean data or prior attack information, exhibiting good versatility, robustness, and scalability, and can be applied to training data security protection in various security-critical scenarios. In cloud-based machine learning service platforms, this method can be used to detect and filter poisoned samples from training data uploaded or shared by third parties, reducing the risk of backdoor implantation from untrusted data sources. In autonomous driving scenarios, it can be used to identify backdoor poisoned samples constructed from traffic signs, road images, or environmental perception data, preventing models from making erroneous decisions under specific triggering conditions. In medical image diagnosis scenarios, it can help filter medical image samples that have been maliciously tampered with or implanted with triggering patterns, reducing the risk of backdoor attacks misleading diagnostic results. In summary, this invention not only improves the accuracy and reliability of training data purification but also provides effective support for the trusted deployment of deep learning models in practical security-critical fields.

[0080] This specification provides a data sample poisoning detection system, including: The alignment feature map generation module is used to obtain the untrusted training dataset and the deep neural network trained on the untrusted training dataset; extract the activation maps of multiple hidden layers of the deep neural network for each sample in the untrusted training dataset, and normalize the activation maps of different hidden layers to a uniform numerical scale to obtain the alignment feature map of each hidden layer.

[0081] The cross-layer extreme value feature extraction and splicing module is used to extract extreme values ​​along the spatial dimension and by channel for the aligned feature map of each hidden layer, as the abnormal features of that channel; splice the abnormal features of each channel of the current activation map to obtain the hidden layer channel abnormal splicing features; and splice the splicing features of all hidden layers to generate the cross-layer extreme value feature vector of each sample.

[0082] The layer-weighted particle swarm optimization module employs a particle swarm optimization algorithm, using the number of hidden layers as the search space dimension. It maps the attention weights of all hidden layers to particle positions within the search space. For each particle, the attention weights corresponding to that particle are used to weight and concatenate the cross-layer extreme feature vectors of all samples, resulting in a weighted feature matrix. After dimensionality reduction of the weighted feature matrix, density-based hierarchical clustering is performed to filter out anomalous micro-clusters whose sample proportion in the training set is less than a preset threshold. The stability value with the highest stability score among the anomalous micro-clusters is used as the fitness value of that particle. The particle positions are iteratively updated until convergence, yielding the globally optimal particle positions as the optimal layer attention weight distribution.

[0083] The poisoning sample determination module re-weights and concatenates the cross-layer extreme feature vectors of all samples using the optimal layer attention weight distribution to obtain the final weighted feature matrix. After dimensionality reduction of the final weighted feature matrix, density-based hierarchical clustering is performed to obtain a secondary screening abnormal micro-cluster set. The sample with the highest stability score in the secondary screening abnormal micro-cluster set, whose proportion of the training set is less than a preset threshold, is determined to be a poisoning sample and output.

[0084] For specific limitations regarding a data sample poisoning detection system, please refer to the limitations of a data sample poisoning detection method described above, which will not be repeated here. Each module in the aforementioned data sample poisoning detection system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.

[0085] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods described above.

[0086] This specification provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described above.

[0087] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0088] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for detecting poisoning in data samples, characterized in that, Specifically, it includes: Obtain an untrusted training dataset and a deep neural network trained on the untrusted training dataset; Extract the activation maps of multiple hidden layers of the deep neural network for each sample in the untrusted training dataset, and normalize the activation maps of different hidden layers to a uniform numerical scale to obtain the aligned feature map of each hidden layer; For the aligned feature map of each hidden layer, extract the extreme values ​​along the spatial dimension by channel, and use the extreme values ​​as the abnormal features of that channel; By splicing the abnormal features of each channel in the current activation image, we obtain spliced ​​features that represent the abnormalities of all channels in the hidden layer; by splicing the spliced ​​features of all hidden layers, we generate multiple cross-layer extreme value feature vectors that represent the abnormalities of each sample. A particle swarm optimization algorithm is employed, using the number of hidden layers as the search space dimension. Attention weights of all hidden layers are mapped to particle positions within the search space. For each particle, the attention weights corresponding to that particle are used to weight and concatenate the cross-layer extreme feature vectors of all samples, resulting in a weighted feature matrix. After dimensionality reduction of the weighted feature matrix, density-based hierarchical clustering is performed to filter out anomalous micro-clusters whose sample proportion in the training set is less than a preset threshold. The stability value with the highest stability score in the anomalous micro-cluster set is used as the fitness value of that particle. The particle positions are iteratively updated until convergence, yielding the globally optimal particle positions as the optimal layer attention weight distribution. The cross-layer extreme feature vectors of all samples are reweighted and concatenated using the optimal layer attention weight distribution to obtain the final weighted feature matrix. After dimensionality reduction of the final weighted feature matrix, density-based hierarchical clustering is performed to obtain a set of anomalous micro-clusters after secondary screening. The sample with the highest stability score among the abnormal micro-clusters in the secondary screening abnormal micro-clusters, whose proportion in the training set is less than a preset threshold, is identified as a poisoned sample.

2. The method for detecting poisoning in data samples as described in claim 1, characterized in that, The method for normalizing the activation maps of different hidden layers to a uniform numerical scale to obtain the aligned feature map of each hidden layer is as follows: Regarding the first Hidden layer Channel 1 The original activation values ​​of the activation map of the hidden layer at the location, and the moving average of the batch-normalized values ​​of that hidden layer. and variance Normalization is performed to obtain the normalized activation response value. ; ; in, For the first Layer Channel 1 The original activation value of the location, To prevent division by zero smoothing constant; The normalized activation response values ​​of all channels of the hidden layer are combined into an aligned feature map.

3. The method for detecting poisoning in data samples as described in claim 1, characterized in that, For each hidden layer's aligned feature map, along the spatial dimension, extreme values ​​are extracted by channel as the abnormal features of that channel. Specifically, for each channel of the activation map of each hidden layer, the global minimum value of the response values ​​at all spatial locations on the feature map of that channel is extracted as the abnormal feature of that channel.

4. The method for detecting poisoning in data samples as described in claim 1, characterized in that, The particle swarm optimization algorithm employs a linearly decreasing inertia weight strategy, specifically: ; in, For the first The inertia weight of the next iteration, and These are the maximum inertia weight and the minimum inertia weight, respectively. This represents the maximum number of iterations.

5. The method for detecting poisoning in data samples as described in claim 1, characterized in that, The step of using the highest stability score among the abnormal micro-clusters as the fitness value of a particle is as follows: for any particle in each iteration, if its corresponding abnormal micro-cluster set is not empty, the score corresponding to the micro-cluster with the highest stability evaluation in the set is selected as the fitness value of the particle; if the abnormal micro-cluster set is empty, the fitness value of the particle is set to zero.

6. The method for detecting poisoning in data samples as described in claim 1, characterized in that, For each particle, the attention weight corresponding to that particle is used to weight and concatenate the cross-layer extreme feature vectors of all samples to obtain a weighted feature matrix. Specifically, for the current attention weight of each particle, the attention weight is multiplied by the cross-layer extreme feature vector of the corresponding layer in sequence and then concatenated to form a weighted feature matrix of all samples. The multiple hidden layers are selected from at least two convolutional layers of different depths in the deep neural network.

7. The method for detecting poisoning in data samples as described in claim 1, characterized in that: The untrusted training dataset is an autonomous driving domain dataset or a medical image dataset that exhibits at least one attack mode; the attack mode includes at least All-to-One, A2O, All-to-All, A2A, or Untargeted, UT; the poisoned sample is a sample with a trigger injected into the untrusted training set; The label of the poisoned sample was tampered with and changed to the attack target label associated with the trigger, so that the deep neural network could establish a hidden mapping relationship between the trigger features and the attack target label during the training process. When the untrusted training set is an autonomous driving domain dataset, the backdoor trigger of the poisoned sample is an incorrect traffic signal instruction; when the untrusted training set is a medical image dataset, the backdoor trigger of the poisoned sample is an incorrect lesion classification result.

8. The method for detecting poisoning in data samples as described in claim 1, characterized in that: The preset threshold value ranges from greater than 0.05 to less than 0. When screening micro-clusters, only clusters containing a number of samples less than the product of the preset threshold and the total number of samples in the training set are identified as micro-clusters that meet the size condition, thereby excluding the main cluster composed of a large number of benign samples.

9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 7.

10. A computer device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any one of claims 1 to 7.