A backdoor trigger detection method and system based on disturbance separation

CN122333462BActive Publication Date: 2026-09-01HANGZHOU ANQUAN DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610779405.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2026-03-26
Filing Date
2026-06-02
Publication Date
2026-09-01
Estimated Expiration
2046-06-02

AI Technical Summary

Technical Problem

其中,基于神经元激活分析的防御技术其假设后门效应仅由少数被感染的神经元体现,通过刺激每个神经元并观察其对模型输出的影响来识别异常;然而,后门效应仅由少数被感染的神经元体现这一假设在实际复杂模型中往往不成立,存在未知失效场景,易被新型攻击绕过,当后门特征与正常特征深度交织时,该方法难以精确定位究竟是哪部分特征分量在诱发恶意决策,系统鲁棒性比较差

Benefits of technology

[0008] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333462B_ABST
    Figure CN122333462B_ABST
Patent Text Reader

Abstract

This specification discloses a backdoor trigger detection method and system based on perturbation separation. The method includes accessing a pre-trained two-stream decoupling coding model, which is determined through decoupling training based on mutual information minimization constraints; determining a baseline distribution of normal perturbation features based on a first perturbation feature; acquiring a test sample, determining a second perturbation feature of the test sample through a perturbation encoder, and determining the offset between the second perturbation feature and the baseline distribution of normal perturbation features; when a backdoor trigger is detected in the test sample based on the offset, separating the second perturbation feature from the test sample and inputting the semantic features of the test sample into a shared decoder for decoding to obtain a first model output of the test sample. This specification improves model security and reliability while enhancing backdoor trigger detection accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence security technology, specifically to a backdoor trigger detection method and system based on disturbance separation. Background Technology

[0002] With the widespread application of deep learning models in fields such as computer vision and natural language processing, their security has become a research focus. Backdoor attacks, also known as backdoor triggers, have become one of the most threatening attack methods in the field of artificial intelligence security. These attacks involve attackers injecting specific trigger patterns into the training data during model training, causing the model to produce a pre-set erroneous output when encountering inputs containing that trigger pattern, while behaving normally when processing normal samples, thus possessing extremely high stealth capabilities. Current research on defense against backdoor attacks mainly includes defense techniques based on neuron activation analysis and defense techniques based on activation space cleanup. Neuron activation analysis-based defense techniques assume that the backdoor effect is only manifested by a few infected neurons, identifying anomalies by stimulating each neuron and observing its impact on the model output. However, this assumption that the backdoor effect is only manifested by a few infected neurons often does not hold true in complex real-world models, leading to unknown failure scenarios and susceptibility to new attacks. When backdoor features are deeply intertwined with normal features, this method struggles to accurately pinpoint which feature component is inducing malicious decision-making, resulting in poor system robustness. Defense techniques based on activation space cleanup aim to clean up backdoor samples by bringing aberrant activation values ​​closer to the optimized minimum clean activation distribution interval within the activation space. However, the effectiveness of this method is highly dependent on the detectability of the activation distribution shift. If an attacker designs a backdoor trigger that causes the activation distribution shift to fall precisely within the optimized clean activation distribution interval, or if adversarial training makes the activation distribution of backdoor-related neurons indistinguishable from the normal distribution, the defense may be bypassed. Therefore, there is an urgent need for a backdoor trigger detection method that can improve model robustness and security while simultaneously increasing detection accuracy. Summary of the Invention

[0003] This specification provides a backdoor trigger detection method and system based on perturbation separation, the technical solution of which is as follows:

[0004] In a first aspect, embodiments of this specification provide a backdoor trigger detection method based on perturbation separation, comprising: accessing a pre-trained two-stream decoupling coding model, wherein the two-stream decoupling coding model is determined by decoupling training based on mutual information minimization constraints; the two-stream decoupling coding model includes at least a semantic encoder, a perturbation encoder, and a shared decoder; the semantic encoder is used to determine the semantic features of the input sample; the perturbation encoder is used to determine the perturbation features of the input sample; acquiring a normal sample, acquiring a first perturbation feature of the normal sample through the perturbation encoder, and determining a baseline distribution of the normal perturbation features based on the first perturbation feature; the normal sample is sample data without a backdoor trigger; acquiring a test sample, determining a second perturbation feature of the test sample through the perturbation encoder, and determining the offset between the second perturbation feature and the baseline distribution of the normal perturbation features; when a backdoor trigger is detected in the test sample based on the offset, separating the second perturbation feature from the test sample and inputting the semantic features of the test sample into the shared decoder for decoding to obtain a first model output of the test sample.

[0005] Secondly, embodiments of this specification provide a backdoor trigger detection method based on perturbation separation, comprising: a model training module for accessing a pre-trained two-stream decoupling coding model, wherein the two-stream decoupling coding model is determined by decoupling training based on mutual information minimization constraints; the two-stream decoupling coding model includes at least a semantic encoder, a perturbation encoder, and a shared decoder; the semantic encoder is used to determine the semantic features of the input sample; the perturbation encoder is used to determine the perturbation features of the input sample; a baseline distribution determination module is used to acquire normal samples, acquire the first perturbation feature of the normal samples through the perturbation encoder, and determine the baseline distribution of the normal perturbation features based on the first perturbation feature; the normal samples are sample data without backdoor triggers; an offset determination module is used to acquire the test sample, determine the second perturbation feature of the test sample through the perturbation encoder, and determine the offset between the second perturbation feature and the baseline distribution of the normal perturbation features; and a detection and purification module is used to separate the second perturbation feature from the test sample when the test sample is detected to carry a backdoor trigger based on the offset, and input the semantic features of the test sample to the shared decoder for decoding to obtain the first model output of the test sample.

[0006] Thirdly, embodiments of this specification provide an electronic device, including a processor and a memory; the processor is connected to the memory; the memory is used to store executable program code; the processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to perform the steps of the backdoor trigger detection method based on perturbation separation in the first aspect of the above embodiments.

[0007] Fourthly, embodiments of this specification provide a computer storage medium storing multiple instructions adapted for loading by a processor and executing the steps of the backdoor trigger detection method based on perturbation separation described in the first aspect of the above embodiments.

[0008] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:

[0009] The dual-stream decoupled coding model in this specification is determined based on decoupled training using mutual information minimization constraints. By introducing mutual information minimization constraints, semantic features and interference features tend to be independent in the representation space, reducing the possibility of backdoor trigger interference entering the semantic channel from the source. This achieves effective separation of semantic features and interference features, improving the robustness and security of the model.

[0010] Furthermore, the embodiments of this specification can also obtain the first perturbation feature of normal samples through an interference encoder and determine the normal perturbation feature baseline distribution based on the first perturbation feature; then, the second perturbation feature of the test sample can be determined through the interference encoder, and the offset between the second perturbation feature and the normal perturbation feature baseline distribution can be determined. The embodiments of this specification establish the normal perturbation feature baseline distribution and calculate the feature offset to make the backdoor triggers significantly different in the distance space, thereby achieving accurate differentiation between normal noise and backdoor triggers and improving detection accuracy. Moreover, by determining the second perturbation feature of the test sample and calculating its offset, the embodiments of this specification can limit the detection process to the interference branch and not perform destructive processing on the semantic branch, thus not reducing the prediction performance of normal samples and maintaining the semantic integrity of the task while detecting triggers.

[0011] Furthermore, when the embodiment of this specification detects that the sample under test carries a backdoor trigger based on the offset, the second perturbation feature can be separated from the sample under test, and the semantic features of the sample under test can be input to the shared decoder for decoding. The embodiment of this specification can remove the attack component without damaging the effective information by separating the interference feature and reconstructing it using only the semantic features, thereby achieving feature-level purification.

[0012] Moreover, the embodiments in this specification only involve feature extraction and distance calculation in the detection process, without the need for complex backpropagation, thereby meeting the requirements for online deployment, reducing additional data requirements and training costs, and facilitating deployment. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a schematic diagram illustrating an application scenario of the backdoor trigger detection method based on disturbance separation provided in this manual.

[0015] Figure 2 This is a flowchart illustrating the backdoor trigger detection method based on disturbance separation provided in this manual.

[0016] Figure 3 This is a flowchart illustrating the decoupled training process based on mutual information minimization constraints provided in this manual.

[0017] Figure 4 This is a flowchart illustrating the process for determining the baseline distribution of normal disturbance characteristics provided in this specification.

[0018] Figure 5 This is a schematic diagram of the backdoor trigger detection system based on disturbance separation provided in this manual.

[0019] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this specification. Detailed Implementation

[0020] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings.

[0021] The terms "first," "second," etc., in the description, claims, and accompanying drawings are used to distinguish different objects and not to describe a particular order. Furthermore, the term "comprising" and any variations thereof are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0022] This specification provides a backdoor trigger detection method based on perturbation separation in several embodiments. The execution subject of this backdoor trigger detection method based on perturbation separation can be the backdoor trigger detection system based on perturbation separation provided in the embodiments of this invention.

[0023] Before this specification describes in detail the backdoor trigger detection method based on perturbation separation with reference to one or more embodiments, it first introduces the application scenarios of this backdoor trigger detection method based on perturbation separation.

[0024] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of a backdoor trigger detection method based on perturbation separation provided in an embodiment of the present invention. In this embodiment, the backdoor trigger detection system 100 based on perturbation separation can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC); the server can be a single server or a server cluster composed of multiple servers.

[0025] In some embodiments, the backdoor trigger detection system 100 based on perturbation separation can also be integrated into multiple electronic devices. For example, the backdoor trigger detection system 100 based on perturbation separation can be integrated into multiple servers, and the backdoor trigger detection method based on perturbation separation of this application can be implemented by multiple servers.

[0026] In some embodiments, the server may also be implemented as a terminal. The terminal may be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC), etc. The terminal includes a central processing unit (CPU), a graphics processing unit (GPU), memory, storage devices, a network communication module, sensors, a display screen, a battery and power management module, etc.

[0027] For example, refer to Figure 1 The electronic device may include a server 110, a storage terminal 120, etc. The storage terminal 120 stores a pre-trained dual-stream decoupled coding model of the hallucination induction detection set, normal samples, test samples, etc. The server 110 and the storage terminal 120 communicate with each other, which will not be described in detail here.

[0028] Server 110 may include a processor and memory. Server 110 can access a pre-trained two-stream decoupling coding model, which is determined through decoupling training based on mutual information minimization constraints. The two-stream decoupling coding model includes at least a semantic encoder, a disturbance encoder, and a shared decoder. The semantic encoder is used to determine the semantic features of the input sample; the disturbance encoder is used to determine the disturbance features of the input sample; normal samples are acquired, and the first disturbance feature of the normal sample is obtained through the disturbance encoder, and a baseline distribution of the normal disturbance feature is determined based on the first disturbance feature; normal samples are sample data without backdoor triggers; a test sample is acquired, and the second disturbance feature of the test sample is determined through the disturbance encoder, and the offset between the second disturbance feature and the baseline distribution of the normal disturbance feature is determined; when a backdoor trigger is detected in the test sample based on the offset, the second disturbance feature is separated from the test sample, and the semantic features of the test sample are input to the shared decoder for decoding to obtain the first model output of the test sample, etc.

[0029] It should be noted that, Figure 1 The schematic diagram of the backdoor trigger detection system based on perturbation separation shown is merely an example. The backdoor trigger detection system and scenario based on perturbation separation described in this embodiment are for the purpose of more clearly illustrating the technical solutions of this embodiment and do not constitute a limitation on the technical solutions provided by this embodiment. As those skilled in the art will know, with the evolution of the backdoor trigger detection system based on perturbation separation and the emergence of new scenarios, the technical solutions provided by this embodiment are also applicable to similar technical problems.

[0030] Please see Figure 2 , Figure 2 This is a flowchart illustrating a backdoor trigger detection method based on perturbation separation provided in an embodiment of the present invention. This backdoor trigger detection method based on perturbation separation can be... Figure 1 The backdoor trigger detection system 100 based on perturbation separation shown is executed. This backdoor trigger detection method based on perturbation separation may include at least the following steps:

[0031] 200. Connect to a pre-trained two-stream decoupling coding model, which is determined by decoupling training based on mutual information minimization constraints;

[0032] 210. Obtain normal samples, acquire the first perturbation feature of the normal samples through the interference encoder, and determine the baseline distribution of the normal perturbation feature based on the first perturbation feature;

[0033] 220. Obtain the sample to be tested, determine the second perturbation feature of the sample to be tested through the interference encoder, and determine the offset between the second perturbation feature and the normal perturbation feature baseline distribution;

[0034] 230. When the backdoor trigger is detected in the sample under test based on the offset, the second perturbation feature is separated from the sample under test and the semantic feature of the sample under test is input to the shared decoder for decoding to obtain the first model output of the sample under test.

[0035] In this embodiment, the two-stream decoupled coding model may include at least a semantic encoder, a interference encoder, and a shared decoder; the semantic encoder is used to determine the semantic features of the input sample; the interference encoder is used to determine the perturbation features of the input sample. Normal samples may be sample data without a backdoor trigger.

[0036] In this embodiment, the semantic encoder may include a deep convolutional structure, the interference encoder may include a lightweight convolutional structure, and the shared decoder may include a deconvolutional network structure. The input to the shared decoder is concatenated data or semantic features of the test sample. The shared decoder is used for self-supervised reconstruction tasks during the training phase. For example, the semantic encoder may adopt the ResNet-18 architecture. The interference encoder may be a custom lightweight convolutional structure, such as containing four convolutional blocks, each of which may include a 3×3 convolution, a LeakyReLU activation layer, a normalization layer, etc. The shared decoder may adopt a 5-layer deconvolutional network symmetrical to the encoder. Each deconvolutional network includes at least a deconvolutional layer, a batch normalization layer, and a ReLU activation layer. The last layer may use the Tanh activation function to map the output to the range [-1, 1]. This embodiment is based entirely on convolutional networks and lightweight structures, and can be directly trained and accelerated using existing deep learning frameworks. It can be deployed on general-purpose hardware and has high engineering feasibility.

[0037] In some embodiments, please refer to Figure 3 , Figure 3 This is a schematic diagram of the decoupling training process based on mutual information minimization constraints provided in an embodiment of the present invention. The decoupling training based on mutual information minimization constraints includes:

[0038] 300. Determine the mutual information minimization loss function. The mutual information minimization loss function is used to minimize the mutual information between semantic features and perturbation features.

[0039] 310. Determine the reconstruction consistency loss function, which is the reconstruction error between the shared decoder output and the input samples;

[0040] 320. Determine the semantic invariance loss function. The semantic invariance loss function is used to constrain input samples with the same semantics but carrying different backdoor triggers to remain consistent in the corresponding output of the semantic encoder.

[0041] 330. Determine the joint loss function based on the mutual information minimization loss function, the reconstruction consistency loss function, and the semantic invariance loss function;

[0042] 340. Based on the joint loss function, the two-stream decoupling coding model is decoupled and trained to obtain a well-trained two-stream decoupling coding model.

[0043] This embodiment can determine the mutual information minimization loss function. By minimizing the mutual information between semantic features and perturbation features, the semantic features and perturbation features are orthogonal in the feature space, thereby ensuring that the semantic branch does not contain any interfering information.

[0044] This embodiment can determine the reconstruction consistency loss function to constrain the reconstruction error, that is, the error between the shared decoder output and the input sample, so as to minimize the reconstruction error and thus ensure the integrity and reversibility of feature decomposition.

[0045] This embodiment can determine the semantic invariance loss function, use different variants of known poisoned samples for training, require the semantic encoder to maintain a high degree of consistency in the extracted features when the same semantics are accessed but different triggers are transformed, and constrain input samples with the same semantics but carrying different backdoor triggers to remain consistent in the corresponding output of the semantic encoder.

[0046] In this embodiment, the joint loss function can be determined by weighting the mutual information minimization loss function, the reconstruction consistency loss function, and the semantic invariance loss function.

[0047] In some embodiments, please refer to Figure 4 , Figure 4 This is a schematic flowchart illustrating the process of determining the baseline distribution of normal perturbation features according to an embodiment of the present invention. The process includes: acquiring normal samples, obtaining the first perturbation features of the normal samples through an interference encoder, and determining the baseline distribution of normal perturbation features based on the first perturbation features; and further including:

[0048] 400. Obtain a sample set that includes several normal samples;

[0049] 410. A first perturbation feature set is determined by an interference encoder. The first perturbation feature set includes the first perturbation features corresponding to several normal samples.

[0050] 420. Determine the statistical characteristics corresponding to the first disturbance feature set. The statistical characteristics are used to determine the baseline distribution of normal disturbance features.

[0051] In this embodiment, normal samples can be sample data without backdoor triggers. Statistical features may include at least the mean center and covariance matrix.

[0052] In this embodiment, the first perturbation feature can be the perturbation feature of normal samples. During the pre-training phase, this embodiment can extract the perturbation features of normal samples using a perturbation encoder, thereby establishing the statistical distribution manifold of normal background noise, i.e., the normal perturbation feature baseline distribution. The normal perturbation feature baseline distribution is used to characterize the statistical characteristics and distribution range of various perturbation factors in the natural environment. Perturbation factors can include ambient light and shadow, filters, etc. This embodiment can use a self-supervised reconstruction method to establish the normal perturbation feature baseline distribution, without relying on specially labeled poisoned samples, reducing training data costs and facilitating deployment.

[0053] In some embodiments, determining the offset between the second perturbation feature and the normal perturbation feature reference distribution includes: determining the Mahalanobis distance or cosine similarity between the second perturbation feature and the normal perturbation feature reference distribution to obtain the offset between the second perturbation feature and the normal perturbation feature reference distribution.

[0054] In this embodiment, the mean vector corresponding to the mean center of the normal perturbation feature baseline distribution can be determined first, and then the Mahalanobis distance or cosine similarity between the second perturbation feature and the mean vector corresponding to the normal perturbation feature baseline distribution can be calculated to obtain the offset between the second perturbation feature and the normal perturbation feature baseline distribution.

[0055] In some embodiments, before detecting that the sample under test carries a backdoor trigger based on the offset, the method includes: acquiring a second perturbation feature and a semantic feature of the sample under test; inputting the second perturbation feature into a shared decoder to obtain first prediction data corresponding to the second perturbation feature; inputting the semantic feature of the sample under test into the shared decoder to obtain second prediction data corresponding to the semantic feature of the sample under test; determining the similarity between the first prediction data and the second prediction data to obtain similarity data; and determining that the sample under test carries a backdoor trigger when the offset is greater than an offset threshold and the similarity data is greater than a similarity threshold.

[0056] In this embodiment, the second perturbation feature can be the perturbation feature of the sample to be tested. In this embodiment, the perturbation feature can be determined to be logically conflicted with the semantic feature based on the similarity between the first prediction data corresponding to the second perturbation feature and the second prediction data corresponding to the semantic feature. When the offset is greater than the offset threshold and the similarity data is greater than the similarity threshold, it is determined that the sample to be tested carries a backdoor trigger.

[0057] In this embodiment, the energy value can be used to further determine whether the sample to be tested carries a backdoor trigger. Specifically, when the energy value exceeds the energy threshold range, the offset is greater than the offset threshold, and the similarity data is greater than the similarity threshold, it is determined that the sample to be tested carries a backdoor trigger. The energy value can be the L2 norm corresponding to the second perturbation feature.

[0058] This embodiment can detect various types of triggers by using offset, characteristic energy value, and degree of logical conflict as the basis for judgment, rather than the style of the backdoor trigger. This embodiment also exhibits good generalization ability by reducing reliance on prior information such as the shape, size, and position of backdoor triggers.

[0059] In some embodiments, the backdoor trigger detection method based on perturbation separation further includes: when the backdoor trigger is not detected in the sample under test based on the offset, concatenating the semantic features of the sample under test with the second perturbation features of the sample under test to obtain concatenated data; and inputting the concatenated data into a shared decoder for decoding to obtain the second model output of the sample under test.

[0060] In this embodiment, if the test sample is not detected to carry a backdoor trigger based on the offset, it means that the test sample is a normal sample. The system concatenates the semantic features of the test sample with the second perturbation features of the test sample and inputs them into the shared decoder. This means that when the test sample is a normal sample, the system retains the environmental details corresponding to the sample.

[0061] The dual-stream decoupled coding model in this specification is determined based on decoupled training using mutual information minimization constraints. By introducing mutual information minimization constraints, semantic features and interference features tend to be independent in the representation space, reducing the possibility of backdoor trigger interference entering the semantic channel from the source, effectively separating semantic features from interference features, and improving the robustness of the model.

[0062] Furthermore, the embodiments of this specification can also obtain the first perturbation feature of normal samples through an interference encoder and determine the normal perturbation feature baseline distribution based on the first perturbation feature; then, the second perturbation feature of the test sample can be determined through the interference encoder, and the offset between the second perturbation feature and the normal perturbation feature baseline distribution can be determined. The embodiments of this specification establish the normal perturbation feature baseline distribution and calculate the feature offset to make the backdoor triggers significantly different in the distance space, thereby achieving accurate differentiation between normal noise and backdoor triggers and improving detection accuracy. Moreover, by determining the second perturbation feature of the test sample and calculating its offset, the embodiments of this specification can limit the detection process to the interference branch and not perform destructive processing on the semantic branch, thus not reducing the prediction performance of normal samples and maintaining the semantic integrity of the task while detecting triggers.

[0063] Furthermore, when the embodiment of this specification detects that the sample under test carries a backdoor trigger based on the offset, the second perturbation feature can be separated from the sample under test, and the semantic features of the sample under test can be input to the shared decoder for decoding. The embodiment of this specification can remove the attack component without damaging the effective information by separating the interference feature and reconstructing it using only the semantic features, thereby achieving feature-level purification.

[0064] Moreover, the embodiments in this specification only involve feature extraction and distance calculation in the detection process, without the need for complex backpropagation, thereby meeting the requirements for online deployment, reducing additional data requirements and training costs, and facilitating deployment.

[0065] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0066] Please see Figure 5 , Figure 5 This is a schematic diagram of the backdoor trigger detection system based on disturbance separation provided in the embodiments of this specification.

[0067] like Figure 5 As shown, the backdoor trigger detection system based on perturbation separation may include at least a model training module 500, a baseline distribution determination module 510, an offset determination module 520, and a detection and purification module 530, wherein:

[0068] The model training module 500 is used to input a pre-trained two-stream decoupling coding model. The two-stream decoupling coding model is determined by decoupling training based on mutual information minimization constraints. The two-stream decoupling coding model includes at least a semantic encoder, a disturbance encoder, and a shared decoder. The semantic encoder is used to determine the semantic features of the input samples. The disturbance encoder is used to determine the perturbation features of the input samples.

[0069] The baseline distribution determination module 510 is used to acquire normal samples, acquire the first perturbation feature of the normal samples through the interference encoder, and determine the baseline distribution of the normal perturbation feature based on the first perturbation feature; the normal samples are sample data without backdoor triggers;

[0070] The offset determination module 520 is used to acquire the sample to be tested, determine the second perturbation feature of the sample to be tested through the interference encoder, and determine the offset between the second perturbation feature and the normal perturbation feature reference distribution.

[0071] The detection and purification module 530 is used to separate the second perturbation feature from the sample under test when the backdoor trigger is detected based on the offset, and input the semantic feature of the sample under test to the shared decoder for decoding to obtain the first model output of the sample under test.

[0072] In some embodiments, the model training module 500 includes a training submodule, which is used to: determine a mutual information minimization loss function, which minimizes the mutual information between semantic features and perturbation features; determine a reconstruction consistency loss function, which is the reconstruction error between the shared decoder output and the input sample; determine a semantic invariance loss function, which constrains input samples with the same semantics but carrying different backdoor triggers to maintain consistency in the corresponding output of the semantic encoder; determine a joint loss function based on the mutual information minimization loss function, the reconstruction consistency loss function, and the semantic invariance loss function; and perform decoupling training on the two-stream decoupled coding model based on the joint loss function to obtain a trained two-stream decoupled coding model.

[0073] In some embodiments, the baseline distribution determination module 510 includes a baseline establishment module, which is used to: acquire a sample set including a plurality of normal samples; determine a first perturbation feature set by an interference encoder, the first perturbation feature set including first perturbation features corresponding to the plurality of normal samples respectively; determine statistical features corresponding to the first perturbation feature set, the statistical features being used to determine the baseline distribution of normal perturbation features; the statistical features include at least the mean center and the covariance matrix.

[0074] In some embodiments, the offset determination module 520 includes an offset calculation module, which is used to: determine the Mahalanobis distance or cosine similarity between the second perturbation feature and the normal perturbation feature reference distribution, and obtain the offset between the second perturbation feature and the normal perturbation feature reference distribution.

[0075] In some embodiments, the backdoor trigger detection system based on perturbation separation further includes a detection module.

[0076] The detection module is used to: acquire the second perturbation feature and the semantic feature of the sample to be tested; input the second perturbation feature into the shared decoder to obtain the first prediction data corresponding to the second perturbation feature; input the semantic feature of the sample to be tested into the shared decoder to obtain the second prediction data corresponding to the semantic feature of the sample to be tested; determine the similarity between the first prediction data and the second prediction data to obtain similarity data; when the offset is greater than the offset threshold and the similarity data is greater than the similarity threshold, it is determined that the sample to be tested carries a backdoor trigger.

[0077] In some embodiments, the backdoor trigger detection system based on perturbation separation further includes a splicing module, which is used to: splice the semantic features of the test sample with the second perturbation features of the test sample when the offset does not detect that the test sample carries a backdoor trigger, to obtain spliced ​​data; and input the spliced ​​data into a shared decoder for decoding to obtain the second model output of the test sample.

[0078] In some embodiments, the semantic encoder includes a deep convolutional structure, the interference encoder includes a lightweight convolutional structure, and the shared decoder includes a deconvolutional network structure. The input to the shared decoder is concatenated data or semantic features of the sample to be tested. The shared decoder is used for self-supervised reconstruction tasks during the training phase.

[0079] Based on the backdoor trigger detection system based on perturbation separation in several embodiments of this specification, it can be seen that the dual-stream decoupled coding model in the embodiments of this specification is determined by decoupling training based on mutual information minimization constraints. By introducing mutual information minimization constraints, semantic features and interference features tend to be independent in the representation space, reducing the possibility of backdoor trigger interference entering the semantic channel from the source, realizing the effective separation of semantic features and interference features, and improving the robustness of the model.

[0080] Furthermore, the embodiments of this specification can also obtain the first perturbation feature of normal samples through an interference encoder and determine the normal perturbation feature baseline distribution based on the first perturbation feature; then, the second perturbation feature of the test sample can be determined through the interference encoder, and the offset between the second perturbation feature and the normal perturbation feature baseline distribution can be determined. The embodiments of this specification establish the normal perturbation feature baseline distribution and calculate the feature offset to make the backdoor triggers significantly different in the distance space, thereby achieving accurate differentiation between normal noise and backdoor triggers and improving detection accuracy. Moreover, by determining the second perturbation feature of the test sample and calculating its offset, the embodiments of this specification can limit the detection process to the interference branch and not perform destructive processing on the semantic branch, thus not reducing the prediction performance of normal samples and maintaining the semantic integrity of the task while detecting triggers.

[0081] Furthermore, when the embodiment of this specification detects that the sample under test carries a backdoor trigger based on the offset, the second perturbation feature can be separated from the sample under test, and the semantic features of the sample under test can be input to the shared decoder for decoding. The embodiment of this specification can remove the attack component without damaging the effective information by separating the interference feature and reconstructing it using only the semantic features, thereby achieving feature-level purification.

[0082] Moreover, the embodiments in this specification only involve feature extraction and distance calculation in the detection process, without the need for complex backpropagation, thereby meeting the requirements for online deployment, reducing additional data requirements and training costs, and facilitating deployment.

[0083] Furthermore, the embodiments of this specification can implement hierarchical and precise intervention through topological isolation technology, realizing a flexible response system composed of different levels of intervention and emergency intervention. This hierarchical system can dynamically adjust the intensity of intervention according to the severity of hallucination, the risk level of the scene, and the real-time context, achieving an optimal balance between safety and effectiveness. Moreover, the embodiments of this specification can also use topological isolation technology to ensure that the function of polysemous neurons on non-hallucinatory pathways is completely unaffected through hierarchical intervention methods such as targeted noise injection, attention rerouting, and conditional gating.

[0084] Moreover, the embodiments in this specification can be integrated into various large language models based on the Transformer architecture, and are applicable to many application scenarios such as intelligent question answering, text generation, code programming, machine translation, and summary generation. Especially in key fields such as financial analysis, medical diagnosis, and legal consultation, where the accuracy of facts is extremely important, this invention can significantly improve the credibility and security of AI systems, and has broad market application value and industrial promotion prospects.

[0085] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the embodiment of the backdoor trigger detection system based on perturbation separation is relatively simple in description because it is fundamentally similar to the embodiment of the backdoor trigger detection method based on perturbation separation; relevant parts can be referred to the description of the method embodiment.

[0086] Please see Figure 6 The diagram shown is a structural schematic of an electronic device provided in an embodiment of this specification.

[0087] like Figure 6 As shown, the electronic device 600 may include at least one processor 610, at least one network interface 640, a user interface 630, a memory 650, and at least one communication bus 620.

[0088] The communication bus 620 can be used to realize the connection and communication of the above components.

[0089] The user interface 630 may include buttons, and the optional user interface may also include a standard wired interface or a wireless interface.

[0090] The network interface 640 may include, but is not limited to, Bluetooth modules, NFC modules, Wi-Fi modules, etc.

[0091] The processor 610 may include one or more processing cores. The processor 610 connects to various parts within the electronic device 600 using various interfaces and lines. It executes various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 650, and by calling data stored in the memory 650. Optionally, the processor 610 may be implemented using at least one hardware form selected from DSP, FPGA, and PLA. The processor 610 may integrate one or more of the following: CPU, GPU, and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 610 and may be implemented as a separate chip.

[0092] The memory 650 may include RAM or ROM. Optionally, the memory 650 may include a non-transitory computer-readable medium. The memory 650 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 650 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 650 may also be at least one storage device located remotely from the aforementioned processor 610. The memory 650, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a backdoor trigger detection application based on perturbation separation. The processor 610 may be used to call the backdoor trigger detection application based on perturbation separation stored in the memory 650 and execute the steps of the backdoor trigger detection method based on perturbation separation mentioned in the foregoing embodiments.

[0093] This specification also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform the above-described instructions. Figures 2 to 4 One or more steps in the illustrated embodiment. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0094] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this specification is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).

[0095] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.

[0096] The above embodiments are merely preferred embodiments described in this specification and are not intended to limit the scope of this specification. Any modifications and improvements made by those skilled in the art to the technical solutions of this specification without departing from the spirit of this specification should fall within the protection scope defined by the claims of this specification.

Claims

1. A backdoor trigger detection method based on perturbation separation, characterized in that, include: The system accesses a pre-trained two-stream decoupling coding model, which is obtained through decoupling training based on a joint loss function determined by a mutual information minimization loss function, a reconstruction consistency loss function, and a semantic invariance loss function; the two-stream decoupling coding model includes at least a semantic encoder, an interference encoder, and a shared decoder. The semantic encoder is used to determine the semantic features of the input sample; the interference encoder is used to determine the perturbation features of the input sample. Obtain normal samples, acquire the first perturbation feature of the normal samples through the interference encoder, and determine the baseline distribution of normal perturbation features based on the first perturbation feature; The normal samples are sample data that do not carry backdoor triggers; A sample to be tested is acquired, and a second perturbation feature of the sample to be tested is determined by the interference encoder. The offset between the second perturbation feature and the normal perturbation feature baseline distribution is also determined. When the test sample is detected to carry a backdoor trigger based on the offset, the second perturbation feature is separated from the test sample, and the semantic features of the test sample are input to the shared decoder for decoding to obtain the first model output of the test sample.

2. The method according to claim 1, characterized in that, The decoupled training of the joint loss function determined by the mutual information minimization loss function, the reconstruction consistency loss function, and the semantic invariance loss function includes: Determine a mutual information minimization loss function, which is used to minimize the mutual information between the semantic features and the perturbation features; Determine the reconstruction consistency loss function, which is the reconstruction error between the shared decoder output and the input sample; A semantic invariant loss function is determined, which is used to constrain input samples with the same semantics but carrying different backdoor triggers to remain consistent in the corresponding output of the semantic encoder; The joint loss function is determined based on the mutual information minimization loss function, the reconstruction consistency loss function, and the semantic invariance loss function; The dual-stream decoupled coding model is decoupled and trained based on the joint loss function to obtain a trained dual-stream decoupled coding model.

3. The method according to claim 1, characterized in that, The step of obtaining normal samples, acquiring the first perturbation feature of the normal samples through the interference encoder, and determining the baseline distribution of normal perturbation features based on the first perturbation feature includes: Obtain a sample set that includes several normal samples; The first perturbation feature set is determined by the interference encoder, and the first perturbation feature set includes the first perturbation features corresponding to the plurality of normal samples respectively; Determine the statistical features corresponding to the first perturbation feature set, the statistical features being used to determine the baseline distribution of normal perturbation features; the statistical features include at least the mean center and the covariance matrix.

4. The method according to claim 3, characterized in that, Determining the offset between the second disturbance feature and the normal disturbance feature baseline distribution includes: Determine the Mahalanobis distance or cosine similarity between the second perturbation feature and the normal perturbation feature reference distribution to obtain the offset between the second perturbation feature and the normal perturbation feature reference distribution.

5. The method according to claim 1, characterized in that, Before detecting that the sample under test carries a backdoor trigger based on the offset, the process includes: Obtain the second perturbation feature and the semantic feature of the sample to be tested; The second perturbation feature is input into the shared decoder to obtain the first prediction data corresponding to the second perturbation feature; The semantic features of the sample to be tested are input into the shared decoder to obtain the second prediction data corresponding to the semantic features of the sample to be tested; Determine the similarity between the first predicted data and the second predicted data to obtain similarity data; When the offset is greater than the offset threshold and the similarity data is greater than the similarity threshold, it is determined that the sample to be tested carries a backdoor trigger.

6. The method according to claim 1, characterized in that, Also includes: When the offset does not detect that the sample under test carries a backdoor trigger, the semantic features of the sample under test are concatenated with the second perturbation features of the sample under test to obtain concatenated data. The spliced ​​data is input into the shared decoder for decoding to obtain the second model output of the sample to be tested.

7. The method according to claim 6, characterized in that, The semantic encoder includes a deep convolutional structure, the interference encoder includes a lightweight convolutional structure, and the shared decoder includes a deconvolutional network structure. The input of the shared decoder is the concatenated data or the semantic features of the sample to be tested. The shared decoder is used for the self-supervised reconstruction task during the training phase.

8. A backdoor trigger detection system based on perturbation separation, characterized in that, include: The model training module is used to access a pre-trained two-stream decoupled coding model. This model is obtained through decoupling training based on a joint loss function determined by mutual information minimization, reconstruction consistency, and semantic invariance. The two-stream decoupled coding model includes at least a semantic encoder, a perturbation encoder, and a shared decoder. The semantic encoder is used to determine the semantic features of the input samples; the perturbation encoder is used to determine the perturbation features of the input samples. The baseline distribution determination module is used to acquire normal samples, obtain the first perturbation feature of the normal samples through the interference encoder, and determine the baseline distribution of normal perturbation features based on the first perturbation feature; the normal samples are sample data without backdoor triggers; The offset determination module is used to acquire the sample to be tested, determine the second perturbation feature of the sample to be tested through the interference encoder, and determine the offset between the second perturbation feature and the normal perturbation feature reference distribution. The detection and purification module is used to separate the second perturbation feature from the sample under test when the sample under test is detected to carry a backdoor trigger based on the offset, and input the semantic feature of the sample under test to the shared decoder for decoding to obtain the first model output of the sample under test.

9. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor runs the application program within the memory to perform the steps in the backdoor trigger detection method based on perturbation separation as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the backdoor trigger detection method based on perturbation separation as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Backdoor attack defense method and system

    CN113792289A

  • Deep reinforcement learning traffic signal control poisoning defense method based on strong disturbance detection and model retraining

    CN115361224A