Model-level backdoor inversion and detection method and device of visual language pre-training model
By constructing a joint loss function and analyzing the geometric properties of the loss plane, we achieved detection of visual language pre-trained models without prior information, solving the problem of text modality backdoor detection under unknown attacks in existing technologies, and improving detection accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF COMPUTING TECH CHINESE ACAD OF SCI
- Filing Date
- 2025-12-30
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies are insufficient to effectively detect text modal backdoors under unknown attacks in visual language pre-trained models, and they rely on triggers or specific task structures, making them unsuitable for real-world scenarios.
This paper proposes a model-level backdoor inversion and detection method that does not require prior knowledge. By constructing a joint loss function of assimilation loss, offset loss and anchor loss, the learnable features are optimized, and the geometric properties of the loss plane are used to determine whether the model has a backdoor. This includes gradient inversion of implicit backdoor features and orthogonal perturbation analysis.
It significantly improves detection accuracy, supports model detection under different architectures and attack paradigms, enhances detection efficiency and robustness, and is suitable for real-world deployment environments.
Smart Images

Figure CN121902145A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence security technology, and in particular to a model-level backdoor inversion, detection method, device, electronic device, computer-readable storage medium, and computer program product for a visual language pre-trained model. Background Technology
[0002] In recent years, visual language pre-trained models, such as open-source models like LLaVA and closed-source models like GPT-4o, have achieved outstanding performance in tasks such as zero-shot classification and cross-modal retrieval. With the rapid development of the open-source ecosystem, a large number of models have been downloaded, fine-tuned, and re-uploaded in untrusted environments, significantly increasing their security risks. Visual language pre-trained models heavily rely on the semantic alignment between their text encoder and visual encoder. Attackers can embed samples with trigger words during the fine-tuning stage, forcing the model to learn specific latent mapping relationships and form activatable backdoor behaviors. Once the trigger word appears, the model output will be biased towards the attacker's target semantics, while maintaining good performance on normal samples, making it highly stealthy and dangerous for attacks.
[0003] Most existing backdoor detection methods based on deep learning rely on downstream classifiers, such as patents CN113204745B and CN114021136B. The main technical approach involves extracting features from backdoor and benign image test samples in the frequency or spatial domains, while simultaneously scanning different categories to detect backdoor samples or abnormal model behavior. However, visual language pre-trained models are feature-aligned models, lacking fixed category outputs, and attackers do not rely on downstream classifiers or classifiers being benign when injecting, making direct application of the aforementioned methods difficult. In backdoor detection targeting visual language pre-trained models, existing technologies utilize decoders to reconstruct the abnormal distribution of input features to identify backdoor samples. However, these sample-level backdoor detection methods rely on trigger or target priors, i.e., scanning based on known trigger patterns or attack targets, but their detection capability is limited when facing unknown attacks. Existing technologies also utilize the separation of backdoor sample features from benign sample features to design a model-level backdoor inversion and detection technique. However, this work only addresses the case of image modality injection backdoors and is difficult to apply effectively to the case of text modality injection backdoors. Summary of the Invention
[0004] This invention aims to solve the problem that existing technologies rely on data, triggers, or specific task structures and are difficult to apply to real-world scenarios. It proposes a backdoor inversion and detection method and apparatus that can directly perform model-level detection on text encoders of visual language pre-trained models without any prior knowledge.
[0005] To address the shortcomings of existing technologies, such as Figure 2As shown, this invention proposes a model-level backdoor inversion and detection method for visual language pre-trained models, including:
[0006] The initial step involves obtaining the visual language pre-trained model to be detected as the target model and the text prompt words. The text prompt Convert to a character sequence and insert learnable features into that character sequence. The pre-trained text segmenter is then fed into the target model. and the anchor point model used for reference. ;
[0007] The inversion step constructs a joint loss function that includes assimilation loss, offset loss, and anchor point loss. The assimilation loss is used to encourage the learnable feature. The text internal features are induced to exhibit strong semantic assimilation; the offset loss is used to encourage the model output to deviate from the original text features, simulating the semantic shift when the backdoor is triggered; the anchor loss is used to encourage the target model's output to differ from the anchor model, amplifying the anomalous response caused by the backdoor; based on this joint loss function, gradient descent repeatedly optimizes the learnable feature. The features that can induce the target model to produce features that meet the judgment conditions are obtained and used as implicit backdoor features. ;
[0008] Calculation steps, towards implicit backdoor features Apply two orthogonal perturbation directions and A two-dimensional projection plane is constructed as the loss plane, and multiple perturbation points are sampled on the loss plane to calculate the corresponding loss values, thereby obtaining the local loss plane projection; the proportion of positive eigenvalues in the planar geometric properties of the local loss plane projection is calculated.
[0009] The detection steps are based on the anchor point model. The semantic assimilation features, semantic shift, and proportion of positive feature values are used as discrimination conditions for the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model. When the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model meet the discrimination conditions, the target model is a backdoor model; otherwise, the current implicit backdoor feature is removed. As a learnable feature Then return to the inversion step to continue iterative optimization until the number of iterations reaches the preset value. If the discrimination condition is not met when the number of iterations reaches the preset value, then the target model is a benign model.
[0010] The aforementioned model-level backdoor inversion and detection method for visual language pre-trained models, wherein the initial step includes:
[0011] For a containing A text dataset of text prompt words. Learnable features can be inserted Obtain inversion data from text datasets .
[0012] Assimilation loss in this inversion step :
[0013]
[0014] in, It is inversion data The m-th text prompt word The One text character, It is a text prompt word The sequence length, It calculates the cosine similarity between features;
[0015] This offset loss :
[0016]
[0017] The anchor point loss:
[0018]
[0019] The final total loss is:
[0020]
[0021] in, All of them are floating-point numbers in the range [0,1].
[0022] The aforementioned model-level backdoor inversion and detection method for visual language pre-trained models, wherein the calculation steps include:
[0023] Features of implicit backdoors The loss plane is used for calculation. ;in, Integers in the range [1, 20] The direction of the loss is the negative gradient, i.e. ; Calculate the spectral eigenvalues of the loss plane to reflect its geometric properties; for the first in the loss plane Construct a second-order Hessian matrix using points:
[0024]
[0025] Through calculation Hessen spectral characteristics The proportion of its positive eigenvalues is:
[0026]
[0027] in, It reflects the quantity of elements.
[0028] The aforementioned method for model-level backdoor inversion and detection of a visual language pre-trained model, wherein the detection step includes:
[0029] Determining the characteristics of implicit backdoors In the test set Does it meet the discrimination criteria? To determine whether the target model contains a backdoor, in the case of... Test datasets Above, simultaneously satisfying:
[0030]
[0031] This is the backdoor model; among which, , The spectral characteristics of the Hessian matrix are , The range is [0.6, 1.0]. The range is [0.95, 1.0]. The range is [-0.1, 0.1]. The range is [0.9, 1.0]. The range is [-0.1, 0.1]. The range is [0.7, 1.0].
[0032] like Figure 3 As shown, this invention also proposes a model-level backdoor inversion and detection device for a visual language pre-trained model, including:
[0033] The initial module obtains the visual language pre-trained model to be detected as the target model and text prompts. The text prompt Convert to a character sequence and insert learnable features into that character sequence. The pre-trained text segmenter is then fed into the target model. and the anchor point model used for reference. ;
[0034] The inversion module constructs a joint loss function that includes assimilation loss, offset loss, and anchor point loss. The assimilation loss is used to encourage the learnable feature. The text internal features are induced to exhibit strong semantic assimilation; the offset loss is used to encourage the model output to deviate from the original text features, simulating the semantic shift when the backdoor is triggered; the anchor loss is used to encourage the target model's output to differ from the anchor model, amplifying the anomalous response caused by the backdoor; based on this joint loss function, gradient descent repeatedly optimizes the learnable feature. The features that can induce the target model to produce features that meet the judgment conditions are obtained and used as implicit backdoor features. ;
[0035] The computation module, targeting implicit backdoor features Apply two orthogonal perturbation directions and A two-dimensional projection plane is constructed as the loss plane, and multiple perturbation points are sampled on the loss plane to calculate the corresponding loss values, thereby obtaining the local loss plane projection; the proportion of positive eigenvalues in the planar geometric properties of the local loss plane projection is calculated.
[0036] The detection module, based on the anchor point model The semantic assimilation features, semantic shift, and proportion of positive feature values are used as discrimination conditions for the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model. When the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model meet the discrimination conditions, the target model is a backdoor model; otherwise, the current implicit backdoor feature is removed. As a learnable feature Then return to the inversion module to continue iterative optimization until the number of iterations reaches the preset value. If the discrimination condition is not met when the number of iterations reaches the preset value, then the target model is a benign model.
[0037] The aforementioned model-level backdoor inversion and detection device for the visual language pre-trained model, wherein the initial module includes:
[0038] For a containing A text dataset of text prompt words. Learnable features can be inserted Obtain inversion data from text datasets .
[0039] Assimilation loss in this inversion module :
[0040]
[0041] in, It is inversion data The m-th text prompt word The One text character, It is a text prompt word The sequence length, It calculates the cosine similarity between features;
[0042] This offset loss :
[0043]
[0044] The anchor point loss:
[0045]
[0046] The final total loss is:
[0047]
[0048] in, All of them are floating-point numbers in the range [0,1].
[0049] The aforementioned model-level backdoor inversion and detection device for the visual language pre-trained model, wherein the computation module includes:
[0050] Features of implicit backdoors The loss plane is used for calculation. ;in, Integers in the range [1, 20] The direction of the loss is the negative gradient, i.e. ; Calculate the spectral eigenvalues of the loss plane to reflect its geometric properties; for the first in the loss plane Construct a second-order Hessian matrix using points:
[0051]
[0052] Through calculation Hessen spectral characteristics The proportion of its positive eigenvalues is:
[0053]
[0054] in, This reflects the number of elements; the detection module includes:
[0055] Determining the characteristics of implicit backdoors In the test set Does it meet the discrimination criteria? To determine whether the target model contains a backdoor, in the case of... Test datasets Above, simultaneously satisfying:
[0056]
[0057] This is the backdoor model; among which, , The spectral characteristics of the Hessian matrix are , The range is [0.6, 1.0]. The range is [0.95, 1.0]. The range is [-0.1, 0.1]. The range is [0.9, 1.0]. The range is [-0.1, 0.1]. The range is [0.7, 1.0].
[0058] The present invention also proposes an electronic device, including a model-level backdoor inversion and detection device for a visual language pre-trained model, wherein the electronic device may be connected to an information display device, which is used to display the detection results of the target model with user-set display parameters, attributes or through an artificial intelligence model.
[0059] The present invention also proposes a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the model-level backdoor inversion and detection method of the visual language pre-trained model.
[0060] The present invention also proposes a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, it implements the steps of model-level backdoor inversion and detection method of the visual language pre-trained model.
[0061] As can be seen from the above solutions, the advantages of the present invention are:
[0062] This invention proposes a model-level backdoor inversion and detection method and apparatus for visual language pre-trained models. Compared with existing technologies, this invention completely eliminates the need for prior information (trigger words, targets, training sets, and downstream classifiers are all unnecessary), significantly improving detection accuracy. It supports model detection under different architectures and attack paradigms, and boasts high detection efficiency. Therefore, this invention can be widely applied to model security auditing in practical deployment environments. Attached Figure Description
[0063] Figure 1 This is a flowchart of an embodiment of the present invention;
[0064] Figure 2 This is a flowchart of the method of the present invention;
[0065] Figure 3 This is a block diagram of the device of the present invention;
[0066] Figure 4 This is a schematic diagram of the structure of the first electronic device of the present invention;
[0067] Figure 5 This is a schematic diagram of the application environment structure of the first electronic device of the present invention;
[0068] Figure 6 This is a schematic diagram of the structure of the second electronic device of the present invention.
[0069] Figure label:
[0070] A - First electronic device;
[0071] Model-level backdoor inversion and detection device for B-visual language pre-trained models;
[0072] C-Data acquisition equipment;
[0073] D-Information display device;
[0074] 1000 - Second electronic device;
[0075] Ⅰ-Computational Unit;
[0076] II-ROM;
[0077] III-RAM;
[0078] N-bus;
[0079] V-Interface;
[0080] VI - Input Unit;
[0081] VII - Output Unit;
[0082] VIII - Storage medium;
[0083] IX - Communication Unit. Detailed Implementation
[0084] It should be noted that, in this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus.
[0085] In the absence of further restrictions, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0086] The processor described in this invention is the control center of an electronic device. It can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of this invention, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0087] Alternatively, the processor can perform various functions of the electronic device by running or executing software programs stored in memory and by calling data stored in memory.
[0088] In a specific implementation, as one example, the processor may include one or more CPUs. Each of these processors may be a single-core processor or a multi-core processor. Here, "processor" can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Electronic devices may include servers, desktop computers, laptops, smartphones, tablets, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.
[0089] The memory is used to store the software program that executes the solution of the present invention, and the execution is controlled by the processor. For specific implementation methods, please refer to the above method embodiments, which will not be repeated here.
[0090] It should be noted that the structure of the electronic device shown in the accompanying drawings of this invention does not constitute a limitation thereof. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0091] The above embodiments can be implemented, in whole or in part, by software, hardware (such as circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0092] It should also be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0093] In this invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of a single item or a plurality of items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be a single item or multiple items.
[0094] It should also be understood that, in various embodiments of the present invention, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0095] In the several embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0096] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0097] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0098] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0099] To achieve the above-mentioned technical effects, the present invention proposes the following key technical points:
[0100] Key Point 1: This invention reveals for the first time the "feature assimilation" mechanism in backdoor models, namely, in a text encoder with a backdoor implanted, the representations of all characters within the same sample exhibit abnormally high similarity. Because backdoor implantation causes the model to focus its attention on trigger words, this leads to abnormally high similarity in the representations of all characters within the same sample. Technical Effects: This phenomenon demonstrates strong cross-architecture stability and is commonly observed on mainstream architectures such as CLIP, SigLIP, and LongCLIP; it provides an observable unified discriminative signal for model-level detection.
[0101] Key Point 2: This invention is the first to propose a backdoor feature inversion mechanism based on gradient inversion. Utilizing the aforementioned feature assimilation properties, this invention proposes a trigger feature inversion mechanism based on parametric gradient inversion, which can directly inversely invert potential backdoor activation features from the model. Technical advantages: No training data, trigger samples, or trigger targets are required; it can recover trigger features implanted by attackers; it can scan a model in a short time (approximately 5 minutes on a single RTX 4090 GPU).
[0102] Key Point 3: Identifying and Removing "Natural Backdoor Features". This invention discovers a natural feature in the official CLIP that, while not artificially implanted, exhibits a backdoor effect. This natural backdoor feature causes a benign CLIP model to ignore other textual semantics and only output the semantics corresponding to that feature. Because this feature is difficult to reconstruct from real discrete text for triggering, it is not harmful. To avoid false positives, this invention uses "loss plane analysis" to distinguish natural backdoor features from real backdoors. Technical effects: Significantly reduces the false positive rate; improves the reliability of model-level detection; enhances robustness in real-world applications.
[0103] To make the above-mentioned features and effects of the present invention clearer and easier to understand, specific embodiments are described below in conjunction with the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are merely illustrative. The scope of protection of the present invention is not limited to the disclosed embodiments, but is defined by the appended claims.
[0104] Please see Figure 1 A preferred embodiment of the present invention provides a backdoor inversion and detection method for model-level detection of visual language models.
[0105] In large models, discrete text is mapped to a continuous feature space. For example, the word "cat" is mapped to a continuous feature space of dimensions [1, 768]. Existing techniques require explicit recovery to the discrete text, while the technique of this application can determine whether the model has been implanted with a backdoor without needing to access the discrete trigger text. This application significantly improves the universality and robustness of backdoor scanning by inverting "implicit backdoor features" and utilizing their loss plane geometry properties for discrimination.
[0106] The method of this invention mainly includes the following steps:
[0107] S1, Input text prompts and construct implicit inversion tasks to invert the implicit backdoor features of the model. ; Prompt words, which can be self-created or derived from the public dataset DiffusionDB;
[0108] S2, based on implicit backdoor features Extracting the geometric properties of the loss plane;
[0109] S3 compares the assimilation degree, feature offset, and curvature index with the threshold to determine whether it is a backdoor model.
[0110] In step S1, the input text prompt word First, the text is converted into a character sequence using a pre-trained text encoder, and then an optimizable implicit backdoor feature is inserted at the beginning of the sequence. Subsequently, the input sequence with implicit backdoor features is fed into a pre-trained text segmenter. and the anchor point model used for reference. (Typically, this refers to officially released models with the same architecture and no backdoors). To invert potential backdoor triggering characteristics, a joint loss function is constructed. It comprises the following three sub-modules: 1) Assimilation Loss: Encourages implicit backdoor features to induce strong semantic assimilation features within the text; 2) Offset Loss: Encourages the model output to shift relative to the original text features, simulating the semantic shift when the backdoor is triggered; 3) Anchor Loss: Encourages the output of the target model to differ from the anchor model, amplifying the anomalous response caused by the backdoor. It is repeatedly optimized using gradient descent. Ultimately, a model is obtained that can induce the model to produce results that conform to the discrimination criterion. The implicit backdoor feature is denoted as The inversion process is a gradient optimization process based on the joint loss function. While the concept of inversion is not original to this application, the guiding signal of the inversion, namely the content of the loss function, is the original core invention of this application.
[0111] In this embodiment, in step S2, the present invention is based on the optimized implicit backdoor features. This further characterizes the loss plane structure of the model near this implicit backdoor feature. First, by... Apply two orthogonal perturbation directions and A two-dimensional projection plane is constructed, and moving points subject to disturbances are sampled on this plane. The corresponding loss values are calculated to obtain the local loss plane projection. Subsequently, the eigenvalues of this loss plane are calculated to reflect its structural characteristics. Experiments show that if the model contains a backdoor, its implicit backdoor features often fall into a stable "bowl-shaped" structure formed by artificial optimization, exhibiting a high proportion of positive eigenvalues; conversely, normal models only have "natural backdoor features," and their loss plane exhibits an irregular, multi-saddle-point structure with a significantly lower proportion of positive eigenvalues.
[0112] In this embodiment, in step S3, the present invention utilizes three key features obtained from the preceding steps: assimilation intensity. : Reflects whether implicit features assimilate the features of the entire input text; feature shift : Reflects whether implicit backdoor features cause a shift in the overall semantic representation; proportion of positive feature values This reflects the stability of the loss plane near the implicit backdoor feature. Thresholds are set for these indicators based on the statistical distribution of a large number of normal models. , , , , , This constitutes the final judgment criterion. Many normal models are derived from well-trained models of the official model. When the discrimination condition is met, the model is identified as a backdoor model; otherwise, the process returns to step S1 to continue optimization until the maximum step size is reached. Stop. If the condition is still not met after reaching the maximum step size. Then the model to be tested This is a benign model.
[0113] Assimilation intensity : Reflects the proportion of feature assimilation that implicit features make the entire input text exhibit; feature shift : The proportion of implicit backdoor features that cause a shift in the overall semantic representation; the proportion of positive feature values. : Reflects the stability of the loss plane near the implicit backdoor feature.
[0114] Working principle:
[0115] Please see Figure 1 This invention effectively distinguishes between text encoder models with implanted backdoors and benign models by introducing an optimizable implicit backdoor feature at the text input end and combining multiple loss constraints with its spectral features in the local loss plane.
[0116] We observed a phenomenon called "feature assimilation" in backdoor samples. Specifically, in the text encoder implanted with the backdoor, the character feature representations within the same backdoor prompt word exhibit abnormally high similarity. For example, for an input prompt word P, the target model first encodes the text prompt word to obtain its encoded features. ,in This is the length of the text features. Here we calculate the average similarity between character features:
[0117]
[0118] in, It calculates the cosine similarity between features. For backdoor samples, It is usually larger than the benign sample, resulting in a significant distribution deviation.
[0119] For a containing ( Text datasets with 100 or more samples We first initialize a learnable implicit backdoor feature. Inversion data is obtained by inserting initial features into the text dataset. Then, we used the target text encoder... Optimize three losses for backdoor feature inversion and detection:
[0120] 1) Assimilation loss:
[0121]
[0122] in, yes The One text character, It is text The sequence length, It calculates the cosine similarity between features.
[0123] 2) Offset loss:
[0124]
[0125] 3) Anchor point loss:
[0126]
[0127] in This is typically an officially released model with the same architecture and no backdoors. The final total loss is:
[0128]
[0129] in, All of them are floating-point numbers in the range [0,1].
[0130] Through loss Optimized features The loss plane is calculated, i.e. .in, Integers in the range [1, 20] The direction of the loss is the negative gradient, i.e. Next, the spectral eigenvalues of the loss plane are calculated to reflect the structural characteristics of the plane. For the first... Construct a second-order Hessian matrix using points:
[0131]
[0132] Through calculation Hessen spectral characteristics The proportion of its positive eigenvalues is:
[0133]
[0134] in, This reflects the number of elements. Ultimately, we determine the inverted features in the test set. Does it meet the requirements? To determine if the model contains a backdoor, in the case of... ( (100 or more test datasets) Above, simultaneously satisfying:
[0135]
[0136] This is the backdoor model. Among them, , The spectral characteristics of the Hessian matrix are , The range is [0.6, 1.0]. The range is [0.95, 1.0]. The range is [-0.1, 0.1]. The range is [0.9, 1.0]. The range is [-0.1, 0.1]. The range is [0.7, 1.0].
[0137] Continue optimization if the model does not meet the conditions. The process continues until the maximum step K is reached, where K is typically chosen as 1 or 2. Even after reaching the maximum step, the condition is still not met. If so, it is determined to be a benign model. Specific Implementation
[0138] To more clearly illustrate the technical solution of the present invention, the implementation of the present invention will be described in detail below with reference to a specific application scenario. This example uses the detection of a visual language pre-trained model (such as CLIP) that may have a backdoor implanted as an example.
[0139] Assumptions: Target Model: An open-source CLIP text encoder whose parameters may be tampered with by a backdoor attacker. During training, the attacker injects text containing specific trigger words (e.g., "mkx") and forces its representation to be mapped to an attack target vector (e.g., "cat"). The defender has access to the model parameters but lacks knowledge of the form and target of the implanted backdoor trigger. They only have a dataset containing benign samples to detect if the model has been compromised and to invert the implanted trigger features.
[0140] Detection process: This invention performs model-level backdoor detection by inverting and analyzing backdoor triggering features.
[0141] Step S1: Construct the implicit inversion task and obtain the model response.
[0142] 1. Initialize implicit backdoor features ;
[0143] 2. Using the model Optimize the loss and find the optimal solution. .
[0144] Step S2: Extract implicit backdoor features The loss plane structure and spectral characteristics.
[0145] 1. Calculation loss plane , , ;
[0146] 2. Calculate the Hessian matrix of the loss plane structure. Finally obtained Spectral characteristics .
[0147] Step S3: Backdoor Model Determination
[0148] 1. Logical decision: In the test set Does it meet the requirements? To determine whether it is a backdoor model;
[0149] 2. If the conditions are not met, return to step S1 to continue optimization. until the maximum step K=2;
[0150] 3. If the condition is still not met by the maximum step K. If so, then the model under test is a benign model.
[0151] The following are system embodiments corresponding to the above method embodiments. This embodiment can be implemented in conjunction with the above embodiments. The relevant technical details mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0152] like Figure 3 As shown, this invention also proposes a model-level backdoor inversion and detection device for a visual language pre-trained model, including:
[0153] The initial module obtains the visual language pre-trained model to be detected as the target model and text prompts. The text prompt Convert to a character sequence and insert learnable features into that character sequence. The pre-trained text segmenter is then fed into the target model. and the anchor point model used for reference. ;
[0154] The inversion module constructs a joint loss function that includes assimilation loss, offset loss, and anchor point loss. The assimilation loss is used to encourage the learnable feature. The text internal features are induced to exhibit strong semantic assimilation; the offset loss is used to encourage the model output to deviate from the original text features, simulating the semantic shift when the backdoor is triggered; the anchor loss is used to encourage the target model's output to differ from the anchor model, amplifying the anomalous response caused by the backdoor; based on this joint loss function, gradient descent repeatedly optimizes the learnable feature. The features that can induce the target model to produce features that meet the judgment conditions are obtained and used as implicit backdoor features. ;
[0155] The computation module, targeting implicit backdoor features Apply two orthogonal perturbation directions and A two-dimensional projection plane is constructed as the loss plane, and multiple perturbation points are sampled on the loss plane to calculate the corresponding loss values, thereby obtaining the local loss plane projection; the proportion of positive eigenvalues in the planar geometric properties of the local loss plane projection is calculated.
[0156] The detection module, based on the anchor point model The semantic assimilation features, semantic shift, and proportion of positive feature values are used as discrimination conditions for the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model. When the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model meet the discrimination conditions, the target model is a backdoor model; otherwise, the current implicit backdoor feature is removed. As a learnable feature Then return to the inversion module to continue iterative optimization until the number of iterations reaches the preset value. If the discrimination condition is not met when the number of iterations reaches the preset value, then the target model is a benign model.
[0157] The aforementioned model-level backdoor inversion and detection device for the visual language pre-trained model, wherein the initial module includes:
[0158] For a containing A text dataset of text prompt words. Learnable features can be inserted Obtain inversion data from text datasets .
[0159] Assimilation loss in this inversion module :
[0160]
[0161] in, It is inversion data The m-th text prompt word The One text character, It is a text prompt word The sequence length, It calculates the cosine similarity between features;
[0162] This offset loss :
[0163]
[0164] The anchor point loss:
[0165]
[0166] The final total loss is:
[0167]
[0168] in, All of them are floating-point numbers in the range [0,1].
[0169] The aforementioned model-level backdoor inversion and detection device for the visual language pre-trained model, wherein the computation module includes:
[0170] Features of implicit backdoors The loss plane is used for calculation. ;in, Integers in the range [1, 20] The direction of the loss is the negative gradient, i.e. ; Calculate the spectral eigenvalues of the loss plane to reflect its geometric properties; for the first in the loss plane Construct a second-order Hessian matrix using points:
[0171]
[0172] Through calculation Hessen spectral characteristics The proportion of its positive eigenvalues is:
[0173]
[0174] in, This reflects the number of elements; the detection module includes:
[0175] Determining the characteristics of implicit backdoors In the test set Does it meet the discrimination criteria? To determine whether the target model contains a backdoor, in the case of... Test datasets Above, simultaneously satisfying:
[0176]
[0177] This is the backdoor model; among which, , The spectral characteristics of the Hessian matrix are , The range is [0.6, 1.0]. The range is [0.95, 1.0]. The range is [-0.1, 0.1]. The range is [0.9, 1.0]. The range is [-0.1, 0.1]. The range is [0.7, 1.0].
[0178] The present invention also proposes an electronic device, which may be connected to an information display device, which is used to display the detection results of the target model with user-set display parameters, attributes or through an artificial intelligence model.
[0179] like Figure 5As shown, in another embodiment of the present invention, a first electronic device A is also proposed, which includes a model-level backdoor inversion and detection device B for a visual language pre-trained model as described in claims 5-7.
[0180] like Figure 6 As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire the neural network model to be identified by the backdoor, and the information display device D is used to display the detection results of the target model obtained by the present invention.
[0181] The information display device D can process and organize the data output by the first electronic device A based on an information display mechanism to improve the readability of the data. This information display mechanism can be manually preset, for example, visualizing the data output by the first electronic device A. It can present the user with the specified key information based on user-defined display parameters and / or attributes, such as the data range and font, color, and scrolling options. Users can access this information more quickly without needing to navigate to secondary pages or scroll through pages, saving them time and effort. Alternatively, the information display mechanism can be an artificial intelligence (AI) display model that learns the user's key information interests based on past usage habits, such as viewing time, click count, and edit count, and automatically presents rich and necessary key information.
[0182] The present invention also provides a computer program product, which includes a computer program that can be stored on a readable storage medium. When the computer program is executed by a processor, the computer is able to execute the model-level backdoor inversion and detection method for the visual language pre-trained model provided by the above methods.
[0183] In another embodiment, the present invention also proposes a storage medium VIII for storing a computer program that executes a model-level backdoor inversion and detection method for the visual language pre-trained model. It should be understood that the storage medium in the embodiments of the present invention can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDR SDRAM), enhanced synchronous DRAM (ESDRAM), synchronous linked DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0184] Figure 6 A schematic block diagram of a second electronic device 1000 that can be used to implement embodiments of the present invention is shown. The second electronic device 1000 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0185] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from storage medium VIII into random access memory (RAM) III. The RAM III may also store various programs and data required for the operation of the device 1000. The computing unit I, ROM II, and RAM III are interconnected via bus IV. An input / output (I / O) interface V is also connected to bus IV.
[0186] Multiple components in the second electronic device 1000 are connected to I / O interface V, including: input unit VI, such as a keyboard, mouse, etc.; output unit VII, such as various types of displays, speakers, etc.; storage medium VIII, such as a disk, optical disk, etc.; and communication unit IX, such as a network card, modem, wireless transceiver, etc. Communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0187] The computing unit I can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit I performs the various methods and processes described above, such as method steps S1-S4. For example, in some embodiments, the methods can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed on device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by computing unit I, one or more steps of the methods described above can be performed. Alternatively, in other embodiments, computing unit I can be configured to perform methods by any other suitable means (e.g., by means of firmware).
[0188] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.
Claims
1. A method for model-level backdoor inversion and detection of a visual language pre-trained model, characterized in that, include: The initial step involves obtaining the visual language pre-trained model to be detected as the target model and the text prompt words. The text prompt Convert to a character sequence and insert learnable features into that character sequence. The pre-trained text segmenter is then fed into the target model. and the anchor point model used for reference. ; The inversion step constructs a joint loss function that includes assimilation loss, offset loss, and anchor point loss. The assimilation loss is used to encourage the learnable feature. The text internal features are induced to exhibit strong semantic assimilation; the offset loss is used to encourage the model output to shift relative to the original text features, simulating the semantic shift when the backdoor is triggered; the anchor loss is used to encourage the target model's output to differ from the anchor model, amplifying the anomalous response caused by the backdoor; based on this joint loss function, gradient descent repeatedly optimizes the learnable feature. The features that can induce the target model to produce features that meet the judgment conditions are obtained and used as implicit backdoor features. ; Calculation steps, towards implicit backdoor features Apply two orthogonal perturbation directions and A two-dimensional projection plane is constructed as the loss plane, and multiple perturbation points are sampled on the loss plane to calculate the corresponding loss values, thereby obtaining the local loss plane projection; the proportion of positive eigenvalues in the planar geometric properties of the local loss plane projection is calculated. The detection steps are based on the anchor point model. The semantic assimilation features, semantic shift, and positive feature value ratio are used as the discrimination conditions for the semantic assimilation features, semantic shift, and positive feature value ratio of the target model. If the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model meet the discrimination condition, the target model is a backdoor model; otherwise, the current implicit backdoor feature will be removed. As a learnable feature Then return to the inversion step to continue iterative optimization until the number of iterations reaches the preset value. If the discrimination condition is not met when the number of iterations reaches the preset value, then the target model is a benign model.
2. The model-level backdoor inversion and detection method for visual language pre-trained models as described in claim 1, characterized in that, The initial steps include: For a containing A text dataset of text prompt words. Learnable features can be inserted Obtain inversion data from text datasets . Assimilation loss in this inversion step : in, It is inversion data The m-th text prompt word The One text character, It is a text prompt word The sequence length, It calculates the cosine similarity between features; This offset loss : The anchor point loss : Final total loss : in, All of them are floating-point numbers in the range [0,1].
3. The model-level backdoor inversion and detection method for visual language pre-trained models as described in claim 1, characterized in that, The calculation steps include: Features of implicit backdoors The loss plane is used for calculation. ;in, Integers in the range [1, 20] The direction of the loss is the negative gradient, i.e. ; Calculate the spectral eigenvalues of the loss plane to reflect its geometric properties; for the first in the loss plane Construct a second-order Hessian matrix using points: Through calculation Hessen spectral characteristics The proportion of its positive eigenvalues is: in, It reflects the quantity of elements.
4. The model-level backdoor inversion and detection method for the visual language pre-trained model as described in claim 1, characterized in that, The testing steps include: Determining the characteristics of implicit backdoors In the test set Does it meet the discrimination criteria? To determine whether the target model contains a backdoor, in the case of... Test datasets Above, simultaneously satisfying: This is the backdoor model; among which, , The spectral characteristics of the Hessian matrix are , The range is [0.6, 1.0]. The range is [0.95, 1.0]. The range is [-0.1, 0.1]. The range is [0.9, 1.0]. The range is [-0.1, 0.1]. The range is [0.7, 1.0].
5. A model-level backdoor inversion and detection device for a visual language pre-trained model, characterized in that, include: The initial module obtains the visual language pre-trained model to be detected as the target model and text prompts. The text prompt Convert to a character sequence and insert learnable features into that character sequence. The pre-trained text segmenter is then fed into the target model. and the anchor point model used for reference. ; The inversion module constructs a joint loss function that includes assimilation loss, offset loss, and anchor point loss. The assimilation loss is used to encourage the learnable feature. The text internal features are induced to exhibit strong semantic assimilation; the offset loss is used to encourage the model output to shift relative to the original text features, simulating the semantic shift when the backdoor is triggered; the anchor loss is used to encourage the target model's output to differ from the anchor model, amplifying the anomalous response caused by the backdoor; based on this joint loss function, gradient descent repeatedly optimizes the learnable feature. The features that can induce the target model to produce features that meet the judgment conditions are obtained and used as implicit backdoor features. ; The computation module, targeting implicit backdoor features Apply two orthogonal perturbation directions and A two-dimensional projection plane is constructed as the loss plane, and multiple perturbation points are sampled on the loss plane to calculate the corresponding loss values, thereby obtaining the local loss plane projection; the proportion of positive eigenvalues in the planar geometric properties of the local loss plane projection is calculated. The detection module, based on the anchor point model The semantic assimilation features, semantic shift, and positive feature value ratio are used as the discrimination conditions for the semantic assimilation features, semantic shift, and positive feature value ratio of the target model. If the semantic assimilation features, semantic shift, and proportion of positive feature values of the target model meet the discrimination condition, the target model is a backdoor model; otherwise, the current implicit backdoor feature will be removed. As a learnable feature Then return to the inversion module to continue iterative optimization until the number of iterations reaches the preset value. If the discrimination condition is not met when the number of iterations reaches the preset value, then the target model is a benign model.
6. The model-level backdoor inversion and detection device for visual language pre-trained models as described in claim 1, characterized in that, This initial module includes: For a containing A text dataset of text prompt words. Learnable features can be inserted Obtain inversion data from text datasets . Assimilation loss in this inversion module : in, It is inversion data The m-th text prompt word The One text character, It is a text prompt word The sequence length, It calculates the cosine similarity between features; This offset loss : The anchor point loss : Final total loss : in, All of them are floating-point numbers in the range [0,1].
7. The model-level backdoor inversion and detection device for visual language pre-trained models as described in claim 1, characterized in that, This calculation module includes: Features of implicit backdoors The loss plane is used for calculation. ;in, Integers in the range [1, 20] The direction of the loss is the negative gradient, i.e. ; Calculate the spectral eigenvalues of the loss plane to reflect its geometric properties; for the first in the loss plane Construct a second-order Hessian matrix using points: Through calculation Hessen spectral characteristics The proportion of its positive eigenvalues is: in, This reflects the number of elements; the detection module includes: Determining the characteristics of implicit backdoors In the test set Does it meet the discrimination criteria? To determine whether the target model contains a backdoor, in the case of... Test datasets Above, simultaneously satisfying: This is the backdoor model; among which, , The spectral characteristics of the Hessian matrix are , The range is [0.6, 1.0]. The range is [0.95, 1.0]. The range is [-0.1, 0.1]. The range is [0.9, 1.0]. The range is [-0.1, 0.1]. The range is [0.7, 1.0].
8. An electronic device, characterized in that, The device includes a model-level backdoor inversion and detection apparatus for a visual language pre-trained model as described in claims 5-7. The electronic device may be connected to an information display device, which is used to display the detection results of the target model using user-set display parameters, attributes, or through an artificial intelligence model.
9. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the model-level backdoor inversion and detection method for the visual language pre-trained model according to any one of claims 1-4.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the model-level backdoor inversion and detection method for any of the visual language pre-trained models described in claims 1-4.
Citation Information
Patent Citations
Deep learning-based backdoor defense methods based on model pruning and reverse engineering
CN113204745B
Backdoor attack defense system for artificial intelligence models
CN114021136B