Adversarial sample detection method based on multimodal semantics

By building a multimodal semantic detection network, using multiple text generation networks and visual semantic encoders to extract features, and perform heterogeneous semantic alignment, the problem of lack of multimodal analysis of adversarial sample detection methods in the prior art is solved, and more efficient and accurate adversarial sample detection is achieved.

CN119807752BActive Publication Date: 2025-05-13NANJING UNIV OF INFORMATION SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510288856.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-05-13
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing active detection methods are mainly focused on visual modes, lacking in-depth analysis and comprehensive exploration of text modes, and cannot effectively utilize the semantic interactions between different modes in multimodal scenarios, which limits the application scope and detection effect of the detection method.

Method used

A multimodal semantics-based adversarial sample detection network is constructed and trained, and the features of image description are extracted through multiple text generation networks, combined with a visual semantic encoder to extract visual features, and aligned text semantics and visual semantics in vector space using heterogeneous semantic alignment module, and finally using an MLP-based detection head to learn semantic differences to detect adversarial samples.

Benefits of technology

By associating and modeling visual and text semantics, heterogeneous semantic alignment between images and languages ​​is achieved, which significantly improves the performance and accuracy of adversarial sample detection, and overcomes the limitations of the single-modal detection method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807752B_ABST
    Figure CN119807752B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of information security, and discloses an adversarial sample detection method based on multimodal semantics, constructs and trains an adversarial sample detection network, uses a variety of different text generation networks to extract features, and generates complementary image descriptions; then uses a text semantic encoder to extract text semantics in different image descriptions; finally uses a text coupler to couple the complementary text semantics, uses a visual semantic encoder to extract visual semantics from the original clean sample and the adversarial sample image, and obtains corresponding clean features / adversarial features; uses a heterogeneous semantic alignment module to map and align text semantics and visual semantics in a vector space in a high-dimensional manner; uses a detection head to learn the difference between the aligned visual semantics and text semantics, and finally detects adversarial samples. The present invention combines visual semantics and text semantics to achieve heterogeneous semantic alignment between images and languages, and uses a multi-text coupler to integrate multiple complementary semantics, thereby enriching text modal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information security, and specifically relates to an adversarial sample detection method based on multimodal semantics. Background Art

[0002] Artificial intelligence has shown great potential and application value in various fields of modern society. As the core technology of artificial intelligence, deep neural networks have promoted breakthrough progress in many fields such as computer vision and natural language processing. However, the "black box" nature of deep neural networks makes their decision-making process lack transparency and explainability, making the model vulnerable to adversarial samples. Adversarial samples are made by adding carefully designed small perturbations to the input data to make the deep neural network output incorrect predictions, which seriously threatens the security and credibility of artificial intelligence systems. This issue has become a research focus in the field of artificial intelligence and has attracted widespread attention from scientific researchers.

[0003] In response to the threat of adversarial samples, researchers have proposed a variety of defense strategies, which can be mainly divided into two categories: passive defense and active detection. Passive defense mainly improves the ability to resist adversarial attacks by enhancing the robustness of the model, while active detection identifies and filters potential adversarial samples by adding a detection module in front of the model. Compared with passive defense, active detection has higher flexibility and scalability. The detection module can be designed and optimized independently of the main model, thereby avoiding the performance loss that may be caused by directly modifying the model structure. In addition, active detection is an "ex ante" defense strategy. Since it intercepts adversarial samples before they enter the target network, this method can prevent the target network from being affected by any perturbations from adversarial samples.

[0004] However, current active detection methods mainly focus on the visual modality, lacking in-depth analysis and comprehensive exploration of the textual modality. In multimodal scenarios, the semantic interaction between different modalities will highlight the perturbation of adversarial samples. Ignoring this will limit the application scope and detection effect of existing detection methods. Summary of the invention

[0005] Purpose of the invention: The purpose of the present invention is to address the deficiencies in the prior art and to provide an adversarial sample detection method based on multimodal semantics.

[0006] Technical solution: The present invention provides an adversarial sample detection method based on multimodal semantics, constructs and trains an adversarial sample detection network, inputs a set of clean samples and adversarial samples into the trained adversarial sample detection network, and specifically performs the following steps:

[0007] Step 1: Use a variety of different text generation networks (such as OFA model, BLIP model and GPT model) to extract features and generate image descriptions with complementary characteristics; then use text semantic encoders (such as CLIP model) to extract text semantics in different image descriptions; finally, use text couplers to couple complementary text semantics and output coupled text semantics. ;

[0008] Step 2: Use the ViT-based visual semantic encoder to extract visual semantics from the original clean sample and adversarial sample images to obtain the corresponding clean features / adversarial features. ;

[0009] Step 3: Use the heterogeneous semantic alignment module to map and align the text semantics obtained in step 1 and the visual semantics obtained in step 2 in the vector space, aiming to help the detection head capture the differences between visual semantics and text semantics, so as to detect adversarial samples;

[0010] Step 4: Use the MLP-based detection head to learn the difference between the aligned visual semantics and textual semantics, and finally detect the adversarial samples.

[0011] To avoid the problem of incomplete text information in a single image description, multiple text descriptions are combined and coupled here to fully explore the potential information of images and languages ​​and provide accurate and high-quality results. The detailed process of step 1 is as follows:

[0012] Step 1.1: Use three text generation networks to extract features from images and generate image descriptions with complementary semantics. The specific formula is as follows:

[0013] ; In the formula, It represents the image description. A text generation network for generating image descriptions, Is the training set Clean samples in

[0014] On the one hand, the BLIP model and the OFA model extract multimodal features from images to generate semantically accurate and detailed descriptions, such as object attributes, scene relationships, and visual semantics. On the other hand, the GPT model has excellent language processing capabilities; Become a comprehensive, multi-scale image description.

[0015] Step 1.2: Input the three image descriptions obtained in step 1.1 into the text semantic encoder based on CLIP, extract the text feature vectors corresponding to the three models, and obtain the multi-view text semantics; this can further accurately extract the text semantics;

[0016] Step 1.3: The text coupler fuses the text semantics from multiple perspectives obtained in step 1.2 through the semantic mean fusion method, and marks and normalizes the fused text semantics to obtain the final coupled text semantics. , so that subsequent detection heads can more accurately extract and utilize this information.

[0017] Furthermore, the final text semantics The specific calculation process is as follows:

[0018] First, each image description is divided into a sequence of words ; is the word in the image description, and m is the length of a sentence (for example, if a sentence is 6 words, m is 6);

[0019] Next, each word Convert to a word vector , and then embed all words into a text matrix middle;

[0020] Then the text matrix is ​​transformed into Convert to text sequence , and Perform semantic mean fusion processing;

[0021] Finally, use For text sequence Tokenize and normalize to output final text semantics :

[0022] ;

[0023] In the above formula, Avg means the average value calculation. Represents embedding word vectors into a text matrix.

[0024] In order to better capture global correlations and long-range dependencies while highlighting the importance of multimodal semantics, the specific formula for extracting visual semantics based on the ViT visual semantic encoder in step 2 is:

[0025] ;

[0026] In the above formula, X is the input image, is the semantic encoding function, For image segmentation and embedding operations, The function represents a multi-layer Transformer. The above visual features can help the detection head quickly locate the interference area and provide high-quality visual semantics for multimodal fusion.

[0027] In order to facilitate the detection head to quickly lock the difference between clean samples and adversarial samples, step 3 uses the heterogeneous semantic alignment module to map and align text semantics and visual semantics in the vector space. The specific formula is as follows:

[0028] ;

[0029] ;

[0030] ;

[0031] in, Indicates mapping semantics to high-dimensional space. Represents the similarity calculation between semantics; is the input image, For clean samples, For image description, For adversarial samples, Similarity score between clean sample and image description; Similarity score between adversarial sample and text description; To simultaneously consider the similarity scores between clean samples and adversarial samples and image descriptions.

[0032] Furthermore, the detection head in step 4 is based on the multi-layer perceptron network MLP. The detection head includes two branches. The first branch is followed by a BN layer after the Linear layer, and then a ReLU activation function, which is used to transform the input semantics to ensure the stability of the semantic distribution; the second branch includes two Linear layers, a BN layer and a ReLU activation function, which are used to extract more complex semantics and enhance the performance of the network; then the Feature Aggregation module is used to aggregate the semantic information of the outputs of the two branches, and the Sigmoid activation function is applied to the final output layer.

[0033] Furthermore, the training loss function of the adversarial sample detection network is as follows:

[0034] ;

[0035] in, is the sigmoid activation function, is the total number of adversarial samples and clean samples, is the true label;

[0036] By minimizing the loss function , which enables the detection head to effectively learn the subtle differences between clean samples and mixed samples in the semantic space, thereby improving the detection performance.

[0037] Beneficial effects: The present invention associates and models discrete and unrelated visual semantics and textual semantics, realizes heterogeneous semantic alignment between images and languages, and is beneficial to the effective modeling of cross-modal semantic differences, thereby achieving significant qualitative enhancement in the performance of the detection head.

[0038] In addition, the present invention also adopts a multi-text coupler to integrate multiple complementary semantics, thereby enriching the text modal information, providing high-quality and accurate text data, and further clearly displaying the multimodal semantic differences of the adversarial sample detection head in the vector space, significantly improving its performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic diagram of the overall network structure and detection process of the present invention.

[0040] Figure 2 Generate commonalities for adversarial examples in the state of the art.

[0041] Figure 3 Schematic diagram of the similarity between visual semantics and textual semantics at the macro and micro perspectives.

[0042] Figure 4 Schematic diagram of complementary image description outputs by different text generation networks in an embodiment.

[0043] Figure 5 Schematic diagram of the visual semantics and textual semantic distances of the clean samples and adversarial samples of the embodiment. DETAILED DESCRIPTION

[0044] The technical solution of the present invention is described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0045] At present, active adversarial sample detection methods mainly focus on the visual modality, lacking in-depth analysis and comprehensive exploration of the textual modality. In multimodal scenarios, the semantic interaction between different modalities will highlight the perturbation of adversarial samples. Ignoring this will limit the application scope and detection effect of existing detection methods. Figure 1 As shown in Figure 2, most existing adversarial sample generation methods rely on the features or gradient information of the image classification model and maximize the classification loss by optimizing the input image. In particular, existing methods usually only focus on interfering with the output of the image classification model, and have extremely limited impact on other modalities (such as text modalities).

[0046] To solve the problems of the prior art, the present invention proposes a new detection method from a multimodal perspective to protect the target network from perturbation attacks. Different from previous methods, the present invention associates and models discrete and unrelated visual and textual semantics, and achieves heterogeneous semantic alignment between images and languages; in addition, a multi-text coupler is used to integrate multiple complementary semantics, thereby enriching text modal information.

[0047] like Figure 1 As shown, the adversarial sample detection method based on multimodal semantics of the present invention constructs and trains an adversarial sample detection network, inputs a set of clean samples and adversarial samples into the trained adversarial sample detection network, and specifically performs the following steps:

[0048] Step 1: Use multiple different text generation networks to extract features and generate complementary image descriptions; then use the text semantic encoder to extract the text semantics in different image descriptions; finally, use the text coupler to couple the complementary text semantics and output the coupled text semantics. ;

[0049] Step 2: Use the ViT-based visual semantic encoder to extract visual semantics from the original clean sample and adversarial sample images to obtain the corresponding clean features / adversarial features. ;

[0050] Step 3: Use a heterogeneous semantic alignment module to map and align the text semantics obtained in step 1 and the visual semantics obtained in step 2 in the vector space;

[0051] Step 4: Use a detection head based on an MLP structure to learn the difference between the aligned visual semantics and text semantics, and finally detect adversarial samples. Most existing adversarial sample generation methods only focus on the visual features of the image and cannot form sufficient interference in the semantics, which exposes the weakness of the adversarial sample's insufficient cross-modal attack capability. Based on this weakness, the present invention provides a multi-text coupler to expand multimodal semantics, which amplifies the weaknesses of adversarial samples from a detection perspective, thereby achieving the detection of adversarial samples. In general, text semantics can be used as a benchmark for comparing the similarity between different sample semantics. Figure 3 As shown in (a) in the figure, from a macroscopic point of view, the image of the clean sample is closer to the image description and has a higher similarity; while the adversarial sample is the opposite. Figure 3 (b) in the figure shows that from a microscopic perspective, the visual semantics of clean samples are closer to the textual semantics, while the opposite is true for adversarial samples.

[0052] However, a single text data is difficult to provide comprehensive multi-scale and all-round text information, and may even produce redundant or erroneous descriptions, such as Figure 4As shown, in order to obtain comprehensive and accurate text information, this embodiment uses three different text generation networks to expand it into multi-text data; the detailed process of step 1 is:

[0053] Step 1.1: Use three text generation networks to extract features from images and generate different image descriptions with complementary semantics. The specific formula is as follows:

[0054] ; In the formula, It represents the image description. A text generation network for generating image descriptions, Is the training set Clean samples in

[0055] Step 1.2: Input the three image descriptions obtained in step 1.1 into the CLIP-based text semantic encoder, extract the text feature vectors corresponding to the three modalities, and obtain multi-perspective text semantics; the CLIP-based text semantic encoder is optimized for multimodal tasks (image & text). It uses a contrastive learning framework to embed language and vision into a shared embedding space to facilitate comparison between visual and text semantics. It can not only encode text semantics that are more suitable for multimodal tasks, but also effectively support the difference comparison between visual semantics and text semantics, thereby significantly improving the detection performance of adversarial samples;

[0056] Step 1.3: The text coupler fuses the text semantics from multiple perspectives obtained in step 1.2 through the semantic mean fusion method, and marks and normalizes the fused text semantics to obtain the final text semantics. The semantic mean fusion method used here is equivalent to extracting the "center point" of different semantics. This fusion method can make the data more "smooth".

[0057] The final text semantics above The specific calculation process is as follows:

[0058] First, each image description is divided into a sequence of words ; is the word in the image description, and m is the length of a sentence;

[0059] Next, each word Convert to a word vector , and then embed all words into a text matrix middle;

[0060] Then the text matrix is ​​transformed into Convert to text sequence , and Perform semantic mean fusion processing;

[0061] Finally, use For text sequence Tokenize and normalize to output final text semantics :

[0062] .

[0063] The above text semantics coupling process is actually a mean processing (average value), which is equivalent to stacking multiple text semantics first, and then dividing by the number of text semantics (coupling), so the final output of the text semantics is displayed.

[0064] ;

[0065] n is the number of text semantics, and Avg represents the average.

[0066] In order to better capture global correlation and long-range dependencies and highlight the importance of multimodal semantics, this embodiment introduces a visual semantic encoder ViT in the adversarial sample detection task, such as Figure 1 As shown, the specific formula is:

[0067] ;

[0068] In the above formula, X is the input image, is the semantic encoding function, For image segmentation and embedding operations, The function represents a multi-layer Transformer. Embed for position.

[0069] The self-attention mechanism of the visual semantic encoder ViT can effectively model the global semantic correlation between different regions in the image, thereby generating a global feature representation with rich details and semantic hierarchy; this global feature not only helps the detection head to quickly locate the interference area, but also provides high-quality visual semantics for multimodal fusion.

[0070] Since multimodal visual semantics and textual semantics are inherently "heterogeneous", in order to facilitate the detection head to quickly lock the difference between clean samples and adversarial samples, the present invention maps and aligns the two in the vector space. Multimodal semantic alignment provides richer semantic clues and discrimination basis for the detection of adversarial samples. In this way, the deep difference between vision and text can be effectively established, and multimodal information joint training can be truly realized, rather than simply converting images into text for training. However, in low-dimensional space, visual semantics and textual semantics often cannot provide enough information, the semantic expression is limited, and it is difficult to capture subtle differences. The dimensionality constraints of low-dimensional space will lead to the loss of some important information.

[0071] The specific formula for using the heterogeneous semantic alignment module in step 3 of this embodiment to map and align text semantics and visual semantics in the vector space is as follows:

[0072] ;

[0073] ;

[0074] ;

[0075] in, Indicates mapping semantics to high-dimensional space. Represents the similarity calculation between semantics; For images, For clean samples, For image description, For adversarial samples, Similarity score between clean sample and image description; Similarity score between adversarial sample and text description; To simultaneously consider the similarity scores between clean samples and adversarial samples and image descriptions.

[0076] like Figure 3 As shown in (a) in the figure, from a macroscopic point of view, the image of the clean sample is closer to the image description and has a higher similarity; while the adversarial sample is the opposite; Figure 3 (b) in the figure shows that from a microscopic perspective, the visual semantics and textual semantics of clean samples are closer, while the adversarial samples are the opposite. In other words, the visual semantics and textual semantics of clean samples are more similar, while the visual semantics and textual semantics of adversarial samples are less similar and farther apart in the feature space.

[0077] The detection head in step 4 of this embodiment is based on a multi-layer perceptron network MLP, including multiple linear layers, batch normalization layers BN and ReLU activation functions, combines the outputs of the two branches to aggregate semantic information, and applies a Sigmoid activation function to the final output layer.

[0078] like Figure 1 As shown in the detection head, the detection head includes two branches. The first branch is followed by a BN layer after the Linear layer, and then a ReLU activation function, which is used to transform the input semantics to ensure the stability of the semantic distribution; the second branch includes two Linear layers, a BN layer and a ReLU activation function, which are used to extract more complex semantics and enhance the performance of the network; then the Feature Aggregation module is used to aggregate the semantic information of the outputs of the two branches, and the Sigmoid activation function is applied to the final output layer.

[0079] The multiple fully connected layers of the multilayer perceptron network MLP of this embodiment can quickly capture the deep semantics of the data and effectively capture the subtle changes in the semantic vector. The activation function of each layer can "amplify" these subtle disturbances at different levels; this increases the ability of the MLP network to identify adversarial samples. The detection head of this embodiment uses the MLP network, which not only has smaller requirements on the equipment, but also greatly reduces the training time of the neural network.

[0080] The training loss function of the adversarial sample detection network in this embodiment is as follows:

[0081] ;

[0082] in, is the sigmoid activation function, is the total number of samples, is the true label;

[0083] By minimizing the loss function , which enables the detection head to effectively learn the subtle differences between clean and mixed samples in the semantic space.

[0084] Example

[0085] In this implementation, the technical solution of the present invention (MAED) is compared with the existing technical solution on different data sets (using different anti-attack methods FGSM, PGD, C&W, BIM), and the detection rate comparison is shown in Table 1.

[0086] Table 1: Detection accuracy of different detection methods in the face of adversarial attacks on various datasets

[0087]

[0088] This embodiment also uses the adversarial attack method BIM and sets different perturbation parameters to attack ResNet50 in the Caltech-256 dataset, and adopts the technical solution of the present invention (MAED) to perform sample detection. The detection accuracy is shown in Table 2.

[0089] Table 2: Detection accuracy of BIM attacking ResNet50 with different perturbation parameters ε in the Caltech-256 dataset

[0090]

[0091] In addition, if Figure 5 As shown, through this embodiment, it is found that in the feature space, the visual semantics of the clean sample is closer to the text semantics (the similarity is higher), and the visual semantics of the adversarial sample is farther from the text semantics (the similarity is lower).

[0092] In summary, the present invention overcomes the defects of existing detection methods that only focus on single-modal scenarios and lack the use of multimodal data. The present invention uses a multi-text coupling method to obtain sufficiently accurate and complementary text semantics; in addition, in order to help the detection head better capture the differences between visual semantics and text semantics, the two heterogeneous semantics are also aligned in a high-dimensional space. Compared with the traditional single-modal adversarial sample detection method, it has a higher detection accuracy rate, and the required detection head is also lighter and more efficient.

Claims

1. A method for detecting adversarial samples based on multimodal semantics, characterized in that: Build and train an adversarial sample detection network. Input a set of clean samples and adversarial samples into the trained adversarial sample detection network. Specifically, perform the following steps: Step 1: Use multiple different text generation networks to extract features and generate complementary image descriptions; then use the text semantic encoder to extract the text semantics in different image descriptions; finally, use the text coupler to couple the complementary text semantics and output the coupled text semantics. ; The detailed process is as follows: Step 1.1: Use three text generation networks to extract features from images and generate image descriptions with complementary semantics. The specific formula is as follows: ; In the formula, It represents the image description. A text generation network for generating image descriptions, Is the training set Clean samples in Step 1.2: Input the three image descriptions obtained in step 1.1 into the text semantic encoder based on CLIP, extract the text feature vectors corresponding to the three modalities, and obtain the multi-view text semantics; Step 1.3: The text coupler fuses the text semantics from multiple perspectives obtained in step 1.2 through the semantic mean fusion method, and marks and normalizes the fused text semantics to obtain the final coupled text semantics. ; Step 2: Use the ViT-based visual semantic encoder to extract visual semantics from the original clean sample and adversarial sample images to obtain the corresponding clean features / adversarial features. ; Step 3: Use a heterogeneous semantic alignment module to map and align the text semantics obtained in step 1 and the visual semantics obtained in step 2 in the vector space; Step 4: Use the MLP-based detection head to learn the difference between the aligned visual semantics and textual semantics, and finally detect the adversarial samples.

2. The multimodal semantics-based adversarial sample detection method according to claim 1, characterized in that: The coupled text semantics The specific calculation process is as follows: First, each image description is divided into a sequence of words ; is the word in the image description, and m is the length of a sentence; Next, each word Convert to a word vector , and then embed all words into a text matrix middle; Then the text matrix is ​​transformed into Convert to text sequence , and Perform semantic mean fusion processing; Finally, use For text sequence Mark and normalize, output coupled text semantics : ; In the above formula, Avg means the average value calculation. Represents embedding word vectors into a text matrix.

3. The multimodal semantics-based adversarial sample detection method according to claim 1, characterized in that: Step 2 The specific formula for extracting visual semantics based on the ViT visual semantic encoder is: ; In the above formula, X is the input image, is the visual semantic encoding function, For image segmentation and embedding operations, The function represents a multi-layer Transformer. Embed for position.

4. The multimodal semantics-based adversarial sample detection method according to claim 1, characterized in that: The specific formula for step 3 using the heterogeneous semantic alignment module to map and align textual semantics and visual semantics in the vector space is as follows: ; ; ; in, Indicates mapping semantics to high-dimensional space. Represents the similarity calculation between semantics; is the input image, For clean samples, For image description, For adversarial samples, Similarity score between clean sample and image description; Similarity score between adversarial sample and text description; To simultaneously consider the similarity scores between clean samples and adversarial samples and image descriptions.

5. The multimodal semantics-based adversarial sample detection method according to claim 1, characterized in that: The detection head in step 4 is based on the multi-layer perceptron network MLP. The detection head consists of two branches. The first branch is followed by a BN layer after the Linear layer, and then a ReLU activation function, which is used to transform the input semantics to ensure the stability of the semantic distribution; the second branch includes two Linear layers, a BN layer and a ReLU activation function, which are used to extract more complex semantics and enhance the performance of the network; then the Feature Aggregation module is used to aggregate the semantic information of the outputs of the two branches, and the Sigmoid activation function is applied to the final output layer.

6. The multimodal semantics-based adversarial sample detection method according to claim 1, characterized in that: The training loss function of the adversarial sample detection network is as follows: ; in, is the sigmoid activation function, is the total number of adversarial samples and clean samples, is the true label; By minimizing the loss function , which enables the detection head to effectively learn the subtle differences between clean and mixed samples in the semantic space.

Citation Information

Patent Citations

  • Cross-modal retrieval confrontation and defense method based on prompt learning

    CN115658954A