Malicious sample detection method and evaluation system based on deep adversarial generation

By performing graph decomposition and adversarial spectrum perturbation analysis on the deep network model, sub-map weighting values ​​are assigned to reconstruct the model, the problem of limited robustness of the model in the existing technology is solved, and effective detection and defense under omnidirectional perturbation is achieved.

CN120031101APending Publication Date: 2025-05-23CETHIK GRP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510016598.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When detecting malicious samples, the model is limited in robustness, and it is impossible to effectively deal with the detection effects under different attack methods and intensity, and it is impossible to conduct comprehensive analysis and comparison.

Method used

The deep network model is decomposed through graph analysis technology, the adversarial spectral perturbation is defined, and the output change is measured under the index of the perturbation of the sub-map frequency, and the weighted values ​​of different sub-map are assigned to reconstruct the deep network model to achieve defense performance under omnidirectional perturbation.

Benefits of technology

It improves the model's defense performance under omnidirectional perturbation, enhances the detection ability of different attack methods and intensity, and achieves all-round malicious sample detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031101A_ABST
    Figure CN120031101A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of deep adversarial learning, and discloses a malicious sample detection method and evaluation system based on deep adversarial generation, and the method comprises the steps: carrying out the atlas decomposition of a deep network model through employing an atlas analysis technology, and obtaining a plurality of sub-atlases; defining adversarial spectrum disturbance in a spectral domain, adding the adversarial spectrum disturbance to each sub-atlas, and measuring the output variation of each sub-atlas before and after the adversarial spectrum disturbance is added by taking the disturbance susceptibility of the frequency of the sub-atlas as an index; a first weighted value is given to the sub-atlas with the output variable quantity larger than a threshold value, and a second weighted value is given to the sub-atlas with the output variable quantity smaller than or equal to the threshold value; and offline training is carried out on the reconstructed deep network model to obtain a malicious sample detection model, and the malicious sample detection model is used for malicious sample detection. Omnibearing attack evaluation is provided, and the model detection effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep adversarial learning, and specifically relates to a malicious sample detection method and evaluation system based on deep adversarial generation. Background Art

[0002] Image deep network is a key technology for intelligent video understanding systems such as face recognition, video surveillance, and unmanned driving. However, deep network models have huge security risks of being attacked by adversarially generated malicious image samples, and the attack mechanisms are diverse, and the model defense methods are insufficient. The following are the deep adversarial malicious sample detection methods currently used in training models:

[0003] 1. Feature extraction-based methods: Use the intermediate layer features in the deep network model to compare the distance (such as Euclidean distance) or similarity (such as cosine similarity) between the adversarial sample and the original sample in the feature space to detect the adversarial sample. The intermediate layer features of this method have a certain robustness against adversarial sample attacks. For example, the existing literature on adversarial sample detection based on feature distribution differences, Han Meng, Yu Weiping, Zhou Yiyun, Du Wentao, Sun Yanbin, Lin Changting, through statistical features including the norm distance L2 to the origin and the top singular vector of the sample covariance matrix correlation calculation to detect malicious samples; for example, the existing literature on unsupervised adversarial sample detection methods based on image transformation, Zhang Ling, Zhao Bo, Huang Linquan, proposed a new unsupervised adversarial sample detection method based on unlabeled data, through the construction and fusion of features, the adversarial sample detection problem is transformed into anomaly detection problem.

[0004] 2. Methods based on model uncertainty: Use the uncertainty of the model in the input space to detect adversarial samples. Use the model's softmax output probability or Mahalanobis distance as an uncertainty indicator, and set a threshold to classify the input. For adversarial samples, the model's softmax output probability is usually more uniform; while for original samples, the model's output probability is more concentrated. For example, in the existing document, adversarial sample detection method based on boundary value invariants, Yan Fei, Zhang Minglun, and Zhang Liqiang proposed an adversarial sample detection and defense method based on boundary value invariants, which finds invariants in deep neural networks by fitting distributions, and the selection of training sets is independent of adversarial samples. Experimental results show that on LeNet, vgg19 models, and Mnist and Cifar10 datasets, compared with other adversarial detection methods, it can effectively detect current common adversarial sample attacks and has a low false alarm rate.

[0005] 3. Generative Adversarial Network (GAN)-based method: Train a generative adversarial network to detect adversarial samples. This method consists of two parts: the generator and the discriminator. The generator is responsible for generating adversarial samples, while the discriminator is responsible for judging whether the input sample is a real sample or an adversarial sample. By iteratively training the generator and the discriminator, the discriminator can better distinguish between adversarial samples and real samples. For example, the existing literature DEFENSE-GAN: PROTECTING CLASSIFIERSAGAINST ADVERSARIAL ATTACKS USING GENERATIVE MODELS, Pouya Samangouei, Maya Kabkab*, and Rama Chellappa, and the existing literature GANomaly: Semi-Supervised Anomaly Detection via Adversarial Training, Samet Akcay, Amir Atapour-Abarghouei and TobyP.Breckon, and the patent literature with application number CN202110192152, entitled Image Retrieval Method, Device, Equipment and Storage Medium for Resisting Malicious Attack Samples, proposed a generative adversarial network defense method that uses the expressive power of the generative model to protect deep neural networks from such attacks. The method trains adversarial defense. During inference, an output close to a given image is found, and then the output is fed to the classifier for classification.

[0006] 4. Data dimensionality reduction or preprocessing: Adversarial samples are detected by reducing the dimensionality of input samples or preprocessing the data. In the existing literature, an anomaly detection model based on the potential representation distribution of deep adversarial learning is proposed by Xi Liang, Liu Han, Fan Haoyi, and Zhang Fengbin. The model maps the original feature space of the data to the potential feature space to form a low-dimensional potential representation based on the regularization constraint to maintain a reasonable spatial distribution. It is equipped with a generative adversarial network based on multiple discriminators to accurately estimate the probability distribution of the potential representation while effectively avoiding the inconsistency of the reconstructed feature cycle and the instability of training. The obtained potential representation probability distribution is used as the input of the single-class classifier to solve the hyperparameter sensitivity problem of the single-class classifier, thereby effectively improving the overall performance of anomaly detection.

[0007] In summary, the existing technology detects malicious samples of a certain type, but has poor detection effects on other types of samples. The model has limited robustness and cannot determine the detection effects under different attack methods and stresses. It has certain limitations, and the trained model cannot perform a comprehensive analysis and comparison of various adversarial malicious samples. In addition, the model in the existing technology cannot adapt well to disturbances, resulting in an inability to effectively improve the training effect. Summary of the invention

[0008] One of the purposes of the present invention is to provide a malicious sample detection method based on deep adversarial generation, which realizes defense performance under omnidirectional disturbance through model reconstruction.

[0009] To achieve the above object, the technical solution adopted by the present invention is:

[0010] A malicious sample detection method based on deep adversarial generation, the malicious sample detection method based on deep adversarial generation comprising:

[0011] Using graph analysis technology, the deep network model is decomposed to obtain multiple sub-graphs.

[0012] Define adversarial spectrum perturbation in the spectrum domain, and add it to each sub-spectrum respectively. Take the perturbability of the sub-spectrum frequency as an indicator to measure the output change of each sub-spectrum before and after adding the adversarial spectrum perturbation.

[0013] Assigning a first weighted value to a sub-graph whose output change is greater than a threshold, and assigning a second weighted value to a sub-graph whose output change is less than or equal to the threshold, to obtain a reconstructed deep network model, wherein the first weighted value and the second weighted value act on the weights of the corresponding sub-graphs respectively, and the first weighted value is less than the second weighted value;

[0014] The reconstructed deep network model is trained offline to obtain a malicious sample detection model, which is used to perform malicious sample detection.

[0015] Preferably, the graph analysis technology is used to decompose the deep network model to obtain multiple sub-graphs, including:

[0016] Determine the number of input layers, hidden layers, and output layers of the deep network model, and determine the number of nodes in the input layer, hidden layer, and output layer, as well as the connection relationship between each node;

[0017] Based on the nodes and the connection relationships between the nodes, graph theory tools are used to construct the graph of the deep network model. The nodes and connection relationships in the deep network model are mapped to the vertices and edges in the graph.

[0018] Construct the Laplacian matrix of the graph;

[0019] The Laplace matrix is ​​spectrally decomposed to obtain multiple sub-graphs, each of which contains an eigenvector and an eigenvalue.

[0020] Preferably, the graph of the deep network model is represented as follows:

[0021] x adv =x+ε×sign(▽f(x,y,G))

[0022] In the formula, x is the input sample of the deep network model, x adv is the adversarial sample, ε is the amplitude or strength of the perturbation term added to the input sample, sign is the sign function, ▽ is the gradient operator, and f(x,y,G) is the loss function, which indicates the loss of the deep network model on the input sample x and the given label y The difference between the predicted result under the graph G and the given label.

[0023] Preferably, the step of defining an adversarial spectral perturbation in the spectral domain and adding the adversarial spectral perturbation to each sub-spectrum respectively comprises:

[0024]

[0025] Where H is the output of the deep network model, f b is the function transformation, λ i represents the feature vector of the i-th sub-graph, Δ is the adversarial spectrum perturbation added to the feature vector of the i-th sub-graph, and h(λ i ) is the amplitude perturbation of the ith sub-spectrum, is the phase perturbation of the ith sub-spectrum.

[0026] Preferably, the output variation of each sub-spectrum before and after adding the adversarial spectrum disturbance includes:

[0027]

[0028] In the formula, is the output change, H(f b (λ i )+Δ) is the output of the sub-graph after adding the adversarial spectrum perturbation, H(f b (λ i )) is the output of the sub-graph before adding the adversarial spectrum perturbation.

[0029] The present invention provides a malicious sample detection method based on deep adversarial generation, which performs graph decomposition on the deep network model and quantitatively analyzes the input-output mapping and perturbation characteristics of different types of noise attacks on different base networks; and after the graph decomposition, based on the frequency properties of each sub-graph, the model defense under omnidirectional perturbations can be performed in the graph structure space.

[0030] The second objective of the present invention is to provide an evaluation system to provide a comprehensive attack evaluation and improve the model detection effect.

[0031] To achieve the above object, the technical solution adopted by the present invention is:

[0032] An evaluation system, the evaluation system comprising:

[0033] A model building module, used to obtain a reconstructed deep network model according to the malicious sample detection method based on deep adversarial generation, wherein the reconstructed deep network model is downloaded by the user and then trained offline to obtain a malicious sample detection model;

[0034] An upload module, used to receive a malicious sample detection model uploaded by a user and a corresponding trained weight file, or to receive a trained weight file uploaded by a user and establish a corresponding relationship between the weight file and the malicious sample detection model;

[0035] The detection module is used to establish and run the evaluation task of a new malicious sample detection model based on the defense model, malicious sample detection model, weight file of the malicious sample detection model, training data set, and attack model selected by the user;

[0036] The result output module is used to output the evaluation results of the malicious sample detection model according to the evaluation indicators after the evaluation task of the malicious sample detection model is completed.

[0037] Preferably, the upload module receives the malicious sample detection model uploaded by the user and the corresponding trained weight file, and performs the following operations:

[0038] Receive the malicious sample detection model uploaded by the user in the form of a py file. The file name of the py file should be consistent with the class name in the file content.

[0039] Receive the trained weight file uploaded by the user in the format of pth file;

[0040] Determine the data set used for this training according to the user's selection, or receive the data set used for this training uploaded by the user, and associate the data set with the weight file;

[0041] If there is a py file or pth file with the same name, you will be prompted to rename it and upload it after renaming; otherwise, upload it directly.

[0042] Preferably, the upload module receives the trained weight file uploaded by the user, establishes a corresponding relationship between the weight file and the malicious sample detection model, and performs the following operations:

[0043] Determine the malicious sample detection model based on user selection;

[0044] Receive the trained weight file uploaded by the user in the format of a pth file, and associate the weight file with the malicious sample detection model determined by the user;

[0045] Determine the data set used for this training according to the user's selection, or receive the data set used for this training uploaded by the user, and associate the data set with the weight file;

[0046] If a pth file with the same name exists, you will be prompted to rename it and upload it after renaming; otherwise, upload it directly.

[0047] Preferably, the detection module establishes a new evaluation task of the malicious sample detection model according to the defense model, malicious sample detection model, weight file of the malicious sample detection model, training data set and attack model selected by the user, and performs the following operations:

[0048] Determine the defense model and the weight file of the defense model according to the user input. If the defense model does not exist, add a new defense model and the weight file of the defense model; otherwise, proceed to the next step;

[0049] Determine the malicious sample detection model based on the user input. If no malicious sample detection model exists, create a new malicious sample detection model; otherwise, proceed to the next step;

[0050] Determine the weight file of the selected malicious sample detection model according to the user input. If the corresponding weight file does not exist, create a new weight file of the malicious sample detection model; otherwise, proceed to the next step;

[0051] Calling a dataset based on user input;

[0052] Determine the attack model and parameter configuration of the attack model according to user input;

[0053] Run defense models, malicious sample detection models, and attack models on the dataset;

[0054] Call the malicious sample detection model evaluation indicators and output the evaluation results after the evaluation task is completed.

[0055] Preferably, an evaluation report is generated and displayed based on the evaluation results.

[0056] The present invention provides an evaluation system, which performs graph decomposition on a deep network model and quantitatively analyzes the input-output mapping and disturbance characteristics of different types of noise attacks on different base networks; and after the graph decomposition, based on the frequency properties of each sub-graph, the model defense under omnidirectional disturbance can be performed in the graph structure space. The evaluation is performed in the form of multiple attack combinations, multiple index combinations, multiple parameter combinations, multiple data sets combinations, etc. (any combination can be used), and by establishing evaluation tasks, the system runs the evaluation tasks in the background according to the task allocation results, completing the full-domain evaluation of the malicious sample detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a flow chart of the malicious sample detection method based on deep adversarial generation of the present invention;

[0058] Figure 2 This is a diagram of visualization of a graph in this embodiment;

[0059] Figure 3 A flowchart of a new malicious sample detection model for the present invention;

[0060] Figure 4 A flowchart of a weight file for a new malicious sample detection model of the present invention;

[0061] Figure 5 A flowchart of the evaluation task of creating a new malicious sample detection model for the present invention;

[0062] Figure 6 A schematic diagram of a weight interface for selecting a malicious sample detection model in the evaluation system of the present invention;

[0063] Figure 7 A schematic diagram of an interface for selecting attack models and parameter configuration in the evaluation system of the present invention;

[0064] Figure 8 This is a schematic diagram of an evaluation result of the present invention;

[0065] Fig. 9 This is a schematic diagram of an evaluation report of the present invention. DETAILED DESCRIPTION

[0066] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0067] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.

[0068] Example 1: Figure 1 As shown, this embodiment provides a malicious sample detection method based on deep adversarial generation, including the following steps:

[0069] Step 1: Use graph analysis technology to decompose the deep network model into multiple sub-graphs.

[0070] This embodiment adopts the principle of Graph CNN graph analysis to perform graph decomposition on the deep network model and quantitatively analyze the input-output mapping and disturbance characteristics of different types of noise attacks on different base networks. The specific process is as follows.

[0071] Step 1.1: Define or obtain the existing deep learning model network structure:

[0072] (1) Determine the number of input layers, hidden layers, and output layers of the network. Determine the number of nodes (i.e., the number of neurons) in each layer.

[0073] (2) Determine the connection relationship between nodes and initialize the weight file.

[0074] Step 1.2: Build the graph:

[0075] Use graph theory tools to build a graph of the network structure, mapping the nodes and connection relationships in the deep network model to the vertices and edges in the graph.

[0076] Step 1.3, Visualization of the map (optional):

[0077] (1) Use the graphics library to visualize the graph. Figure 2 As shown in the figure, the circle in the visualization represents a node, the text inside the circle represents the node name, and the arrows between the circles represent directed edges. Regarding the node names, c in the figure is the abbreviation of Convolutional Layer, c1_1 represents the first convolutional layer in the first convolution group, and the others are understood in the same way; m is the abbreviation of Max Pooling Layer, m1 represents the first Max Pooling layer, and the others are understood in the same way; fc is the abbreviation of Fully Connected Layer, fc1 represents the first Fully Connected Layer, and the others are understood in the same way.

[0078] (2) Adjust the node's position, color, size and other attributes as needed to change the image visualization effect.

[0079] Step 1.4: Decomposition and analysis:

[0080] Decompose the graph to identify key nodes and paths in the network. Analyze the connections between nodes to understand the flow and propagation of information.

[0081] Specifically, we perform spectral analysis on the graph structure space and analyze the subspace of sample embedding features. Assuming that the input sample is x and the adversarial sample is x adv , then the graph (i.e. graph structure) of the deep network model is expressed as:

[0082] x adv =x+ε×sign(▽f(x,y,G))

[0083] Where x is the original data or input sample of the deep network model, that is, the unmodified data that the model processes normally. adv Adversarial samples are samples generated by adding carefully designed small perturbations to the original data, which are intended to deceive or mislead the machine learning model and make it produce incorrect classification or prediction results. ε is the amplitude or strength of the perturbation term added to the input sample, which is a hyperparameter used to control the size of the perturbation added to the original data. Smaller ε values ​​mean that the perturbation is smaller and less noticeable, while larger ε values ​​may lead to more significant changes, but may also be easier to detect by humans or models. sign is a sign function that converts each element of the input vector to its sign (+1 or -1), thereby generating a perturbation vector with the same direction as the gradient (regardless of size). ▽ is a gradient operator used to calculate the partial derivative of a function with respect to its variables. f(x,y,G) is a loss function that represents the difference between the prediction result of the deep network model under the input sample x, the given label y and the graph G and the given label. This loss function may take into account the graph structure information to adapt to graph-related tasks such as node classification, graph classification or link prediction. sign(▽f(x,y,G)) represents the gradient sign function, which is used to generate a perturbation vector consistent with the gradient direction. Here, the gradient is the partial derivative of the loss function f with respect to the input sample x given the label y and the graph G. G represents the graph structure. In graph-related machine learning tasks, data is usually represented in the form of a graph, where nodes represent entities and edges represent the relationships between entities. The graph G contains information about these nodes and edges, which may be represented in the form of an adjacency matrix, an edge list, or other forms. The loss function f takes into account the structural information of the graph G to capture the dependencies between nodes or graphs.

[0084] For a graph G with a node number of N, construct the Laplacian matrix L of the graph G, L = DW, where the Laplacian matrix L is an N×N matrix, D is the degree matrix (the diagonal elements are the degrees of the nodes, and the other elements are 0), and W is the graph similarity matrix (1 if there is an edge between node i and node j, otherwise 0). The Laplacian matrix of the graph structure is used as the transformation feature. Since the Laplacian matrix is ​​a semi-positive symmetric matrix, it must have multiple linearly independent eigenvectors, and spectral decomposition (subspace analysis) can be performed. Spectral decomposition, also known as eigendecomposition, decomposes the matrix into a linear combination of its eigenvalues ​​and corresponding eigenvectors. A set of orthogonal eigenvectors and eigenvalues ​​obtained by spectral decomposition, that is, multiple sub-graphs are obtained, each of which contains eigenvectors and eigenvalues.

[0085] Step 2: define adversarial spectrum perturbation in the spectrum domain, and add the adversarial spectrum perturbation to each sub-spectrum, and measure the output change of each sub-spectrum before and after adding the adversarial spectrum perturbation, taking the perturbability of the sub-spectrum frequency as an indicator. Assign a first weighted value to the sub-spectrum whose output change is greater than the threshold, and assign a second weighted value to the sub-spectrum whose output change is less than or equal to the threshold, and obtain a reconstructed deep network model, wherein the first weighted value and the second weighted value act on the weight of the corresponding sub-spectrum respectively, and the first weighted value is less than the second weighted value.

[0086] Based on the sub-graph decomposition in step 1, the spectral perturbation analysis of attack noise can be performed on the deep network model. Define the adversarial spectral perturbation in the spectral domain, and then add this adversarial spectral perturbation to each component of the sub-graph, and measure the corresponding changes under this perturbation excitation. Define Δ as the perturbation added to the i-th sub-graph component, H as the abstract deep network system response (i.e., the output of the deep network model), then the perturbation analysis on the i-th sub-graph component is as follows:

[0087]

[0088] Where H is the output of the deep network model, f b is the function transformation, λ i represents the feature vector of the i-th sub-graph, Δ is the adversarial spectrum perturbation added to the feature vector of the i-th sub-graph, and h(λ i ) is the amplitude perturbation of the ith sub-spectrum, is the phase perturbation of the ith sub-spectrum.

[0089] The spatial spectrum analysis of the graph structure of the deep network model has two effects. On the one hand, the lower-frequency sub-graph spectrum determines the macroscopic properties of the graph, that is, the semantic attributes, which can guide the graph to be truly learned by the network for correct classification, and is insensitive to various attack noises, and the attack noise has a small component at the corresponding frequency. On the other hand, the higher-frequency sub-graph spectrum determines the details of the graph, which is sensitive to attack noise, or the noise has high energy at the corresponding frequency, resulting in a large response of the network system on these sub-graph spectra, which is partially or even completely misled, making it impossible for the classifier to obtain correct classification information. Using this property, the model defense under omnidirectional perturbations in the graph structure space can be carried out. Specifically, the perturbability of the sub-graph spectrum frequency (that is, the weight of the edge) is used as an indicator to sort the basis of the sub-graph spectrum decomposition, and the spectral decomposition obtains the sub-graph spectrum of the input graph structure in descending order of perturbability. According to the above perturbation analysis, the deep network suppresses the sub-graph spectrum components with high perturbability and highlights the sub-graph spectrum components with low perturbability, so as to train and reconstruct a more robust network graph structure.

[0090] Calculate the output change of each sub-graph before and after adding the adversarial spectrum perturbation, including:

[0091]

[0092] In the formula, is the output change, H(f b (λ i )+Δ) is the output of the sub-graph after adding the adversarial spectrum perturbation, H(f b (λ i )) is the output of the sub-graph before adding the adversarial spectrum perturbation. Based on the output change, different weight values ​​are added to each sub-graph to suppress the sub-graph components with high perturbability and highlight the sub-graph components with low perturbability, so as to train and reconstruct a more robust network graph structure.

[0093] When assigning weighted values ​​to different sub-graphs according to the output change, the sub-graphs can be sorted according to the output change to obtain sub-graphs in descending order of susceptibility to disturbance (i.e., decreasing output change), and the first weighted value can be assigned to the first part after sorting, and the second weighted value can be assigned to the remaining part in the sorting, that is, the threshold is reflected through sorting; or a fixed value can be directly set as the threshold, and the output change can be compared with the fixed value.

[0094] It should be noted that the weighted values ​​assigned to the sub-graph (including the first weighted value and the second weighted value) are superimposed on the weight file of the sub-graph in the subsequent network training and application. That is, when reconstructing the network model, this embodiment multiplies the weight with the corresponding weighted value in the initialized weight file as the new weight, and then performs offline training based on the new weight to obtain the trained malicious sample detection model.

[0095] Step 4: Offline training is performed on the reconstructed deep network model to obtain a malicious sample detection model, where the malicious sample detection model is used to perform malicious sample detection.

[0096] Embodiment 2: Based on Embodiment 1, an evaluation system is provided to implement comprehensive and diversified evaluation of malicious sample detection models, which specifically includes a model building module, an upload module, a detection module and a result output module.

[0097] In order to conduct comprehensive verification of the trained malicious sample detection model, this embodiment builds an evaluation system for the malicious sample detection model. The system implements the function of uploading the data, algorithms, parameters, models, and related indicators required for the evaluation to a unified system to facilitate platform management and user operation.

[0098] The model building module is used to obtain a reconstructed deep network model according to steps 1 to 3 of the malicious sample detection method based on deep adversarial generation in Example 1. The reconstructed deep network model is downloaded by the user and then trained offline to obtain a malicious sample detection model.

[0099] Among them, the upload module is used to receive the malicious sample detection model uploaded by the user and the corresponding trained weight file, or to receive the trained weight file uploaded by the user and establish a corresponding relationship between the weight file and the malicious sample detection model.

[0100] like Figure 3 As shown in the figure, the operations of uploading the module to create a new malicious sample detection model are as follows:

[0101] Receive the malicious sample detection model uploaded by the user in the format of a py file. The file name of the py file should be consistent with the class name in the file content. This step is mandatory. The malicious sample detection model is stored in the py file in code form, such as the malicious sample detection model code detect***.py.

[0102] Receive the trained weight file uploaded by the user in the format of a pth file. This step is mandatory and can be selected as a single option. For example, the weight file is detect***.pth. After uploading, the system database automatically records the corresponding malicious sample detection model and the weight file corresponding to the malicious sample detection model.

[0103] The data set used for this training is determined according to the user's selection, or when the system does not have the data set required by the user, the data set used for this training uploaded by the user is received and the data set is associated with the weight file.

[0104] If there are py files and / or pth files with the same name, you will be prompted to rename them and upload them after renaming; otherwise, upload them directly.

[0105] In addition, you can also fill in the model description when uploading the malicious sample detection model to facilitate the subsequent identification and calling of the malicious sample detection model. Figure 4 As shown in the figure, when the same malicious sample detection model is trained to obtain multiple weight files, it is necessary to independently create the malicious sample detection model weights. The operation is as follows:

[0106] Determine the malicious sample detection model based on user selection.

[0107] Receive the trained weight file uploaded by the user in the format of a pth file, and associate the weight file with the malicious sample detection model determined by the user.

[0108] The data set used for this training is determined according to the user's selection, or when the system does not have the data set required by the user, the data set used for this training uploaded by the user is received and the data set is associated with the weight file.

[0109] If a pth file with the same name exists, you will be prompted to rename it and upload it after renaming; otherwise, upload it directly.

[0110] The detection module is used to establish and run a new malicious sample detection model evaluation task based on the defense model, malicious sample detection model, malicious sample detection model weight file, training data set, and attack model selected by the user. The ratio of defense models, malicious sample detection models, data sets, and attack models is 1:1:1:n, where n is the number of selected attack models.

[0111] like Figure 5 As shown in the figure, create a new malicious sample detection model evaluation task and perform the following operations:

[0112] Determine the defense model and the weight file of the defense model based on the user input (the defense model and the weight file of the defense model are collectively referred to as the defense method). This step is a single choice. First select the defense model and then select the defense method. If no defense method exists, add a new defense method; otherwise, proceed to the next step.

[0113] A list of malicious sample detection model names is displayed. The malicious sample detection models provided in the list are the same as the dataset used by the defense method selected in the previous step. The malicious sample detection model is determined based on the user input. If no malicious sample detection model exists, a new malicious sample detection model is created; otherwise, proceed to the next step.

[0114] Displays the existing weight list of the current malicious sample detection model, such as Figure 6As shown, the weight file of the selected malicious sample detection model is determined according to the user input. For example, the weight file is selected as detect_denoise_resnext101_kait_imagenet.pth. If the corresponding weight file does not exist, a new weight file of the malicious sample detection model is created; otherwise, the next step is executed.

[0115] Calls a dataset based on user input.

[0116] Determine the attack model based on user input, for example, the name is fgsm, pgd, Figure 7 As shown, multiple attack models can be selected at the same time, and one or more different parameter configurations can be set for each attack model (the attack model and the corresponding parameter configuration are collectively referred to as the attack method). For example, the parameter configuration files are fgsm_01.xml, fgsm_02.xml, pgd_01.xml, and pgd_02.xml, so as to realize the generation of hyperparameter combinations by different factors, where the attack model and parameter configuration are preset in the evaluation system.

[0117] Run the defense model, malicious sample detection model, and attack model on the data set to perform the evaluation task of the malicious sample detection model.

[0118] Call the malicious sample detection model evaluation indicators, and output the evaluation results after the evaluation task is completed. The evaluation system presets various malicious sample detection model evaluation indicator script packages for calling.

[0119] Among them, the result output module is used to output the evaluation results of the malicious sample detection model according to the evaluation indicators after the evaluation task of the malicious sample detection model is completed, and complete the full-domain evaluation of the malicious sample detection model. The evaluation report can also be generated and displayed based on the evaluation results.

[0120] like Figure 8 As shown in the figure, it is the evaluation result generated after a malicious sample detection model evaluation task is completed. The record evaluation name is allDetect, the evaluation type is malicious sample evaluation, the creator is admin1, the creation time is 2024-11-29 17:28:50, the comprehensive score is 0, and the ranking is 1.

[0121] The attack model names used are: detect_upgd_imagenet, detect_tpgd_imagenet, detect_tifgsm_imagenet, detect_sinifgsm_imagenet, detect_rfgsm_imagenet, detect_pifgsm_plusplus_imagenet, detect_pifgsm_imagenet , detect_pgdl2_imagenet, detect_pgd_imagenet, detect_ntifgsm_imagenet, detect_mifgsm_imagenet, detect_fgsm_imagenet, detect_ffgsm_imagenet, detect_eotpgd_imagenet and detect_difgsm_imagenet.

[0122] The parameter configuration names of the attack models used are: detect_upgd imagenet_01.xml, detect_tpgd_imagenet_01.xml, detect_tifgsm_imagenet_01.xml, detect_sinifgsm_imagenet_01.xml, detect_rfgsm_imagenet_01.xml, detect_pifgsm_plusplus_imagenet_01.xml, detect_pifgsm_imagenet_01.xml, detect_pgdl2_imagenet_01.xml, detect_pgd_imagenet_01.xml, detect_ntifgsm_imagenet_01.xml, detect_mifgsm_imagenet_01.xml, detect_fgsm_imagenet_01.xml, detect_ffgsm_imagenet_01.xml , detect_eotpgd_imagenet_01.xml and detect_difgsm_imagenet_01.xml.

[0123] The name of the defense model used is: detect_defense_resnet101_imagenet.

[0124] The weight name of the defense model is: detect_defense_resnet101_imagenet.pth.

[0125] The name of the malicious sample detection model used is: detect_resnet101_imagenet.

[0126] The weight name of the malicious sample detection model is: detect_denoise_resnext101_kait_imagenet.pth.

[0127] The corresponding dataset of the malicious sample detection model is: Imagenet.

[0128] The evaluation indicator used is ACC (Accuracy).

[0129] The evaluation results are as follows:

[0130] Under the attack model detect_upgd_imagenet and parameter configuration detect_upgd imagenet_01.xml, the value of ACC is 1.0000.

[0131] Under the attack model detect_tpgd_imagenet and parameter configuration detect_tpgd_imagenet_01.xml, the value of ACC is 1.0000.

[0132] Under the attack model detect_tifgsm_imagenet and parameter configuration detect_tifgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0133] Under the attack model detect_sinifgsm_imagenet and parameter configuration detect_sinifgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0134] Under the attack model detect_rfgsm_imagenet and parameter configuration detect_rfgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0135] Under the attack model detect_pifgsm_plusplus_imagenet and parameter configuration detect_pifgsm_plusplus_imagenet_01.xml, the value of ACC is 1.0000.

[0136] Under the attack model detect_pifgsm_imagenet and parameter configuration detect_pifgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0137] Under the attack model detect_pgdl2_imagenet and parameter configuration detect_pgdl2_imagenet_01.xml, the value of ACC is 1.0000.

[0138] Under the attack model detect_pgd_imagenet and parameter configuration detect_pgd_imagenet_01.xml, the value of ACC is 1.0000.

[0139] Under the attack model detect_ntifgsm_imagenet and parameter configuration detect_ntifgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0140] Under the attack model detect_mifgsm_imagenet and parameter configuration detect_mifgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0141] Under the attack model detect_fgsm_imagenet and parameter configuration detect_fgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0142] Under the attack model detect_ffgsm_imagenet and parameter configuration detect_ffgsm_imagenet_01.xml, the ACC value is 0.9999.

[0143] Under the attack model detect_eotpgd_imagenet and parameter configuration detect_eotpgd_imagenet_01.xml, the value of ACC is 1.0000.

[0144] Under the attack model detect_difgsm_imagenet and parameter configuration detect_difgsm_imagenet_01.xml, the value of ACC is 1.0000.

[0145] And an evaluation report in a specified format can be generated according to the above evaluation results. This embodiment provides an evaluation report such as Fig. 9The present invention uses hyperparameter combinations to generate malicious attack samples under different attack intensities, realizes robustness evaluation result values ​​of malicious sample detection models under different hyperparameters and forms evaluation curves, and generates evaluation reports at the same time, providing a basic evaluation system for robustness verification of subsequent models.

[0146] In the evaluation system of the present invention, users can select the tasks of the system according to their own needs based on the models and weights and other files uploaded historically, and the system can complete the whole process of malicious sample evaluation through this process. Different from the previous malicious sample detection method for a single dimension, we can evaluate in a multi-dimensional and multi-combination manner (multi-attack combination, multi-indicator combination, multi-group parameter combination, multi-data set combination, etc.) through the method selected by the system user. By establishing the evaluation task, the system runs the evaluation task in the background according to the task allocation result, obtains the evaluation result after the operation, and can export the evaluation report according to the demand.

[0147] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0148] The above-mentioned embodiments only express several implementation modes of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention. It should be pointed out that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the attached claims.

Claims

1. A malicious sample detection method based on deep adversarial generation, characterized in that: The malicious sample detection method based on deep adversarial generation includes: Using graph analysis technology, the deep network model is decomposed to obtain multiple sub-graphs. Define adversarial spectrum perturbation in the spectrum domain, and add it to each sub-spectrum respectively. Take the perturbability of the sub-spectrum frequency as an indicator to measure the output change of each sub-spectrum before and after adding the adversarial spectrum perturbation. Assigning a first weighted value to a sub-graph whose output change is greater than a threshold, and assigning a second weighted value to a sub-graph whose output change is less than or equal to the threshold, to obtain a reconstructed deep network model, wherein the first weighted value and the second weighted value act on the weights of the corresponding sub-graphs respectively, and the first weighted value is less than the second weighted value; The reconstructed deep network model is trained offline to obtain a malicious sample detection model, which is used to perform malicious sample detection.

2. The method for detecting malicious samples based on deep adversarial generation according to claim 1, characterized in that: The graph analysis technology is used to decompose the deep network model to obtain multiple sub-graphs, including: Determine the number of input layers, hidden layers, and output layers of the deep network model, and determine the number of nodes in the input layer, hidden layer, and output layer, as well as the connection relationship between each node; Based on the nodes and the connection relationships between the nodes, graph theory tools are used to construct the graph of the deep network model. The nodes and connection relationships in the deep network model are mapped to the vertices and edges in the graph. Construct the Laplacian matrix of the graph; The Laplace matrix is ​​spectrally decomposed to obtain multiple sub-graphs, each of which contains an eigenvector and an eigenvalue.

3. The method for detecting malicious samples based on deep adversarial generation according to claim 1, characterized in that: The graph representation of the deep network model is as follows: In the formula, x is the input sample of the deep network model, x adv is the adversarial sample, ε is the amplitude or strength of the perturbation term added to the input sample, sign is the sign function, is the gradient operator, f(x,y,G) is the loss function, which represents the difference between the prediction result of the deep network model under the input sample x, the given label y and the graph G and the given label.

4. The method for detecting malicious samples based on deep adversarial generation according to claim 2, characterized in that: The step of defining the adversarial spectral perturbation in the spectral domain and adding the adversarial spectral perturbation to each sub-spectrum respectively includes: Where H is the output of the deep network model, f b is the function transformation, λ i represents the feature vector of the i-th sub-graph, Δ is the adversarial spectrum perturbation added to the feature vector of the i-th sub-graph, and h(λ i ) is the amplitude perturbation of the ith sub-spectrum, is the phase perturbation of the ith sub-spectrum.

5. The method for detecting malicious samples based on deep adversarial generation according to claim 4 is characterized in that: The output changes of each sub-spectrum before and after adding the adversarial spectrum disturbance include: In the formula, is the output change, H(f b (λ i )+Δ) is the output of the sub-graph after adding the adversarial spectrum perturbation, H(f b (λ i )) is the output of the sub-graph before adding the adversarial spectrum perturbation.

6. An evaluation system, characterized in that: The evaluation system comprises: A model building module, used to obtain a reconstructed deep network model according to the malicious sample detection method based on deep adversarial generation according to any one of claims 1-5, wherein the reconstructed deep network model is downloaded by the user and then trained offline to obtain a malicious sample detection model; An upload module, used to receive a malicious sample detection model uploaded by a user and a corresponding trained weight file, or to receive a trained weight file uploaded by a user and establish a corresponding relationship between the weight file and the malicious sample detection model; The detection module is used to establish and run the evaluation task of a new malicious sample detection model based on the defense model, malicious sample detection model, weight file of the malicious sample detection model, training data set, and attack model selected by the user; The result output module is used to output the evaluation results of the malicious sample detection model according to the evaluation indicators after the evaluation task of the malicious sample detection model is completed.

7. The evaluation system according to claim 6, characterized in that: The upload module receives the malicious sample detection model uploaded by the user and the corresponding trained weight file, and performs the following operations: Receive the malicious sample detection model uploaded by the user in the form of a py file. The file name of the py file should be consistent with the class name in the file content. Receive the trained weight file uploaded by the user in the format of pth file; Determine the data set used for this training according to the user's selection, or receive the data set used for this training uploaded by the user, and associate the data set with the weight file; If there is a py file or pth file with the same name, you will be prompted to rename it and upload it after renaming; Otherwise, the upload is completed directly.

8. The evaluation system according to claim 6, characterized in that: The upload module receives the trained weight file uploaded by the user, establishes a corresponding relationship between the weight file and the malicious sample detection model, and performs the following operations: Determine the malicious sample detection model based on user selection; Receive the trained weight file uploaded by the user in the format of a pth file, and associate the weight file with the malicious sample detection model determined by the user; Determine the data set used for this training according to the user's selection, or receive the data set used for this training uploaded by the user, and associate the data set with the weight file; If there is a file with the same name as the pth file, it will prompt you to rename it and complete the upload after renaming; Otherwise, the upload is completed directly.

9. The evaluation system according to claim 6, characterized in that: The detection module establishes a new malicious sample detection model evaluation task based on the defense model, malicious sample detection model, malicious sample detection model weight file, training data set and attack model selected by the user, and performs the following operations: Determine the defense model and the weight file of the defense model according to the user input. If the defense model does not exist, add a new defense model and the weight file of the defense model; Otherwise, proceed to the next step; Determine a malicious sample detection model based on user input, and create a new malicious sample detection model if one does not exist; Otherwise, proceed to the next step; Determine the weight file of the selected malicious sample detection model according to the user input, and if the corresponding weight file does not exist, create a new weight file of the malicious sample detection model; Otherwise, proceed to the next step; Calling a dataset based on user input; Determine the attack model and parameter configuration of the attack model according to user input; Run defense models, malicious sample detection models, and attack models on the dataset; Call the malicious sample detection model evaluation indicators and output the evaluation results after the evaluation task is completed.

10. The evaluation system according to claim 6, characterized in that: It also includes generating and displaying an evaluation report based on the evaluation results.

Citation Information

Patent Citations

  • Image retrieval methods, devices, equipment, and storage media to resist malicious sample attacks.

    CN112860932B