Deep learning model attack-oriented detection method based on GPU HPCs
By using GPU hardware performance counter event data and naive Bayesian method, a deep learning model attack detection model is constructed, which solves the problem of difficult to identify deep learning model poisoning and backdoor attacks in the existing technology, and achieves efficient and transparent attack detection.
Patent Information
- Application Number
- CN202510065287.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
It is difficult for the prior art to effectively identify and defend against data poisoning and backdoor attacks in deep learning models, especially in the absence of internal information and computing resources of deep learning models.
By collecting GPU hardware performance counter event data during deep learning model training, the Naive Bayes method is used to build a detection model to identify the feature patterns of the GPU performance counter, thereby detecting whether the deep learning model is poisoned or backdoor attacked.
This method can efficiently and transparently identify poisoning and backdoor attacks of deep learning models, and is highly versatile and robust, and does not rely on the internal structure or weight information of the deep learning model.
Smart Images

Figure CN119989343A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer security and artificial intelligence technology, and in particular to a detection method for deep learning model attacks based on GPU HPCs, which is used to identify threats such as data poisoning attacks and backdoor attacks. Background Art
[0002] With the development of artificial intelligence and big data technology, deep learning models have been widely used in computer vision, natural language processing, network security and other fields. However, the security issues of deep learning models have gradually become a research hotspot, especially the increasingly serious threats of data poisoning attacks and backdoor attacks. These attacks take advantage of the training phase of deep learning models and induce deep learning models to output wrong decisions in certain situations by injecting specific malicious data. Since deep learning models are essentially complex multi-layer neural networks, their decision-making process is difficult to explain and is regarded as a "black box". This lack of explainability makes the attack behavior more concealed and also limits the defense capabilities of existing detection technologies.
[0003] At present, attack detection methods for deep learning models are mainly divided into two categories: white-box detection and black-box detection. White-box detection requires access to the internal information of the deep learning model, including weight parameters, activation functions, and gradient values. For example, through the analysis of the weight distribution or gradient analysis of the deep learning model, the significant impact of abnormal samples on the training of the deep learning model is found. Although the white-box detection method can deeply analyze the internal structure of the deep learning model and provide high detection accuracy, its implementation depends on the complete access rights and computing resources of the deep learning model, which is limited in many application scenarios. In contrast, the black-box detection method pays more attention to the external behavior of the deep learning model and infers potential security issues through the relationship between input data and output results. Typical black-box detection methods include analyzing the impact of input samples on the output of the deep learning model through activation clustering, or using entropy detection to identify the distribution characteristics of the triggering input. However, these methods rely on the statistical characteristics of the sample input distribution and are easily interfered by adversarial samples. In addition, traditional black-box detection is mainly based on software features and lacks consideration of the underlying hardware behavior, resulting in limited robustness of the detection method.
[0004] Hardware Performance Counters (HPCs) are a set of registers embedded in hardware that record the usage of underlying hardware resources, such as the number of instruction executions, cache hit rates, and data transfer volume. Unlike data collected at the software layer, the values of hardware performance counters directly reflect the underlying microarchitecture behavior and are difficult to tamper with. In recent years, studies have shown that the values of hardware performance counters can reflect the microarchitecture characteristics of software behavior, and the behavior of deep learning models can significantly affect the values of hardware performance counters, especially when attacked by poisoning or backdoor attacks. The characteristics of GPU performance counters show a pattern that is significantly different from normal training. Therefore, based on the data timing characteristics of GPU hardware performance counters, combined with machine learning methods, an efficient and transparent security detection method can be provided for deep learning models. Summary of the invention
[0005] In order to solve the above problems, the present invention proposes a detection method for deep learning model attacks based on GPU HPCs. This method collects GPU hardware performance counter event data during the training process of the deep learning model and uses machine learning algorithms to build a corresponding detection model, thereby effectively identifying poisoning and backdoor attacks on the deep learning model.
[0006] To achieve the above object, the present invention adopts the following technical solution:
[0007] The detection method for deep learning model attacks based on GPU HPCs includes the following steps:
[0008] 1) Train the deep learning model on the GPU platform. Each time, input a normal data set, a poisoning attack data set, or a backdoor attack data set into the initial deep learning model to train and collect the corresponding GPU data set;
[0009] 2) Using the naive Bayes method as a machine learning classifier, the data set generated in step 1) is used for training to obtain a detection model for deep learning model attacks; the GPU hardware performance counter data generated during the deep learning model training is input into the attack detection model for detection to obtain the detection results.
[0010] The technical solution is further optimized, and the step 1) specifically collects the GPU hardware performance counter data of the entire training process as the initial data; then processes the initial data, filters the features through the information gain method, and generates a data set for training the attack detection model based on the processed and filtered results.
[0011] This technical solution is further optimized, and the specific steps of step 1) are as follows:
[0012] 1.1) Install the corresponding environment on the GPU platform and deploy the initial deep learning model;
[0013] 1.2) Input a normal data set or a poisoning attack data set or a backdoor attack data set to train the initial deep learning model. Each time training is performed, the GPU hardware event counts generated during the entire training process are collected to obtain GPU computing-related, memory and cache-related, and data transmission-related GPU hardware event counts. Call the nsysprofile command through the Nsight tool, use the "ga10x-gfxt" performance indicator set as the standard, set the frequency of 10kHz to collect hardware performance counter data, and add parameters such as the path for storing the output data file and the path for the deep learning model training code. After the collection is completed, a file in the nsys-rep format is generated; repeat steps 1.1) and 1.2) to generate a set of nsys-rep files as the initial data set;
[0014] 1.3) Convert the nsys-rep format file into a sqlite database file. For each sqlite file, extract the GPU_metrics table. The metricId and value fields in the table refer to the number of times the GPU hardware performance counter event numbered metricId occurred in a certain period of time. Sum the value according to metricId to obtain the number of occurrences of each hardware performance counter feature event in the entire deep learning model training process.
[0015] 1.4) Label the data. The data collected by the deep learning model trained with normal data is labeled as 0, and the data collected by the deep learning model trained with poisoning attack data sets or backdoor attack data sets is labeled as 1.
[0016] 1.5) Calculate the corresponding information gain score for the collected hardware performance counter features, and select the top 12 features; use the 12 GPU hardware performance counter counts collected during a training process as a set of GPU hardware performance counter feature data; and obtain a data set for training the attack detection model as a whole.
[0017] The technical solution is further optimized, and the feature selection method of step 1.5) is as follows:
[0018] Calculate the information gain score IG(T,X) corresponding to each hardware performance counter feature using the following formula:
[0019] IG(T,X)=H(T)-H(T,X)
[0020] Where T is the data label 0 or 1, X is the feature variable, i.e., the hardware performance counter feature, H(T) is the entropy of the target variable T, and H(T,X) is the conditional entropy of the target variable T under the given feature X. The calculation formulas of H(T) and H(T,X) are as follows:
[0021]
[0022] Among them, P(t i ) is the category t in the target variable T i The frequency, t 0 Trained on a normal dataset, t 1 Trained for poisoning attack datasets or backdoor attack datasets;
[0023]
[0024] Among them, P(x j ) is the feature X with value x j The sampling frequency of G(T|x j ) is given by X = x j The entropy of the target variable T is as follows:
[0025] H(T|x j )=-P(t 0 |x j )log 2 P(t 0 |x j )-P(t 1 |x j )log 2 P(t 1 |x j )
[0026] Among them, P(t 0 |x j ) and P(t 1 |x j ) is in the feature X = x j When , label T is the frequency of 0 and 1.
[0027] The technical solution is further optimized, and the step 2) is specifically as follows:
[0028] 2.1) Design a machine learning model based on the naive Bayes method as a classifier, and train the machine learning model with the data obtained in step 1) to obtain a detection model for deep learning model attacks;
[0029] 2.2) According to step 1), collect the GPU hardware performance counter data generated by the deep learning model to be tested in a training, and perform the same processing and screening; input the processed data into the detection model to obtain the output result, and judge whether the deep learning model is attacked by security; if it is 1, it means that the deep learning model is attacked, otherwise it is judged as not attacked.
[0030] The technical solution is further optimized, and the step 2.1) of the machine learning model training process is specifically as follows:
[0031] The classifier training data set is D, D = {(X 1 ,y 1 ),(X 2 ,y 2 ),…,(X n ,y n )}, each input data is a vector X=[x 1 ,x 2 ,…,x 12 ], x i represents the i-th hardware performance counter feature, with label y∈{0,1}, where 0 represents that the corresponding X is collected from the deep learning model trained with unattacked data, and 1 represents that the corresponding X is collected from the deep learning model trained with attacked data;
[0032] The first step is to calculate the prior probability P(y) of each label. The calculation formula is:
[0033]
[0034] Where n is the total number of samples in the dataset D, n 0 is the number of samples from unattacked 1 is the number of samples from the attacked ones;
[0035] The second step is to calculate the conditional probability, for each feature X i (i=1,2,…,12) and each class label y∈{0,1}, the conditional probability P(X i The calculation formula of |y) is:
[0036]
[0037] where μ i,y is feature x i The mean of the samples labeled y, is feature x i The variance in the sample labeled y is calculated as follows:
[0038]
[0039] Where n y is the total number of samples with label y, x ij represents the i-th eigenvalue of sample j.
[0040] Different from the prior art, the above technical solution has the following beneficial effects:
[0041] 1) It uses GPU hardware performance counters to collect data directly from the bottom layer, avoiding the prior knowledge of traditional software detection methods
[0042] 2) It does not rely on the internal structure or weight information of the deep learning model and is highly transparent.
[0043] 3) It can show consistent results on different deep learning models and GPU hardware platforms, and has strong versatility and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 The flowchart of the detection method for deep learning model attacks based on GPU HPCs. DETAILED DESCRIPTION
[0045] In order to explain the technical content, structural features, achieved objectives and effects of the technical solution in detail, the following is a detailed description in conjunction with specific embodiments and accompanying drawings.
[0046] like Figure 1 FIG. 1 is a flowchart of a method for detecting anomalies of a virtualization platform based on hardware performance counters. The method includes the following steps performed in sequence:
[0047] 1) Train the deep learning model on the GPU platform. Each time, input the normal data set, poisoning attack data set, or backdoor attack data set into the initial deep learning model to train and collect the corresponding GPU data set. Specifically, use the Nsight tool to collect the GPU hardware performance counter data throughout the training process as the initial data. Then process the initial data, filter the features using the information gain method, and generate a data set for training the attack detection model based on the processed and filtered results.
[0048] 1.1) Install the corresponding environment on the GPU platform and deploy the initial deep learning model.
[0049] 1.2) Input a normal data set or a poisoning attack data set or a backdoor attack data set to train the initial deep learning model. Each time the training is performed, the GPU hardware event counts generated during the entire training process are collected once to obtain GPU computing-related, memory and cache-related, and data transmission-related GPU hardware event counts. Call the nsysprofile command through the Nsight tool, set the "ga10x-gfxt" performance indicator set as the standard, and set the frequency of 10kHz to collect hardware performance counter data. The specific content of the collection command is: nsys profile-trace = none --gpu-metrics-set = ga10x-gfxt --gpu-metrics-device = 0 --gpu-metrics-frequency = 10000 --backtrace = none --cpuctxsw = none -o [output data file storage path] -wtrue [deep learning model training code path]. After the collection is completed, the nsys-rep format file is generated. Repeat steps 1.1) and 1.2) to generate a set of nsys-rep files as the initial data set.
[0050] 1.3) Call the nsys export command to convert the nsys-rep format file into a sqlite database file. The specific content of the command is: nsys export-t sqlite-separate-string=true-o[output sqlite file storage path][input nsys-rep file storage path]. For each sqlite file, extract the GPU_metrics table. The metricId and value fields in the table refer to the number of times the GPU hardware performance counter event numbered metricId occurred in a certain period of time. Sum the value according to metricId. The specific sql statement is: "select metricId, SUM(value) as total_value from GPU_METRICS where metricIdbetween 0and 117group by metricId", and obtain the number of occurrences of each hardware performance counter feature event in the entire process of deep learning model training.
[0051] 1.4) Label the data. The data collected by the deep learning model trained with normal data is labeled as 0, and the data collected by the deep learning model trained with the poisoning attack dataset or the backdoor attack dataset is labeled as 1.
[0052] 1.5) Calculate the corresponding information gain scores for the 118 collected hardware performance counter features, and select the top 12 features. The specific screening method is as follows:
[0053] Calculate the information gain score IG(T,X) corresponding to each hardware performance counter feature using the following formula:
[0054] IG(T,X)=H(T)-H(T,X),
[0055] Where T is the data label 0 or 1, X is the feature variable, i.e., the hardware performance counter feature, H(T) is the entropy of the target variable T, and H(T,X) is the conditional entropy of the target variable T given the feature X. The calculation formulas for H(T) and H(T,X) are as follows:
[0056]
[0057] Among them, P(t i ) is the category t in the target variable T i The frequency, t 0 Trained on a normal dataset, t 1 Trained for poisoning attack datasets or backdoor attack datasets.
[0058]
[0059] Among them, P(x j ) is the feature X with value x j The sampling frequency, H(T|x j ) is given by X = x j The entropy of the target variable T is as follows:
[0060] H(T|x j )=-P(t 0 |x j )log 2 P(t 0 |x j )-P(t 1 |x j )log 2 P(t 1 |x j ),
[0061] Among them, P(t 0 |x j ) and P(t 1 |x j ) is in the feature X = x j When , label T is the frequency of 0 and 1.
[0062] The 12 selected hardware performance counter features are: PCle Write Requests to BAR0 (number of peripheral component high-speed interconnect write requests for base address register 0), PCle Read Requests to BAR0 (number of peripheral component high-speed interconnect read requests for base address register 0), PCle Write Bandwidth (peripheral component high-speed interconnect write bandwidth), PCleRead Bandwidth (peripheral component high-speed interconnect read bandwidth), Compute Warps (number of compute warps), AsyncCopy Engine Active (asynchronous copy engine active state), SM FMA Heavy Pipe Throughput (streaming multiprocessor floating point multiplication and addition heavy load pipeline throughput), SM IssueActive (streaming multiprocessor instruction issuance activity), L1Shared+Attribute Data-Stage Throughput (first-level cache shared+attribute data stage throughput), L1Surface Data-Stage Throughput (first-level cache surface data stage throughput), L1 HitRate (first-level cache hit rate), L2 Bandwidth fromLl (bandwidth from first-level cache to second-level cache). The 12 GPU hardware performance counter counts collected during a training process are used as a set of GPU hardware performance counter feature data. The overall data set used to train the attack detection model is saved in the form of csv.
[0063] 2) Using the naive Bayes method as a machine learning classifier, the dataset generated in step 1) is used for training to obtain a detection model for deep learning model attacks. The GPU hardware performance counter data generated during the deep learning model training is input into the attack detection model for detection to obtain the detection results.
[0064] 2.1) Design a machine learning model based on the naive Bayes method as a classifier, and use the data obtained in step 1) to train the machine learning model to obtain a detection model for deep learning model attacks. The specific process of machine learning model training is as follows:
[0065] The classifier training data set is D, D = {(X 1 ,y 1 ),(X 2 ,y 2 ),…,(X n ,y n )}, each input data is a vector X=[x 1 ,x 2 ,…,x 12 ], xi Represents the i-th hardware performance counter feature, with label y∈{0,1}, where 0 represents that the corresponding X is collected from the deep learning model trained with unattacked data, and 1 represents that the corresponding X is collected from the deep learning model trained with attacked data.
[0066] The first step is to calculate the prior probability P(y) of each label. The calculation formula is:
[0067]
[0068] Where n is the total number of samples in the dataset D, n 0 is the number of samples from unattacked 1 is the number of samples from the attacked.
[0069] The second step is to calculate the conditional probability, for each feature X i (i=1,2,…,12) and each class label y∈{0,1}, the conditional probability P(X i The calculation formula of |y) is:
[0070]
[0071] where μ i,y is feature x i The mean of the samples labeled y, is feature x i The variance in the sample labeled y is calculated as follows:
[0072]
[0073] Where n y is the total number of samples with label y, x ij represents the i-th eigenvalue of sample j.
[0074] 2.2) According to step 1), collect the GPU hardware performance counter data generated by the deep learning model to be tested in a training, and perform the same processing and screening. Input the processed data into the detection model, obtain the output result, and judge whether the deep learning model is under security attack. If it is 1, it means that the deep learning model is under attack, otherwise it is judged as not under attack.
[0075] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "include..." or "comprise..." do not exclude the existence of other elements in the process, method, article or terminal device including the elements. In addition, in this article, "greater than", "less than", "exceed" and the like are understood to exclude the number itself; "above", "below", "within" and the like are understood to include the number itself.
[0076] Although the above embodiments have been described, once those skilled in the art know the basic creative concepts, they can make additional changes and modifications to these embodiments. Therefore, the above description is only an embodiment of the present invention and does not limit the patent protection scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the specification and drawings of the present invention, or directly or indirectly used in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A detection method for deep learning model attacks based on GPU HPCs, characterized in that: The following steps are involved: 1) Train the deep learning model on the GPU platform. Each time, input a normal data set, a poisoning attack data set, or a backdoor attack data set into the initial deep learning model to train and collect the corresponding GPU data set; 2) Using the naive Bayes method as a machine learning classifier, the data set generated in step 1) is used for training to obtain a detection model for deep learning model attacks; the GPU hardware performance counter data generated during the deep learning model training is input into the attack detection model for detection to obtain the detection results.
2. The GPU HPCs-based deep learning model attack detection method according to claim 1, characterized in that: The step 1) specifically collects GPU hardware performance counter data throughout the training process as initial data; then processes the initial data, filters features using the information gain method, and generates a data set for training the attack detection model based on the processed and filtered results.
3. The GPU HPCs-based deep learning model attack detection method according to claim 2, characterized in that: The specific steps of step 1) are as follows: 1.1) Install the corresponding environment on the GPU platform and deploy the initial deep learning model; 1.2) Input a normal data set or a poisoning attack data set or a backdoor attack data set to train the initial deep learning model. Each time you train, collect the GPU hardware event counts generated during the entire training process to obtain GPU computing-related, memory and cache-related, and data transmission-related GPU hardware event counts. Use the Nsight tool to call the nsysprofile command, use the "ga10x-gfxt" performance indicator set as the standard, set the frequency of 10kHz to collect hardware performance counter data, and add parameters such as the path for storing the output data file and the path for the deep learning model training code. After the collection is completed, generate a file in the nsys-rep format; Repeat steps 1.1) and 1.2) to generate a set of nsys-rep files as the initial data set; 1.3) Convert the nsys-rep format file into a sqlite database file. For each sqlite file, extract the GPU_metrics table. The metricId and value fields in the table refer to the number of times the GPU hardware performance counter event numbered metricId occurred in a certain period of time. Sum the value according to metricId to obtain the number of occurrences of each hardware performance counter feature event in the entire deep learning model training process. 1.4) Label the data. The data collected by the deep learning model trained with normal data is labeled as 0, and the data collected by the deep learning model trained with the poisoning attack dataset or the backdoor attack dataset is labeled as 1. 1.5) Calculate the corresponding information gain scores for the collected hardware performance counter features and select the top 12 features; The 12 GPU hardware performance counter counts collected during a training process are used as a set of GPU hardware performance counter feature data; As a whole, we obtain a dataset for training attack detection models.
4. The method for detecting deep learning model attacks based on GPU HPCs as claimed in claim 3, characterized in that: The feature selection method of step 1.5) is as follows: Calculate the information gain score IG(T,X) corresponding to each hardware performance counter feature using the following formula: IG(T,X)=H(T)-H(T,X) Where T is the data label 0 or 1, X is the feature variable, i.e., the hardware performance counter feature, H(T) is the entropy of the target variable T, and H(T,X) is the conditional entropy of the target variable T under the given feature X. The calculation formulas of H(T) and H(T,X) are as follows: Among them, P(t i ) is the category t in the target variable T i The frequency of t0 is trained with the normal data set, and t1 is trained with the poisoning attack data set or the backdoor attack data set; Among them, P(x j ) is the feature X with value x j The sampling frequency, H(T|x j ) is given by X = x j The entropy of the target variable T is as follows: H(T|x j )=-P(t0|x j )log2P(t0|x j )-P(t1|x j )log2P(t1|x j ) Among them, P(t0|x j ) and P(t1|x j ) is in the feature X = x j When , label T is the frequency of 0 and 1.
5. The GPU HPCs-based deep learning model attack detection method according to claim 1, characterized in that: The step 2) is as follows: 2.1) Design a machine learning model based on the naive Bayes method as a classifier, and train the machine learning model with the data obtained in step 1) to obtain a detection model for deep learning model attacks; 2.2) According to step 1), the GPU hardware performance counter data generated by the deep learning model to be tested in a training is collected and processed and screened in the same way; the processed data is input into the detection model to obtain the output result, and determine whether the deep learning model is subject to security attacks; If it is 1, it means that the deep learning model has been attacked, otherwise it is judged as not being attacked.
6. The GPU HPCs-based deep learning model attack detection method according to claim 5, characterized in that: The step 2.1) of the machine learning model training process is as follows: The classifier training data set is D, D = {(X1, y1), (X2, y2), ..., (X n ,y n )}, each input data is a vector X=[x1,x2,…,x 12 ], x i represents the i-th hardware performance counter feature, with label y∈{0,1}, where 0 represents that the corresponding X is collected from the deep learning model trained with unattacked data, and 1 represents that the corresponding X is collected from the deep learning model trained with attacked data; The first step is to calculate the prior probability P(y) of each label. The calculation formula is: Where n is the total number of samples in the dataset D, n0 is the number of samples from unattacked samples, and n1 is the number of samples from attacked samples; The second step is to calculate the conditional probability, for each feature X i (i=1,2,…,12) and each class label y∈{0,1}, the conditional probability P(X i The calculation formula of |y) is: where μ i,y is feature x i The mean of the samples labeled y, is feature x i The variance in the sample labeled y is calculated as follows: where n y is the total number of samples with label y, x ij represents the i-th eigenvalue of sample j.