Robust hardware Trojan horse detection model generation method based on adversarial training

By generating a robust hardware Trojan detection model through adversarial training, the problem of deep learning models being vulnerable to adversarial attacks is solved, thereby improving the accuracy and security of hardware Trojan detection.

CN120974488APending Publication Date: 2025-11-18HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511021084.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing deep learning models are vulnerable to adversarial attacks in hardware Trojan detection, leading to decreased detection accuracy and difficulty in effectively identifying new types of hardware Trojans.

Method used

A robust hardware Trojan detection model is generated through adversarial training, including acquiring an infection dataset, flattening the dataset, generating a control-data flow graph, extracting statement paths, segmenting labeled samples, performing adversarial attacks, and training to generate a robust hardware Trojan detection model.

Benefits of technology

It improves the accuracy of hardware Trojan detection, effectively resists adversarial sample attacks, and enables fine-grained localization and early elimination of security threats from hardware Trojans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974488A_ABST
    Figure CN120974488A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of integrated circuit hardware security, in particular to a robust hardware Trojan horse detection model generation method based on adversarial training, and the method comprises the steps: obtaining an integrated circuit design infected by a hardware Trojan horse; flattening each integrated circuit design, converting the integrated circuit design into a control-data flow diagram, and extracting a statement path from an input signal to an output signal; carrying out segmentation and label labeling on each extracted statement path, generating an original sample set to train the initial network, and further generating an original hardware Trojan horse detection model; on the basis of the signal significance score and an original hardware Trojan horse detection model, carrying out adversarial attack on all original hardware Trojan horse samples in the original sample set, and generating a hardware Trojan horse adversarial sample set; and confrontation training is performed on the original hardware Trojan horse detection model based on the two sample sets to generate the robust hardware Trojan horse detection model, so that confrontation sample attack in the hardware Trojan horse detection process can be effectively resisted, and the detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of integrated circuit hardware security, and in particular to a robust hardware Trojan detection model generation method based on adversarial training. BACKGROUND

[0002] Integrated circuits (ICs) are widely used in various fields, making them an indispensable part of modern society. With the development of the times, people's demand for efficient, secure and intelligent electronic devices is rising, and the progress of IC technology has led to the continuous emergence of new system architectures and circuit designs. Whether in consumer electronics, medical devices or security systems, the performance and security of integrated circuits directly affect the normal operation of these industries.

[0003] Hardware Trojans are malicious design modifications implanted by attackers in integrated circuit designs. Their design forms are flexible and varied, with complex and hidden triggering mechanisms, which makes traditional integrated circuit testing methods face serious challenges in identifying hardware Trojan security threats. In recent years, hardware Trojan detection techniques based on machine learning have received widespread attention, especially in the pre-silicon gate level and register transfer level (RTL) design phase. These techniques extract circuit structures and behaviors of IC designs as features, transforming hardware Trojan detection into a binary classification problem. However, most of these methods rely on expert knowledge for feature extraction, have poor generalization ability, and are difficult to detect new types of hardware Trojans.

[0004] The rise of deep learning (DL) technology has prompted researchers to use deep neural networks (DNNs), graph neural networks (GNNs) and other models to automatically extract and detect hardware Trojan features, significantly improving detection efficiency and accuracy. However, deep learning models themselves are also vulnerable to adversarial attacks. Numerous studies have shown that adding minor perturbations to the original circuit to generate adversarial samples can significantly reduce the accuracy of deep learning models for hardware Trojan detection. Therefore, there is an urgent need for a robust hardware Trojan detection method to improve the adversarial robustness of deep learning models. SUMMARY

[0005] To solve the above technical problems, the present application proposes a robust hardware Trojan detection model generation method based on adversarial training, aiming to train a hardware Trojan detection model with adversarial robustness, which can effectively resist adversarial sample attacks during hardware Trojan detection, thereby improving the accuracy of hardware Trojan detection.

[0006] In order to achieve the above object, the application provides a robust hardware Trojan detection model generation method based on adversarial training, which comprises the following steps: obtaining a large number of integrated circuit designs infected by hardware Trojan with different functions to form an infection data set; performing flattening processing on each integrated circuit design in the infection data set, and converting the integrated circuit design after the flattening processing into a control-data flow graph; extracting a statement path between an input signal and an output signal in each control-data flow graph; segmenting and labeling each extracted statement path to generate original hardware Trojan samples and original normal samples, taking the original hardware Trojan samples as positive classes and the original normal samples as negative classes to form an original sample set, and training an initial network based on the original sample set to generate an original hardware Trojan detection model; performing adversarial attack on all original hardware Trojan samples in the original sample set based on a signal significance score and the original hardware Trojan detection model to generate a hardware Trojan adversarial sample set; and performing adversarial training on the original hardware Trojan detection model based on the original sample set and the hardware Trojan adversarial sample set to generate a robust hardware Trojan detection model.

[0007] In order to achieve the above object, the application further provides an electronic device, which comprises at least one processor and a memory in communication connection with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the robust hardware Trojan detection model generation method based on adversarial training.

[0008] In order to achieve the above object, the application further provides a computer readable storage medium storing a computer program, and the computer program is executable by a processor to implement the robust hardware Trojan detection model generation method based on adversarial training.

[0009] The application provides a robust hardware Trojan detection model generation method based on adversarial training. First, a large number of integrated circuit designs infected by hardware Trojans with different functions are obtained to form an infection dataset, which serves as the basis for training a robust hardware Trojan detection model. In order to better express features, the application needs to perform flattening processing on each integrated circuit design in the infection dataset, convert the flattened integrated circuit design into a control-data flow graph, extract the statement path between the input signal and the output signal in each control-data flow graph, and generate original hardware Trojan samples and original normal samples after segmentation and label annotation. The original hardware Trojan samples are used as positive classes, and the original normal samples are used as negative classes to form an original sample set. This processing fully expresses the semantic features and structural features in the integrated circuit design, and effectively improves the effect of regular training and adversarial training. For the original hardware Trojan detection model obtained by regular training, the application also needs to perform adversarial attacks on all original hardware Trojan samples in the original sample set based on the signal saliency score and the original hardware Trojan detection model to generate a hardware Trojan adversarial sample set. Then, the original hardware Trojan detection model is adversarially trained based on the original sample set and the hardware Trojan adversarial sample set, and a robust hardware Trojan detection model is generated. The finally trained robust hardware Trojan detection model can effectively resist adversarial sample attacks in the hardware Trojan detection process, thereby improving the hardware Trojan detection accuracy and realizing fine-grained positioning of hardware Trojans to eliminate security threats and hidden dangers as soon as possible.

[0010] Optionally, each hardware Trojan infected integrated circuit design in the infection dataset is input in the form of RTL code, and the flattening processing of each integrated circuit design in the infection dataset includes: converting the RTL code of each module in the integrated circuit design into an abstract syntax tree by using a Pyverilog analysis tool, wherein the nodes in the abstract syntax tree represent the basic units of the integrated circuit design, including arithmetic operation units, logic operation units, operands, etc.; starting from the top module in the abstract syntax tree, copying a corresponding abstract syntax tree for each module instantiation statement, pointing the input signal of the instantiation statement to the input signal of the abstract syntax tree, and pointing the output signal of the abstract syntax tree to the output signal of the instantiation statement; after traversing all the instantiation statements of the modules, the abstract syntax tree of the top module is taken as the flattened circuit design.

[0011] Optionally, the flattened integrated circuit design is converted into a control-dataflow graph, including: performing control flow and data flow analysis on the top-level module corresponding to the flattened integrated circuit design, adding the control dependencies and data dependencies between statements to the abstract syntax tree, and obtaining the control-dataflow graph corresponding to the flattened integrated circuit design. The graph data structure of the control-dataflow graph is saved in the form of objects defined internally by Python.

[0012] Optionally, the step of extracting the statement path from the input signal to the output signal in each control-dataflow graph includes: using a depth-first search algorithm to extract the statement path from the input signal to the output signal in each control-dataflow graph, where each element in the statement path corresponds to a node in the abstract syntax tree.

[0013] Optionally, the extracted statement paths are segmented and labeled to generate original hardware Trojan samples and original normal samples. The original hardware Trojan samples are designated as positive classes, and the original normal samples are designated as negative classes to form an original sample set. This includes: segmenting the extracted statement paths using a sliding window algorithm to obtain segmented statement paths; labeling the segmented statement paths according to whether they contain hardware Trojan logic, assigning category label 1 to segmented statement paths containing hardware Trojan logic (denoted as original hardware Trojan samples), and assigning category label 0 to segmented statement paths not containing hardware Trojan logic (denoted as original normal samples); and treating all original hardware Trojan samples as positive classes and all original normal samples as negative classes to form an original sample set.

[0014] Optionally, the initial network is a code2vec model, whose model structure consists of three parts: a word embedding layer, an attention layer, and an output layer. Based on the original sample set, the initial network is trained to generate the original hardware Trojan detection model, including: The original samples from the original sample set are input into the initial network; the original samples include original hardware Trojan samples as positive classes and original normal samples as negative classes. The initial network performs a serialization operation on the original input samples. The serialized sentence path is expressed by the formula: , This indicates the number of words in the statement path. Indicates the first One word; The initial word embedding layer of the network will The words in the text are converted into continuous vector representations, thereby preserving and enabling the initial network to learn the structural and semantic features of integrated circuit design. The sentence path after word embedding is represented as word vectors. , ; The initial network's attention layer quantifies the importance of each word within its sentence path and represents the sentence path through aggregation operations, initializing the global attention vector with random values. Based on the following formula Calculate the attention weight for each word : ; in, and They represent The first in The element and the first elements, global attention vector Attention weights for each word Updated during the initial network training; Based on attention weight calculation The weighted average is used to generate a single path vector. , to represent the entire statement path, ; The output layer of the initial network utilizes The output of the hardware trojan detection is as follows: ; in, and These are the weight matrix and bias vector of the output layer, respectively. This represents the softmax function; based on Using the labels on the original samples and the CBfocal loss function, the initial network is trained to generate the original hardware Trojan detection model.

[0015] Optionally, based on the signal saliency score and the original hardware Trojan detection model, adversarial attacks are performed on all original hardware Trojan samples in the original sample set to generate a hardware Trojan adversarial sample set, including: Extract the signal names from the segmented Trojan statement paths to form a list of Trojan signal names. And calculate the global significance score for each signal name; Trojan signal global significance score The calculation process is as follows: ; in, Indicates a label, This indicates the presence of a Trojan signal. The number of Trojan horse statement paths, This indicates the path to the original Trojan statement. This indicates the path to the modified statement. express The first in One word, Indicates to The revised result; Based on the global significance score, Sort the data in descending order to obtain the list of signal names to be attacked. ; Using the leave-one-out method, the dataset is divided into a training set and a test set. The test set consists of an integrated circuit design. Signal names that do not appear in the integrated circuit design code of the training set are extracted to form a replacement signal name list. ; from Select each signal name in turn ,from Select signal name right Perform a global replacement and calculate the attack success rate. , This represents the proportion of paths predicted as category label 0 by the original hardware Trojan detection model among all Trojan statement paths after an attack. The calculation formula is: ; in, , Indicates an indicator function; Choose the attack with the highest success rate. ,get The final replacement signal , ; All original hardware Trojan samples in the original sample set are globally replaced to generate hardware Trojan adversarial samples, thus forming a hardware Trojan adversarial sample set.

[0016] Optionally, based on the original sample set and the hardware Trojan adversarial sample set, the original hardware Trojan detection model is adversarially trained to generate a robust hardware Trojan detection model. This includes: merging the original sample set and the hardware Trojan adversarial sample set into an adversarial training set; inputting the training samples from the adversarial training set into the original hardware Trojan detection model; performing adversarial training pre-tests on the original hardware Trojan detection model; and then combining the F1 score maximization strategy to finally generate a robust hardware Trojan detection model. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies of this application will be briefly introduced below. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. The drawings described herein are only used to explain this application and are not intended to limit this application.

[0018] Figure 1 This is a flowchart of a robust hardware Trojan detection model generation method based on adversarial training provided in one embodiment of this application; Figure 2 This is a schematic diagram of the process of flattening and control-data flow graph conversion of an integrated circuit design infected by a hardware Trojan, provided in one embodiment of this application. Figure 3 This is a schematic diagram provided in one embodiment of the present application, illustrating the process of extracting statement paths from a control-data flow graph and segmenting and labeling the extracted statement paths; Figure 4 This is a schematic diagram illustrating the statement path extraction process provided in one embodiment of this application; Figure 5 This is a schematic diagram of the structure of the code2vec model provided in one embodiment of this application; Figure 6 This is a schematic diagram of an anti-attack process provided in one embodiment of this application; Figure 7 This is a schematic diagram of an adversarial training process provided in one embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in another embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the various embodiments of this application will be described in detail below with reference to the accompanying drawings. Those skilled in the art will understand that many technical details have been presented in the embodiments of this application to facilitate better understanding. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments. The division of the following embodiments is for ease of description and should not constitute any limitation on the specific implementation of this application. The following embodiments can be combined with and referenced by each other without contradiction.

[0020] One embodiment of this application proposes a robust hardware Trojan detection model generation method based on adversarial training, applied to electronic devices, wherein the electronic device can be a terminal or a server. This embodiment and the following embodiments will use a server as an example for description. The implementation details of the robust hardware Trojan detection model generation method based on adversarial training proposed in this embodiment are described in detail below. The following implementation details are provided for ease of understanding only and are not necessary for implementing this solution.

[0021] The specific process of the robust hardware Trojan detection model generation method based on adversarial training proposed in this embodiment can be described as follows: Figure 1 As shown, it includes: Step 11: Obtain a large number of integrated circuit designs with different functions that have been infected by hardware Trojans, and form an infection dataset.

[0022] In practice, the server first needs to acquire a large number of integrated circuit designs infected with hardware trojans, each with different functions, to form an infection dataset as the basis for model training. Understandably, the server needs to acquire as many integrated circuit designs infected with different hardware trojans as possible to enhance the model's general applicability.

[0023] In one example, the various integrated circuit designs infected by the hardware trojan in the infected dataset are input in the form of RTL code, such as Verilog HDL code.

[0024] In one example, an integrated circuit design in RTL code form can be as follows: Figure 2 As shown.

[0025] Step 12: Flatten each integrated circuit design in the infection dataset and convert the flattened integrated circuit design into a control-data flow graph.

[0026] In the specific implementation, after the server obtains the infection dataset, it needs to flatten each integrated circuit design in the infection dataset to obtain an abstract syntax tree, and then convert the flattened integrated circuit design into a control-data flow graph, thereby converting the integrated circuit design in RTL code form into a control-data flow graph form that is easy to view and process.

[0027] In one example, the flattening process can be as follows: Figure 2 As shown, the server uses the Pyverilog parsing tool to convert the RTL code of each module in the integrated circuit design into an abstract syntax tree. The nodes in the abstract syntax tree represent the basic units of the integrated circuit design, including arithmetic operation units, logic operation units, operands, etc.

[0028] Subsequently, the server traverses from the top-level module in the abstract syntax tree of the RTL code. For each module's instantiation statement, it copies its corresponding abstract syntax tree representation, points the input signal of the instantiation statement to the input signal of the module's abstract syntax tree, and points the output signal of the module's abstract syntax tree to the output signal of the instantiation statement.

[0029] After traversing all the instantiation statements of all modules, the abstract syntax tree representation of the top-level module is used as the flattened circuit design.

[0030] In one example, the control-data flow graph transformation process can be as follows: Figure 2 As shown, the server performs control flow and data flow analysis on the top-level module corresponding to the flattened integrated circuit design, adding the control and data dependencies between statements to the abstract syntax tree to obtain the control-data flow graph corresponding to the flattened integrated circuit design. The graph data structure of the control-data flow graph is stored as objects defined internally by Python. Control flow represents the control dependencies between statements in the RTL code, such as conditions and branches, while data flow represents the data flow between signals in the RTL code, such as assignment operations.

[0031] Step 13: Extract the statement paths from input signals to output signals in each control-data flow diagram.

[0032] In the specific implementation, after the server flattens all the integrated circuit designs in the infection dataset and converts them into control-data flow graphs, it needs to extract the statement paths from the input signals to the output signals in each control-data flow graph.

[0033] In one example, such as Figure 3 As shown, the server uses a depth-first search algorithm to extract the statement paths from input signals to output signals in each control-dataflow graph. Each element in the statement path corresponds to a node in the abstract syntax tree.

[0034] In one example, the specific details of the statement path extraction process can be as follows: Figure 4 As shown, the distinction between data flow and control flow is clearly visible.

[0035] Step 14: Segment and label the extracted statement paths to generate original hardware Trojan samples and original normal samples. Use the original hardware Trojan samples as positive classes and the original normal samples as negative classes to form an original sample set. Based on the original sample set, train the initial network to generate the original hardware Trojan detection model.

[0036] In the specific implementation, after the server completes the extraction of all statement paths, it can segment and label each extracted statement path to generate original hardware Trojan samples and original normal samples. The original hardware Trojan samples are used as positive classes and the original normal samples are used as negative classes to form the original sample set. Then, based on the original sample set, the initial network is trained in a regular manner to generate the original hardware Trojan detection model.

[0037] In one example, the process of splitting a statement path can be as follows: Figure 3 As shown, the server uses a sliding window algorithm (with a preset window size and step size, e.g., a window size of 6 and a step size of 4) to segment the extracted statement paths, resulting in segmented statement paths. Next, the segmented statement paths are labeled based on whether they contain hardware Trojan logic. Segmented statement paths containing hardware Trojan logic are assigned category label 1, which is the original hardware Trojan sample; segmented statement paths not containing hardware Trojan logic are assigned category label 0, which is the original normal sample. Finally, all original hardware Trojan samples are treated as positive classes, and all original normal samples are treated as negative classes, forming the original sample set as the basis for regular training.

[0038] In one example, the initial network is a code2vec model, whose model structure can be as follows: Figure 5 As shown, it consists of three parts: word embedding layer, attention layer and output layer.

[0039] In one example, when the server performs regular training on the initial network based on the original sample set, it needs to input the original samples from the original sample set into the initial network. The original sample set includes original hardware Trojan samples as positive classes and original normal samples as negative classes.

[0040] The initial network performs a serialization operation on the original input samples. The serialized sentence path can be expressed by the formula: , , This indicates the number of words in the statement path. Indicates the first One word.

[0041] The initial word embedding layer of the network will The words in the text are converted into continuous vector representations, thereby preserving and enabling the initial network to learn the structural and semantic features of integrated circuit design. The sentence path after word embedding is represented as word vectors. , .

[0042] The initial attention layer of the network quantifies the importance of each word in its sentence path and represents the sentence path through aggregation operations. The global attention vector is then initialized with random values. , Based on the following formula Calculate the attention weight for each word : ; in, and They represent The first in The element and the first 1 element, calculate attention weight The process is actually calculation and The normalized inner product between them. Specific global attention vectors. Attention weights for each word Updated during the initial network training.

[0043] Next, the attention layer calculates based on attention weights. The weighted average is used to generate a single path vector. , to represent the entire statement path, .

[0044] The output layer of the initial network utilizes The output of the hardware trojan detection is as follows: ; in, and These are the weight matrix and bias vector of the output layer, respectively. This represents the softmax function.

[0045] The server can set a Trojan detection threshold of 0.5, which means that if the predicted probability is greater than the threshold, the path is detected as a Trojan path; otherwise, it is a normal path.

[0046] Finally, the server is based on Using the labels on the original samples and the CBfocal loss function, the initial network is trained to generate the original hardware Trojan detection model.

[0047] Step 15: Based on the signal saliency score and the original hardware Trojan detection model, perform adversarial attacks on all original hardware Trojan samples in the original sample set to generate a hardware Trojan adversarial sample set.

[0048] In practice, after completing the regular training, the server also needs to perform adversarial attacks on all the original hardware Trojan samples in the original sample set based on the signal saliency score and the original hardware Trojan detection model, to generate a hardware Trojan adversarial sample set as the basis for adversarial training.

[0049] In one example, the adversarial attack process can be as follows: Figure 6 As shown. The server extracts the signal names from the segmented Trojan statement paths and assembles a list of Trojan signal names. , And calculate the global significance score for each signal name. Trojan signals global significance score The calculation process is as follows: ; in, Indicates a label, This indicates the presence of a Trojan signal. The number of Trojan horse statement paths, This indicates the path to the original Trojan statement. This indicates the path to the modified statement. express The first in One word, Indicates to The modified results; next, the server will sort the results according to the global significance scores. Sort the data in descending order to obtain the list of signal names to be attacked. .

[0050] Next, the server divides the dataset into a training set and a test set using the leave-one-out method. The test set consists of an integrated circuit design. Signal names that do not appear in the integrated circuit design code of the training set are extracted from the integrated circuit design code of the test set to form a list of replacement signal names. .

[0051] Subsequently, the server from Select each signal name in turn ,from Select signal name right Perform a global replacement and calculate the attack success rate. , This represents the proportion of paths predicted as category label 0 by the original hardware trojan detection model among all trojan statement paths after an attack.

[0052] The calculation formula is: ; in, ; Indicates the indicator function. Specific global attention vector. Attention weights for each word Update during network model training.

[0053] Finally, choose the attack with the highest success rate. ,get The final replacement signal The original hardware Trojan samples in the original sample set are globally replaced to generate hardware Trojan adversarial samples, thus forming a hardware Trojan adversarial sample set. .

[0054] In one example, the server can set an attack success rate threshold (e.g., 80%), first from the list of signal names to be attacked. Sort by global saliency score, select a single signal as the attack target, and the attack is successful if the success rate reaches a threshold. If the attack on a single signal does not reach the specified threshold, then select from the list of signals to be attacked. The algorithm selects two signal names based on their importance scores and attacks them simultaneously, calculating the attack success rate. An attack is considered successful if the success rate reaches a specified threshold. If the success rate does not reach the threshold, the replacement combination with the highest success rate is selected as the final attack choice, and all original hardware Trojan samples in the original sample set are globally replaced to generate adversarial hardware Trojan samples.

[0055] Step 16: Based on the original sample set and the hardware Trojan adversarial sample set, perform adversarial training on the original hardware Trojan detection model to generate a robust hardware Trojan detection model.

[0056] In practice, after obtaining the hardware Trojan adversarial sample set, the server can perform adversarial training on the original hardware Trojan detection model based on the original sample set and the hardware Trojan adversarial sample set to generate a robust hardware Trojan detection model, thereby achieving high-precision hardware Trojan detection.

[0057] In one example, the adversarial training process can be as follows: Figure 7 As shown, the server first combines the original sample set and the set of adversarial samples of hardware Trojans into an adversarial training set. Then, the training samples in the adversarial training set are input into the original hardware Trojan detection model to perform adversarial training pre-test on the original hardware Trojan detection model. Finally, combined with the F1 score maximization strategy, a robust hardware Trojan detection model is generated.

[0058] This embodiment proposes a robust hardware Trojan detection model generation method based on adversarial training. First, it acquires a large number of integrated circuit designs with different functions infected by hardware Trojans, forming an infection dataset as the foundation for training a robust hardware Trojan detection model. To better represent features, this application needs to flatten each integrated circuit design in the infection dataset and convert the flattened integrated circuit designs into control-data flow graphs. Then, it extracts the statement paths from input signals to output signals in each control-data flow graph. After segmentation and labeling, it generates original hardware Trojan samples and original normal samples. The original hardware Trojan samples are designated as the positive class, and the original normal samples as the negative class, forming the original sample set. This processing fully expresses the semantic and structural features of the integrated circuit designs, significantly improving the effectiveness of both conventional and adversarial training. For the original hardware Trojan detection model obtained through conventional training, this application further requires adversarial attacks on all original hardware Trojan samples in the original sample set based on the signal saliency score and the original hardware Trojan detection model. This generates a hardware Trojan adversarial sample set. Based on the original sample set and the hardware Trojan adversarial sample set, the original hardware Trojan detection model is then adversarially trained to generate a robust hardware Trojan detection model. The finally trained robust hardware Trojan detection model can effectively resist adversarial attacks during the hardware Trojan detection process, thereby improving the accuracy of hardware Trojan detection and achieving fine-grained localization of hardware Trojans to eliminate security threats and hidden dangers as early as possible.

[0059] The steps described above are for clarity only. In implementation, they can be combined into one step, or some steps can be broken down into multiple steps, as long as they involve the same logical relationship, they are all within the scope of protection of this application. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the scope of protection of this application.

[0060] In one embodiment, taking the Trust-Hub hardware Trojan test set as an example, the effectiveness and superiority of the robust hardware Trojan detection model generation method based on adversarial training proposed in this application are illustrated.

[0061] 1. Constructing the dataset. The dataset includes 11 hardware trojan circuits from Trust-Hub. These hardware trojans are conditionally triggered and are implanted into AES, RS232, and PIC16F84 reference circuits described in Verilog HDL code.

[0062] 2. Flatten the integrated circuit design and convert the register-transfer level (RTL) code into an abstract syntax tree (AST). Specifically, use Pyverilog tools for syntax parsing, reflecting the structural information of the RTL code in the branching structure of the AST, thereby quickly identifying the dependencies between various syntactic elements (including but not limited to modules, signals, expressions, and control statements) in the code. After conversion, the control dependencies between RTL code statements can be represented by the branching structure in the AST, thus reflecting the execution order of statements in the RTL code.

[0063] 3. Convert each circuit under test into a control-dataflow graph based on the abstract syntax tree (API). Considering that the API itself contains control flow information, constructing the control-dataflow graph only requires analyzing the data flow relationships between statements and connecting statements with signal declarations and calls within the API. After this step, the nodes in the control-dataflow graph represent statements in the circuit design, while data flow and control flow are represented by the edges connecting the nodes.

[0064] 4. A depth-first search algorithm is used for the control-dataflow graph of each integrated circuit design. By analyzing the control flow and data flow, the statement paths between the main input and output signals of the hardware design are extracted. The extracted paths completely represent the signal propagation paths inside the circuit, effectively preserving the circuit structure and semantic information in the hardware design.

[0065] 5. Using a sliding window algorithm to segment the extracted statement paths helps improve the granularity of hardware Trojan detection. This process involves two parameters: window size and stride. In practice, the specific window size is 6 and the stride is 4, where the window size represents the number of statements contained in the window. Figure 4 This paper demonstrates a complete process for extracting statement paths. The Trojan triggering module of the AES-T500 is converted into a control-data flow graph, and segmented statement paths are obtained based on the sliding window algorithm. Since hardware Trojan detection is a supervised learning process, before training the hardware Trojan detection model, class labels need to be assigned to the segmented statement paths. Specifically, paths containing hardware Trojan logic are assigned class label 1, denoted as the original hardware Trojan sample, while paths not containing hardware Trojan logic are assigned class label 0, denoted as the original normal sample. For example, in... Figure 4 In part (d), the two paths are the normal path and the path infected by the Trojan, respectively, and are therefore assigned category label 0 and category label 1. By performing statement path segmentation and labeling on the Trust-Hub hardware Trojan test set circuit, an original sample set is generated. The original samples in the original sample set include original hardware Trojan samples and original normal samples.

[0066] 6. Based on the original sample set, perform routine training of the model to generate the original hardware Trojan detection model. Specifically, the model used is the attention-based code2vec model, the structure of which is as follows: Figure 5 As shown, the code2vec model mainly consists of three parts: a word embedding layer, an attention layer, and an output layer. The performance of the code2vec model was evaluated using leave-one-out cross-validation, with a total of 11 rounds of experiments. In each round, the statement path extracted by one circuit was used as the test data, while the statement paths of the remaining circuits were used for model training. The evaluation metrics for the hardware Trojan detection code2vec model include: True Positive Rate (TPR), True Negative Rate (TNR), Accuracy, F1 score, and Area Under the Curve (AUC). These metrics fully reflect the performance of the original hardware Trojan detection model from different dimensions.

[0067] Based on the pre-trained original code2vec model, hardware Trojan detection was performed on the test data. Table 1 lists the detection results of the pre-trained original code2vec model on the original samples. Specifically, the average TPR, TNR, Accuracy, F1 score, and AUC values ​​were 96.51%, 99.31%, 99.13%, 90.06%, and 97.91%, respectively. This indicates that the pre-trained model can correctly classify and accurately detect hardware Trojans in the original samples.

[0068] Table 1. Hardware Trojan Detection Results of the Original Code2Vec Model on the Original Samples

[0069] 7. In the adversarial attack experiments, a leave-one-out method was used to attack each hardware Trojan circuit, generating adversarial examples of hardware Trojans. The attack effect was verified by calculating TPR, Accuracy, F1 score, AUC, and Attack Success Rate (ASR). Notably, a lower TPR indicates poor model robustness, while a higher ASR means a higher attack success rate, indicating adversarial vulnerability. Since the attacker's goal is to attack the Trojan examples, the TNR value remains essentially unchanged. The attack method proposed in this application achieved an attack success rate of 99.53% on the original code2vec model. Furthermore, the adversarial examples successfully reduced the TPR and AUC of the original code2vec model, with the TPR, F1 score, and AUC of the original code2vec model on the adversarial examples decreasing to 0.47%, 0.85%, and 50.19%, respectively. The results show that the original code2vec model is not robust to adversarial examples.

[0070] Table 2: Adversarial attack results of the original code2vec model

[0071] 8. In the adversarial training phase, firstly, adversarial attacks are performed on all original hardware Trojan samples in the original sample set to generate a hardware Trojan adversarial sample set. Then, the original sample set and the hardware Trojan adversarial sample set are merged to generate an adversarial training set. Next, based on this adversarial training set, a code2vec model with the same architecture is retrained using the leave-one-out method. A robust hardware Trojan detection model is generated based on the F1 score maximization strategy.

[0072] Tables 3 and 4 show the hardware Trojan detection results of the robust code2vec model obtained after adversarial training on adversarial and original samples, respectively. According to the experimental results, the robust code2vec model achieves average TPR, TNR, Accuracy, F1 score, and AUC of 96.02%, 99.04%, 99.03%, 91.14%, and 96.20% on adversarial samples, respectively, and average TPR, TNR, Accuracy, F1 score, and AUC of 94.79%, 99.59%, 99.50%, 89.60%, and 96.14% on original samples, respectively. The robust code2vec model obtained after adversarial training can achieve high-precision detection of hardware Trojans in both original and adversarial samples.

[0073] Table 3: Results of robust code2vec model in detecting hardware Trojans in adversarial examples.

[0074] Table 4: Hardware Trojan Detection Results of Robust Code2Vec Model on Original Samples

[0075] Another embodiment of this application provides an electronic device, such as Figure 8 As shown, it includes: at least one processor 21; and a memory 22 communicatively connected to the at least one processor 21; wherein the memory 22 stores instructions executable by the at least one processor 21, the instructions being executed by the at least one processor 21 to enable the at least one processor 21 to execute a robust hardware Trojan detection model generation method based on adversarial training as described in the above method embodiment.

[0076] The memory and processor are connected via a bus, which includes any number of interconnecting buses and bridges, connecting various circuits of one or more processors and the memory. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.

[0077] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.

[0078] Another embodiment of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, can implement a robust hardware Trojan detection model generation method based on adversarial training as described in the above method embodiments.

[0079] That is, those skilled in the art will understand that all or part of the steps in the above method embodiments can be implemented by a program instructing related hardware. The program is stored in a storage medium and includes several instructions to cause a device (such as a microcontroller, chip, etc.) or processor to execute all or part of the steps of the method in the method embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.

[0080] Those skilled in the art will understand that the above embodiments are specific implementations of this application, and in practical applications, various changes can be made in form and detail without departing from the spirit and scope of this application. For those skilled in the art, several improvements and modifications can be made without departing from the principles of this application, and these improvements and modifications are also considered to be within the scope of protection of this application.

Claims

1. A robust hardware Trojan detection model generation method based on adversarial training, applicable to integrated circuit design, characterized in that, The method includes: Obtain a large number of integrated circuit designs with different functions that have been infected by hardware Trojans, and form an infection dataset; The integrated circuit designs in the infection dataset are flattened, and the flattened integrated circuit designs are converted into control-data flow graphs. Extract the statement paths from input signals to output signals in each control-dataflow graph; The extracted statement paths are segmented and labeled to generate original hardware Trojan samples and original normal samples. The original hardware Trojan samples are used as positive classes and the original normal samples are used as negative classes to form an original sample set. Based on the original sample set, the initial network is trained to generate an original hardware Trojan detection model. Based on the signal saliency score and the original hardware Trojan detection model, adversarial attacks are performed on all original hardware Trojan samples in the original sample set to generate a hardware Trojan adversarial sample set. Based on the original sample set and the hardware Trojan adversarial sample set, the original hardware Trojan detection model is adversarially trained to generate a robust hardware Trojan detection model.

2. The method for generating a robust hardware Trojan detection model based on adversarial training according to claim 1, characterized in that, Each integrated circuit design infected by the hardware trojan in the infection dataset is input in the form of RTL code. The integrated circuit designs in the infection dataset are then flattened, including: The Pyverilog parsing tool is used to convert the RTL code of each module in the integrated circuit design into an abstract syntax tree. The nodes in the abstract syntax tree represent the basic units of the integrated circuit design, including arithmetic operation units, logic operation units, and operands. Starting from the top-level module in the abstract syntax tree, for each module's instantiation statement, copy its corresponding abstract syntax tree representation, point the input signal of the instantiation statement to the input signal of the abstract syntax tree, and point the output signal of the abstract syntax tree to the output signal of the instantiation statement. After traversing all the instantiation statements of all modules, the abstract syntax tree of the top-level module is used as the flattened circuit design.

3. The method for generating a robust hardware Trojan detection model based on adversarial training according to claim 1, characterized in that, Converting the flattened integrated circuit design into a control-dataflow graph includes: Control flow and data flow analysis are performed on the top-level module corresponding to the flattened integrated circuit design. The control and data dependencies between statements are added to the abstract syntax tree to obtain the control-data flow graph corresponding to the flattened integrated circuit design. The graph data structure of the control-data flow graph is saved in the form of objects defined internally by Python.

4. The method for generating a robust hardware Trojan detection model based on adversarial training according to claim 1, characterized in that, The extraction of statement paths from input signals to output signals in each control-data flow graph includes: A depth-first search algorithm is used to extract the statement paths from input signals to output signals in each control-dataflow graph. Each element in the statement path corresponds to a node in the abstract syntax tree.

5. The method for generating a robust hardware Trojan detection model based on adversarial training according to claim 1, characterized in that, The extracted statement paths are segmented and labeled to generate original hardware Trojan samples and original normal samples. The original hardware Trojan samples are designated as positive classes, and the original normal samples are designated as negative classes, forming the original sample set, which includes: The extracted statement path is segmented using a sliding window algorithm to obtain the segmented statement path; Based on whether the segmented statement path contains hardware Trojan logic, the segmented statement path is labeled. The segmented statement path containing hardware Trojan logic is assigned category label 1 and is recorded as the original hardware Trojan sample. The segmented statement path that does not contain hardware Trojan logic is assigned category label 0 and is recorded as the original normal sample. All original hardware Trojan samples are treated as positive classes, and all original normal samples are treated as negative classes, forming the original sample set.

6. The method for generating a robust hardware Trojan detection model based on adversarial training according to claim 1, characterized in that, The initial network is a code2vec model, whose structure consists of three parts: a word embedding layer, an attention layer, and an output layer. Based on the original sample set, the initial network is trained to generate the original hardware Trojan detection model, including: The original samples from the original sample set are input into the initial network; the original samples include original hardware Trojan samples as positive classes and original normal samples as negative classes. The initial network performs a serialization operation on the original input samples. The serialized sentence path is expressed by the formula: , This indicates the number of words in the statement path. Indicates the first One word; The initial word embedding layer of the network will The words in the text are converted into continuous vector representations, thereby preserving and enabling the initial network to learn the structural and semantic features of integrated circuit design. The sentence path after word embedding is represented as word vectors. , ; The initial network's attention layer quantifies the importance of each word within its sentence path and represents the sentence path through aggregation operations, initializing the global attention vector with random values. Based on the following formula Calculate the attention weight for each word : ; in, and They represent The first in The element and the first elements, global attention vector Attention weights for each word Updated during the initial network training; Based on attention weight calculation The weighted average is used to generate a single path vector. , to represent the entire statement path, ; The output layer of the initial network utilizes The output of the hardware trojan detection is as follows: ; in, and These are the weight matrix and bias vector of the output layer, respectively. This represents the softmax function; based on Using the labels on the original samples and the CBfocal loss function, the initial network is trained to generate the original hardware Trojan detection model.

7. The method for generating a robust hardware Trojan detection model based on adversarial training according to claim 1, characterized in that, Based on signal saliency scores and the original hardware Trojan detection model, adversarial attacks are performed on all original hardware Trojan samples in the original sample set to generate a hardware Trojan adversarial sample set, including: Extract the signal names from the segmented Trojan statement paths to form a list of Trojan signal names. And calculate the global significance score for each signal name; Trojan signal global significance score The calculation process is as follows: ; in, Indicates a label, This indicates the presence of a Trojan signal. The number of Trojan horse statement paths, This indicates the path to the original Trojan statement. This indicates the path to the modified statement. express The first in One word, Indicates to The revised result; Based on the global significance score, Sort the data in descending order to obtain the list of signal names to be attacked. ; Using the leave-one-out method, the dataset is divided into a training set and a test set. The test set consists of an integrated circuit design. Signal names that do not appear in the integrated circuit design code of the training set are extracted to form a replacement signal name list. ; from Select each signal name in turn ,from Select signal name right Perform a global replacement and calculate the attack success rate. , This represents the proportion of paths predicted as category label 0 by the original hardware Trojan detection model among all Trojan statement paths after an attack. The calculation formula is: ; in, , Indicates an indicator function; Choose the attack with the highest success rate. ,get The final replacement signal , ; All original hardware Trojan samples in the original sample set are globally replaced to generate hardware Trojan adversarial samples, thus forming a hardware Trojan adversarial sample set.

8. The method for generating a robust hardware Trojan detection model based on adversarial training according to claim 1, characterized in that, Based on the original sample set and the hardware Trojan adversarial sample set, the original hardware Trojan detection model is adversarially trained to generate a robust hardware Trojan detection model, including: Combine the original sample set and the hardware Trojan adversarial sample set into an adversarial training set; The training samples from the adversarial training set are input into the original hardware Trojan detection model to perform adversarial training pre-tests on the original hardware Trojan detection model. Then, by combining the F1 score maximization strategy, a robust hardware Trojan detection model is finally generated.

9. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute a robust hardware Trojan detection model generation method based on adversarial training as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it can implement a robust hardware Trojan detection model generation method based on adversarial training as described in any one of claims 1 to 8.