Information-missing-oriented super-large-scale integrated circuit security design detection method
The graph variational autoencoder completes the incomplete circuit structure and uses the graph neural network model to extract features, solving the problem of difficult to accurately detect hardware Trojans under incomplete circuit structure in the prior art, and achieving more efficient hardware Trojan detection.
Patent Information
- Application Number
- CN202510158954.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-06-03
AI Technical Summary
Existing hardware Trojan detection technology is difficult to accurately detect potential hardware Trojans when facing incomplete circuit structures, especially when critical paths are damaged or features are missing.
GraphVAE is used to perform structural reasoning and completion of circuit structures with missing information, to construct completed circuit directed graph data, and extract deep-level features through graph neural network model (GNN) to realize the classification of hardware Trojans.
Through the completion of circuit structure and feature extraction, comprehensive information closer to the original design is provided, the accuracy and applicability of hardware Trojan detection is improved, and more targeted defense measures can be carried out in the subsequent process of integrated circuit design.
Smart Images

Figure CN120087302A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of hardware Trojan detection, and particularly relates to a security design detection method for very large scale integrated circuits (VLSI) facing information loss. Background Art
[0002] Integrated circuits are the core of cyber-physical systems, driving the development of cutting-edge fields such as intelligent manufacturing, autonomous driving, and virtual reality. As the scale of integrated circuits gradually expands to very large scale integrated circuits, the circuit design and development cycle has been greatly extended, and the design and production costs have continued to rise. At the same time, the urgency of product time-to-market has become increasingly prominent. To improve efficiency and shorten the development cycle, the design and manufacturing processes of integrated circuits increasingly rely on third-party intellectual property (IP) cores and electronic design automation (EDA) tools. However, the introduction of these third-party resource tools brings potential security risks. Attackers may implant hardware Trojans in third-party IP cores through the supply chain or use insecure design tools to carry out Trojan attacks, resulting in significant security hazards and damage.
[0003] To address the threats posed by hardware Trojans, many detection techniques have been proposed:
[0004] Among them, a patent for invention with publication number CN115204078A discloses a method and system for detecting hardware Trojans in integrated circuits. The method includes: obtaining the integrated circuit layout of a chip to be tested, performing reverse engineering on it, and generating a gate-level netlist; verifying the equivalence of the gate-level netlist generated by reverse engineering with the gate-level netlist of the original design of the chip to be tested; and finally, judging whether a hardware Trojan has been implanted in the chip according to the verification result.
[0005] Another patent for invention with publication number CN119249421A discloses a method for detecting hardware Trojans by fusing structural information and semantic information, which can solve the technical problem of poor hardware Trojan detection ability. The method includes: determining a first file among multiple chip design files; establishing an abstract syntax tree based on the first file; extracting node features and adjacency relationships according to the abstract syntax tree; obtaining structural information based on the node features and node adjacency relationships; obtaining code semantic features based on multiple code files; fusing the structural information and code semantic features based on a cross-attention mechanism to obtain a target feature vector; training a hardware Trojan detection model based on the target feature vector; and inputting the code to be tested into the model to obtain the Trojan detection result.
[0006] Another invention patent with the publication number CN118607018A discloses a hardware Trojan detection method and device based on a gated recurrent neural network. The method includes: obtaining the gate-level netlist circuit of the hardware IP core to be tested; screening out the nodes exceeding the controllability threshold from the gate-level netlist circuit as suspicious circuit nodes; taking the suspicious circuit nodes as root nodes, forming a suspicious circuit with multiple nodes having a connection relationship with the suspicious circuit nodes, and dividing it into multiple node paths, encoding each node in each node path according to the logic gate type as a feature to obtain a circuit sequence feature; inputting the circuit sequence feature into the gated recurrent neural network model, and the gated recurrent neural network model performs hardware Trojan detection on the hardware IP core to be tested and outputs a detection result.
[0007] Most of the existing hardware Trojan detection technologies are used in complete circuit netlists. However, in actual application scenarios, due to reasons such as intellectual property protection and the limitations of reverse engineering technology, the circuit netlists used for hardware Trojan detection are often incomplete. The missing circuit structures in the incomplete netlist will seriously affect the accuracy of Trojan detection. Moreover, the incomplete circuit structure will also prevent the acquisition of global information features in Trojan detection, such as the global positions in the primary inputs, primary outputs, and circuit logic gates, further weakening the accuracy and applicability of the detection method. Especially when the critical path is damaged or the features for clearly identifying Trojans are missing, the existing technologies cannot comprehensively and accurately detect potential hardware Trojans. Summary of the Invention
[0008] The purpose of the present invention is to provide a security design detection method for very large-scale integrated circuits facing information loss. This method can effectively infer and complete the incomplete circuit structure, and provide more comprehensive information closer to the original design based on the completed circuit structure, so as to achieve more effective hardware Trojan detection, which helps to take more targeted defense measures in the subsequent processes of integrated circuit design.
[0009] To achieve the above purpose, the technical solution of the present invention is: a security design detection method for very large-scale integrated circuits facing information loss, including:
[0010] Parsing the gate-level netlist design file of the very large-scale integrated circuit with information loss, creating a directed graph representation and performing feature matrix encoding on the logic gate nodes to construct circuit directed graph data;
[0011] Using the Graph Variational Autoencoder (GraphVAE) to perform structure inference on the directed graph data, and using the decoder to complete the incomplete circuit structure to construct the completed circuit directed graph data;
[0012] Calculate the logic gate type, fan-in and fan-out quantities, betweenness centrality, and neighborhood type distribution for the complemented circuit directed graph data, and combine and splice them into circuit gate features;
[0013] Construct a graph neural network model GNN, take the complemented circuit directed graph data and circuit gate features as inputs, further extract deep features through graph convolution operations, and use a multi-layer perceptron for inference training to achieve hardware Trojan classification.
[0014] In an embodiment of the present invention, the method includes the following steps:
[0015] Step A: Parse the obtained very large scale integrated circuit gate-level netlist design file with missing information, convert the circuit structure into a directed graph form GI=(V, E); in the directed graph, V represents the logic gates in the netlist, and E represents the connection relationships between the logic gates; convert the directed graph data into a directed graph adjacency matrix A; at the same time, perform One-Hot encoding on the types of logic gate nodes in the directed graph as the logic gate node feature matrix X; further, label each logic gate with a hardware Trojan label Y to construct circuit directed graph data including type features and label information;
[0016] Step B: Perform structure inference and complementation on the incomplete circuit structure in the circuit directed graph data; construct a graph variational autoencoder GraphVAE, take the directed graph adjacency matrix A and the logic gate node feature matrix X as the inputs of GraphVAE, perform structure inference training, calculate whether there are missing parts in the circuit structure in terms of function and topological relationship, and use the decoder to complement the missing structure to obtain the complemented adjacency matrix A recon and the logic gate node feature matrix
[0017] Step C: Perform feature selection on the complemented circuit directed graph data, and the extracted features include: logic gate type, logic gate fan-in quantity, logic gate fan-out quantity, betweenness centrality of the logic gate, and neighborhood type distribution of the logic gate; combine and splice the calculated features into the final circuit gate feature F;
[0018] Step D: Perform inference classification on the obtained circuit gate feature F; construct a graph neural network model GNN, take the complemented adjacency matrix A recon and the circuit gate feature F as the inputs of GNN, extract deeper features H through graph convolution operations, and input H into the multi-layer perceptron layer in GNN for inference training; calculate the probabilities that each circuit gate belongs to a normal circuit gate and a Trojan circuit gate; finally, achieve the classification of hardware Trojans.
[0019] In an embodiment of the present invention, Step A includes the following steps:
[0020] Step A1: Obtain a gate-level netlist design file with missing information for training, and parse its structure to collect all circuit gate types involved in the netlist.
[0021] Step A2: Extract a set of logic gates V = {v 0 , v 1 , v 2 ,..., v n} and a set of connections E = {(v i , v j ) | v i , v j ∈ V} from the gate-level netlist, where v i and v j represent the i-th and j-th logic gates extracted from the gate-level netlist respectively, and n represents the total number of logic gates in the gate-level netlist.
[0022] Step A3: Construct a directed graph with the set of logic gates V as vertices and the set of connections E between the logic gates as edges; further, convert the directed graph into an adjacency matrix A. If (v i , v j ) exists in the edge set, then the corresponding element α ij in the adjacency matrix is 1, otherwise it is 0.
[0023] Step A4: Encode the type of each logic gate node in the gate-level netlist using One-Hot encoding to form a feature matrix; the length of the One-Hot encoding is the number of circuit gate types collected in Step A1; for each logic gate, set the position corresponding to the One-Hot type to 1 and the remaining positions to 0; further, merge all the logic gate node feature matrices in the order of the nodes in the adjacency matrix A into a logic gate node feature matrix X.
[0024] Step A5: Parse the Trojan nodes in the netlist, label the Trojan logic gate nodes with Trojan labels and the normal logic gate nodes with normal labels to form the label information Y of the netlist. The initial incomplete directed graph data is composed of the adjacency matrix A, the logic gate node feature matrix X, and the label information Y.
[0025] In an embodiment of the present invention, in Step A4, 32-bit encoding is used as the feature matrix encoding.
[0026] In an embodiment of the present invention, Step B includes the following steps:
[0027] Step B1: Take the circuit directed graph data of Step A as input and perform a normalization operation on the adjacency matrix. The normalization operation is as follows:
[0028]
[0029] Among them, is represented as the adjacency matrix for adding self-loops, and I n is represented as the identity matrix; is the degree matrix calculated through the adjacency matrix ;
[0030] Step B2: Construct a graph convolutional layer to extract node features in the directed graph. The operation formula is as follows:
[0031]
[0032] Among them, H (l) is the node feature matrix of the l-th layer, and H (0) = X is the initial node feature matrix; W (l) represents the training weight matrix of the l-th layer; σ(·) represents the activation function in the graph convolution operation; through multi-layer graph convolution operations, the node feature matrix is updated layer by layer to capture high-order topological information and local structure features;
[0033] Step B3: Aggregate the node feature matrix H (l+1) obtained in Step B2 into a global vector h, representing the structural information of the circuit. The operation formula is as follows:
[0034] h = pooling(H (l+1) )
[0035] Among them, pooling(·) represents the pooling operation function;
[0036] Step B4: Perform a latent space mapping on the global vector. Map the global vector h to the latent space through a fully connected layer to generate distribution parameters. The operation formula is as follows:
[0037] μ = FC μ (h),
[0038] logσ = FC σ (h)
[0039] Among them, μ and logσ respectively represent the mean and log variance of the latent space; FC μ (·) and FC σ (·) respectively represent the fully connected operations for the mean and log variance of the latent space;
[0040] Step B5: Sample latent variables from the latent space distribution through a reparameterization operation. The operation formula is as follows:
[0041] z = μ + σ · ε
[0042] Among them, z represents the latent variable obtained by sampling; ε~N(0, I) is the standard normal distribution;
[0043] Step B6: Use the sampled latent variable to complete the node features and adjacency matrix through the decoder. The operation formula is as follows:
[0044]
[0045] A recon =σ(zz T )
[0046] Among them, represents the reconstructed node feature matrix; A recon represents the completed adjacency matrix; FC node (·) represents the fully connected operation for generating the node feature matrix; σ(·) represents the activation function; z T represents the transpose of the latent variable;
[0047] Step B7: Construct the loss function of the completion model. In the training and inference of incomplete circuit structure completion, the weighted sum of the reconstruction loss and the KL divergence loss is used as the total loss function L total ;
[0048] Step B8: Through the backpropagation mechanism, optimize the model parameters according to the value of the total loss function L total . Among them, the optimization algorithm uses the Adam optimizer; when the number of loss value iteration rounds reaches the preset value, terminate the training of the model;
[0049] Step B9: Input the circuit netlist with missing information to be measured into the trained inference completion model for completion, and output the completed directed graph; at the same time, create an initial netlist label in the completed directed graph to judge whether the node is a completed node; if it is a node generated by completion, the label value is set to 0, otherwise, it is set to 1.
[0050] In an embodiment of the present invention, in step B7, the loss function calculation includes the following content:
[0051] The reconstruction loss is used to measure the difference between the completed graph and the input graph, including the reconstruction errors of the node feature matrix and the adjacency matrix. The operation formula is as follows:
[0052]
[0053] Among them, represents the L2 norm; represents the Frobenius norm;
[0054] The KL divergence is used to constrain the distribution of the latent space to be close to the standard normal distribution. The operation formula is as follows:
[0055]
[0056] Among them, μ i and σ i respectively represent the mean and variance of the latent variables for each dimension i;
[0057] Combining the reconstruction loss and the KL divergence loss, the total loss function is:
[0058] L total = L recon + λL KL
[0059] where λ represents the weight of the loss function.
[0060] In an embodiment of the present invention, step C includes the following steps:
[0061] Step C1: Use the circuit directed graph data after completing step B for feature selection. Traverse each logic gate in the directed graph, collect the input ports and the number of output ports of each logic gate, which respectively correspond to the fan-in number FI of the logic gate and the fan-out number FO of the logic gate;
[0062] Step C2: Use the breadth-first search algorithm to calculate the number of shortest paths from each logic gate node in the directed graph to other logic gate nodes. Starting from the leaf nodes of the shortest path tree, backpropagate the dependency relationships between the nodes, and record the betweenness centrality of all logic gates; its operation formula is as follows:
[0063]
[0064] where σ ij represents the total number of shortest paths from logic gate node i to logic gate node j; σ ij (v) represents the number of paths passing through logic gate node v in the shortest path from i to j;
[0065] Step C3: Traverse the logic gate nodes in the directed graph, count and record the type distribution of the nodes in the neighborhood of each logic gate node; the type distribution characteristics of the neighborhood are represented by integer encoding, where the length of the encoding is the number of circuit gate types collected in step A; the initial value of each encoding bit is 0, and counting is performed according to the number of different types of nodes in the neighborhood, so as to generate the neighborhood type distribution feature vector ND of the complete logic gate;
[0066] Step C4: Obtain the logic gate node feature matrix obtained in step B Each node feature vector XN in; the node feature vector XN of each node, the fan-in number FI of the logic gate, the fan-out number FO of the logic gate, the betweenness centrality BC of the logic gate, and the neighborhood type distribution feature vector ND of the logic gate are concatenated, and finally merged into the feature matrix F for hardware Trojan classification.
[0067] In an embodiment of the present invention, in step C4, the encoding lengths of XN and ND are both 32 bits, FI, FO, and BC are all one bit, and the encoding length of the jointly formed node feature matrix is 67 bits.
[0068] In an embodiment of the present invention, the step D includes the following steps:
[0069] Step D1, construct a graph convolutional layer to extract the feature representation of each node in the complemented circuit directed graph, and its operation formula is as follows:
[0070] h (l+1) =σ(A recon h (l) W (l) )
[0071] Among them, h (l) is the node feature matrix of the l-th layer, and the initially input node feature matrix h (0) =F is the circuit gate feature F obtained in step C; W (l) represents the training weight matrix of the l-th layer; σ(·) represents the activation function in the graph convolution operation; through multi-layer convolution operations, the features of the nodes gradually aggregate the information from their neighborhoods, thereby generating the final feature H for input to the classification model;
[0072] Step D2, construct a multi-layer perceptron; input the final feature H obtained in step D1 into the multi-layer perceptron to infer the probability that the logic gate belongs to a hardware Trojan, and its operation formula is as follows:
[0073] p = MLP(H)
[0074] Among them, MLP(·) is the multi-layer perceptron, and p is the probability vector obtained through the multi-layer perceptron;
[0075] Step D3, input the hardware Trojan probability of the logic gate obtained in step D2 into the Softmax activation function for probability normalization to obtain the classification result of the node, and its operation formula is as follows:
[0076]
[0077] Among them, is the predicted type, and the predicted types are divided into two categories: hardware Trojan gate nodes and normal gate nodes; Softmax(·) is the Softmax activation function;
[0078] Step D4. Calculate the difference between the prediction result and the true label for the nodes with the initial netlist label of 1 in the complemented directed graph using the cross-entropy loss function. The operation formula is as follows:
[0079]
[0080] where N represents the number of non-complemented generated gate nodes in the complemented directed graph; y i represents the true classification label of the i-th gate node in the Trojan label Y; represents the classification result of the i-th node predicted by the model;
[0081] Step D5. Through the backpropagation mechanism, optimize the model parameters according to the value of the loss function L, where the optimization algorithm uses the Adam optimizer; terminate the training of the model when the iteration round of the loss value reaches the preset value;
[0082] Step D6. Input the complemented circuit netlist to be tested into the trained hardware Trojan gate classification model for detection, and output the classification results of the non-complemented generated gate nodes in the complemented directed graph.
[0083] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be run by a processor are stored. When the processor runs the computer program instructions, the method steps described above can be implemented.
[0084] Compared with the prior art, the present invention has the following beneficial effects:
[0085] (1) The present invention proposes a security design detection method for very large-scale integrated circuits facing information loss, effectively addressing the challenges in the prior art that it is impossible to effectively extract netlist features for the gate-level netlist with information loss and the detection results are invalid due to information loss.
[0086] (2) The present invention proposes a framework for complementing the circuit structure with missing information based on graph variational autoencoders. The encoder encodes the circuit netlist with missing structure and maps the encoding result to the latent space, and the decoder is used to complement the incomplete circuit structure for the latent variables, providing more comprehensive information for the detection of hardware Trojans in the circuit netlist and effectively alleviating the problem of invalid detection results caused by insufficient decision-making information.
[0087] (3) The present invention proposes a hardware Trojan classification model using local information as features, effectively solving the defect in the prior art that it is impossible to effectively obtain global information such as the main inputs, main outputs, and global positions of gate nodes due to information loss, thereby effectively detecting hardware Trojans in the gate-level netlist with missing information.
[0088] (4) The present invention can be used for the completion of the information - missing gate - level netlist, the extraction of local features of gate nodes, and the effective detection of hardware Trojan nodes, which helps to take more targeted defense measures in the subsequent processes of integrated circuit design. Description of the Drawings
[0089] Figure 1 It is a flowchart of the method according to an embodiment of the present invention.
[0090] Figure 2 It is a flowchart for implementing step A of an embodiment of the present invention.
[0091] Figure 3 It is a flowchart for implementing step B of an embodiment of the present invention.
[0092] Figure 4 It is a flowchart for implementing step C of an embodiment of the present invention.
[0093] Figure 5 It is a flowchart for implementing step D of an embodiment of the present invention. Detailed Embodiment
[0094] The technical solution of the present invention will be specifically described below with reference to the drawings.
[0095] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0096] It should be noted that the terms used herein are only for describing the specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0097] The present invention provides a security design detection method for very - large - scale integrated circuits (VLSIs) facing information loss, including:
[0098] Parsing the gate - level netlist design file of the information - missing VLSI, creating a directed - graph representation and encoding the feature matrix of logic gate nodes to construct circuit directed - graph data;
[0099] Using the graph variational auto - encoder (GraphVAE) to perform structural inference on the directed - graph data and using the decoder to complete the information - missing circuit structure to construct the completed circuit directed - graph data;
[0100] Calculate the logic gate type, fan-in and fan-out quantity, betweenness centrality, and neighborhood type distribution for the completed circuit directed graph data, and combine and splice them into circuit gate features;
[0101] Construct a graph neural network model GNN, take the completed circuit directed graph data and circuit gate features as inputs, further extract deep features through graph convolution operations, and use a multi-layer perceptron for inference training to achieve hardware Trojan classification.
[0102] The following are specific implementation examples of the present invention.
[0103] As Figures 1-5 shown, this embodiment provides a security design detection method for very large-scale integrated circuits facing information loss, including the following steps:
[0104] Step A: Parse the obtained gate-level netlist design file of the very large-scale integrated circuit with information loss, and convert the circuit structure into a directed graph form GI=(V, E). In the directed graph, V represents the logic gates in the netlist, and E represents the connection relationship between the logic gates. Convert the directed graph data into an adjacency matrix A. At the same time, perform One-Hot encoding on the logic gate types in the directed graph as the type feature representation X. Further, label the hardware Trojan label Y for each logic gate, and construct circuit directed graph data containing type features and label information;
[0105] Step A1: Obtain the gate-level netlist design file of the very large-scale integrated circuit with information loss for training, and parse its structure to collect all the circuit gate types involved in the netlist.
[0106] Step A2: Extract the logic gate set V={v 0 ,v 1 ,v 2 ,...,v n} and the connection set E={(v i ,v j )|v i ,v j ∈V} from the gate-level netlist.
[0107] Step A3: Use the logic gate set V as vertices and the connection set E between the logic gates as edges to construct a directed graph. Further, convert the directed graph into an adjacency matrix A. If (v i ,v j ) exists in the edge set, then the corresponding element α ij in the adjacency matrix is 1, otherwise it is 0.
[0108] Step A4: For the convenience of model processing, the types of each logic gate node in the netlist are characterized using One-Hot encoding. The length of the One-Hot encoding is the number of circuit gate types collected in Step A1. For each logic gate, the position corresponding to the One-Hot type is set to 1, and the remaining positions are set to 0. In this patent, 32-bit encoding is used as the type feature encoding. Further, all the logic gate type encodings are merged into a node feature matrix X in the order of the nodes in the adjacency matrix A.
[0109] Step A5: Parse the Trojan nodes in the netlist, label the Trojan logic gate nodes with Trojan labels and the normal logic gate nodes with normal labels to form the label information Y of the netlist; the initial incomplete directed graph data is composed of the adjacency matrix A, the logic gate type feature representation X, and the label information Y.
[0110] Step B: Perform structure inference and completion on the incomplete circuit structure in the directed graph circuit data. Build a graph variational autoencoder model GraphVAE, use the directed graph adjacency matrix A and the logic gate type X as the input of GraphVAE for structure inference training, calculate whether there are any missing parts in the circuit structure in terms of function and topological relationship, and use the decoder to complete the missing structure to obtain the completed adjacency matrix A recon and the logic gate type matrix
[0111] Step B1: Take the directed graph data in Step A as the input. To improve the model stability, perform a normalization operation on the adjacency matrix. The normalization operation is as follows:
[0112]
[0113] where represents the adjacency matrix with self-loops added, and I n represents the identity matrix; is the degree matrix calculated through the adjacency matrix
[0114] Step B2: Build a graph convolutional layer to extract the node features in the directed graph. Its operation formula is as follows:
[0115]
[0116] where H (l) is the node feature matrix of the l-th layer, and H (0) = X is the initial node feature matrix; W (l) Denote the training weight matrix of the l-th layer; σ(·) represents the activation function in the graph convolution operation. In this patent, the ReLU function is adopted as the activation function in step B2; through multi-layer graph convolution operations, the node feature matrix is updated layer by layer to capture high-order topological information and local structural features.
[0117] Step B3: Aggregate the node feature matrix H obtained in step B2 through a pooling operation (l+1) into a global vector h, representing the structural information of the circuit. The operation formula is as follows:
[0118] h = pooling(H (l+1) )
[0119] where pooling(·) represents the pooling operation function.
[0120] Step B4: Perform a latent space mapping on the global vector. Map the global vector h to the latent space through a fully connected layer to generate distribution parameters. The operation formula is as follows:
[0121] μ = FC μ (h),
[0122] logσ = FC σ (h)
[0123] where μ and logσ represent the mean and log variance of the latent space respectively; FC μ (·) and FC σ (·) represent the fully connected operations for the mean and log variance of the latent space respectively.
[0124] Step B5: Sample latent variables from the latent space distribution through a reparameterization operation. The operation formula is as follows:
[0125] z = μ + σ · ε
[0126] where z represents the sampled latent variable; ε ∼ N(0, I) is the standard normal distribution.
[0127] Step B6: Use the sampled latent variables to complete the node features and adjacency matrix through a decoder. The operation formula is as follows:
[0128]
[0129] A recon = σ(zz T )
[0130] where, denotes the reconstructed node feature matrix; A recon denotes the completed adjacency matrix; FC node(·) represents the fully connected operation for generating the node feature matrix; σ(·) represents the activation function; z T represents the transpose of the latent variable. In this patent, the activation function in step B6 is the ReLU function.
[0131] Step B7: Construct the loss function of the completion model. In the training and inference of the incomplete circuit structure completion, the weighted sum of the reconstruction loss and the KL divergence loss is used as the total loss function.
[0132] The reconstruction loss is used to measure the difference between the completed graph and the input graph, including the reconstruction errors of the node features and the adjacency matrix. Its operation formula is as follows:
[0133]
[0134] where represents the L2 norm; represents the Frobenius norm.
[0135] The KL divergence is used to constrain the distribution in the latent space to be close to the standard normal distribution. Its operation formula is as follows:
[0136]
[0137] where μ i and σ i represent the mean and variance of the latent variable for each dimension i respectively.
[0138] Combining the reconstruction loss and the KL divergence loss, the total loss function is:
[0139] L total = L recon + λL KL
[0140] where λ represents the weight of the loss function.
[0141] Step B8: Through the backpropagation mechanism, optimize the model parameters according to the value of the loss function L total . Among them, the optimization algorithm uses the Adam optimizer. When the iteration round of the loss value reaches the preset value, terminate the training of the model.
[0142] Step B9: Input the circuit netlist with missing information to be measured into the trained inference completion model for completion, and output the completed directed graph; meanwhile, create an initial netlist label in the completed directed graph to determine whether a node is a completed node. If it is a node generated by completion, the label value is set to 0, otherwise, it is set to 1.
[0143] Step C: Feature selection is performed on the complemented circuit directed graph data, and the main features extracted include: logic gate type, logic gate fan-in number, logic gate fan-out number, betweenness centrality of the logic gate, and neighborhood type distribution of the logic gate. The calculated feature combinations are concatenated into the final circuit gate feature F;
[0144] Step C1: The complemented circuit directed graph in step B is used for feature selection. Each logic gate in the directed graph is traversed, and the number of input ports and output ports of each logic gate is collected, corresponding to the logic gate fan-in number FI and fan-out number FO respectively.
[0145] Step C2: Use the breadth-first search algorithm to calculate the number of shortest paths from each gate node to other gate nodes in the directed graph. Starting from the leaf nodes of the shortest path tree, the dependency relationship between nodes is propagated backward, and the betweenness centrality of all gate nodes is recorded. The operation formula is as follows:
[0146]
[0147] where, σ ij represents the total number of shortest paths from gate node i to gate node j; σ ij (v) represents the number of paths passing through gate node v in the shortest path from i to j.
[0148] Step C3: Traverse the gate nodes in the directed graph, and count and record the type distribution of the nodes in the neighborhood of each gate node. The type distribution feature of the neighborhood is represented by integer encoding, where the length of the encoding is the number of circuit gate types collected in step A. The initial value of each encoding bit is 0, and the count is performed according to the number of nodes of different types in the neighborhood, so as to generate the complete neighborhood type distribution feature vector ND.
[0149] Step C4: Obtain each node type feature vector XN in the node feature matrix obtained in step B. The node feature vector XN, fan-in number FI, fan-out number FO, betweenness centrality BC, and neighborhood type distribution feature vector ND of each node are concatenated, and finally merged into the feature matrix F for hardware Trojan classification. In this patent, the encoding lengths of XN and ND are both 32 bits, FI, FO, and BC are all 1 bit, and the jointly constructed node feature representation encoding length is 67 bits.
[0150] Step D: Perform inference classification on the obtained circuit gate feature F. Construct a graph neural network model GNN, and use the information-complemented adjacency matrix A reconAnd the circuit gate feature F is used as the input of the GNN. Deeper features H are extracted through graph convolution operations, and H is input into the multi-layer perceptron layer in the GNN for inference training. Calculate the probabilities that each circuit gate belongs to a normal circuit gate and a Trojan circuit gate. Finally, the classification of hardware Trojans is achieved.
[0151] Step D1: Construct a graph convolution layer to extract the feature representation of each node in the complemented circuit directed graph. Its operation formula is as follows:
[0152] h (l+1) = σ(A recon h (l) W (l) )
[0153] Among them, h (l) is the node feature matrix of the l-th layer. The initial input node feature matrix h (0) = F is the feature matrix F obtained in step C; W (l) represents the training weight matrix of the l-th layer; σ(·) represents the activation function in the graph convolution operation. In this patent, the activation function in step D1 adopts the ReLU function. Through multi-layer convolution operations, the features of nodes gradually aggregate the information from their neighborhoods, thereby generating the final feature H for the input of the classification model.
[0154] Step D2: Construct a multi-layer perceptron. Input the final feature H obtained in step D1 into the multi-layer perceptron to infer the probability that the logic gate belongs to a hardware Trojan. Its operation formula is as follows:
[0155] p = MLP(H)
[0156] Among them, MLP(·) is the multi-layer perceptron, and p is the probability vector obtained through the multi-layer perceptron.
[0157] Step D3: Input the hardware Trojan probability of the logic gate obtained in step D2 into the Softmax activation function for probability normalization to obtain the classification result of the node. Its operation formula is as follows:
[0158]
[0159] Among them, is the predicted type. In this patent, the predicted types are divided into two categories: hardware Trojan gate nodes and normal gate nodes; Softmax(·) is the Softmax activation function.
[0160] Step D4: For the nodes with the initial netlist label of 1 in the complemented directed graph, calculate the difference between the prediction result and the true label using the cross-entropy loss function. Its operation formula is as follows:
[0161]
[0162] Among them, N represents the number of non-complemented generated gate nodes in the complemented directed graph; y i represents the true classification label of the i-th gate node in the Trojan label Y; represents the classification result of the i-th node predicted by the model.
[0163] Step D5: Through the backpropagation mechanism, optimize the model parameters according to the value of the loss function L, where the optimization algorithm uses the Adam optimizer. When the number of loss value iteration rounds reaches the preset value, terminate the training of the model.
[0164] Step D6: Input the complemented circuit netlist to be tested into the trained hardware Trojan gate classification model for detection, and output the classification results of the non-complemented generated gate nodes in the complemented directed graph.
[0165] The present invention also provides a computer-readable storage medium, on which computer program instructions that can be run by a processor are stored. When the processor runs the computer program instructions, the method steps described in any of the above can be implemented.
[0166] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0167] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0168] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions in the process Figure 1one or more processes and / or blocks Figure 1 the functions specified in one or more blocks.
[0169] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more processes and / or blocks Figure 1 one or more blocks.
[0170] As described above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A method for detecting the security design of a very large scale integrated circuit for information loss, characterized in that: include: Parse the VLSI gate-level netlist design files with missing information, create directed graph representations, encode the feature matrix of logic gate nodes, and construct circuit directed graph data; Use the graph variational autoencoder GraphVAE to perform structural reasoning on directed graph data, and use the decoder to complete the circuit structure with missing information to construct the completed circuit directed graph data; The completed circuit directed graph data is used to calculate the logic gate type, fan-in and fan-out number, betweenness centrality, and neighborhood type distribution, and then combined and spliced into circuit gate features; A graph neural network model GNN is constructed, and the completed circuit directed graph data and circuit gate features are taken as input. Deep features are further extracted through graph convolution operations, and multi-layer perceptron is used for inference training to realize hardware Trojan classification.
2. The VLSI security design detection method for information loss according to claim 1, characterized in that: The steps include: Step A, parsing the obtained VLSI gate-level netlist design file with missing information, converting the circuit structure into a directed graph form GI=(V,E); in the directed graph, V represents the logic gate in the netlist, and E represents the connection relationship between the logic gates; converting the directed graph data into a directed graph adjacency matrix A; at the same time, performing One-Hot encoding on the type of the logic gate node in the directed graph as the logic gate node feature matrix X; Furthermore, each logic gate is labeled with a hardware Trojan label Y, and circuit directed graph data including type features and label information is constructed; Step B: Perform structural reasoning and completion on the incomplete circuit structure in the circuit directed graph data; construct a graph variational autoencoder GraphVAE, take the directed graph adjacency matrix A and the logic gate node feature matrix X as the input of GraphVAE, perform structural reasoning training, calculate whether the circuit structure is missing in function and topological relationship, and use the decoder to complete the missing structure to obtain the completed adjacency matrix A recon And the logic gate node feature matrix Step C: feature selection is performed on the completed circuit directed graph data, and the extracted features include: logic gate type, logic gate fan-in number, logic gate fan-out number, logic gate betweenness centrality, and logic gate neighborhood type distribution; the calculated feature combinations are spliced into the final circuit gate feature F; Step D: Infer and classify the obtained circuit gate features F; construct a graph neural network model GNN, and transform the completed adjacency matrix A recon The circuit gate feature F is used as the input of GNN, and the deeper feature H is extracted through graph convolution operation, and H is input into the multi-layer perceptron layer in GNN for inference training; the probability of each circuit gate belonging to a normal circuit gate and a Trojan circuit gate is calculated; finally, the classification of hardware Trojans is realized.
3. The VLSI security design detection method for information loss according to claim 2, characterized in that: The step A comprises the following steps: Step A1, obtaining a VLSI gate-level netlist design file with missing information for training, and parsing its structure to collect all circuit gate types involved in the netlist; Step A2: Extract the logic gate set V = {v0, v1, v2, ..., v n } and the connection set E between logic gates = {(v i ,v j )|v i ,v j ∈V}, where v i 、v j They represent the i-th and j-th logic gates extracted from the gate-level netlist, respectively, and n represents the total number of logic gates in the gate-level netlist; Step A3: construct a directed graph with the logic gate set V as vertices and the connection set E between the logic gates as edges; further, convert the directed graph into an adjacency matrix A. If there is (v i ,v j ), then the corresponding element α in the adjacency matrix ij =1, otherwise 0; Step A4: Use One-Hot coding to encode the type of each logic gate node in the gate-level netlist into a feature matrix; the length of the One-Hot coding is the number of circuit gate types collected in step A1; for each logic gate, the position of the One-Hot corresponding type is set to 1, and the remaining positions are 0; further, all logic gate node feature matrices are merged into a logic gate node feature matrix X according to the node order in the adjacency matrix A; Step A5, parse the Trojan nodes in the netlist, mark the Trojan logic gate nodes with Trojan labels, and mark the normal logic gate nodes with normal labels, to form the label information Y of the netlist; the adjacency matrix A, the logic gate node feature matrix X and the label information Y constitute the initial incomplete directed graph data.
4. The VLSI security design detection method for information loss according to claim 3, characterized in that: In step A4, 32-bit code is used as the feature matrix code.
5. The VLSI security design detection method for information loss according to claim 2, characterized in that: The step B comprises the following steps: Step B1: Take the circuit directed graph data of step A as input and perform normalization operation on the adjacency matrix. The normalization operation is as follows: in, Represented as the adjacency matrix with self-loops added, I n Represented as the identity matrix; To pass the adjacency matrix The calculated degree matrix; Step B2: Construct a graph convolution layer to extract node features in the directed graph. The calculation formula is as follows: Among them, H (l) is the node feature matrix of the lth layer, H (0) =X is the initial node feature matrix; W (l) represents the training weight matrix of the lth layer; σ(·) represents the activation function in the graph convolution operation; through multi-layer graph convolution operations, the node feature matrix is updated layer by layer to capture high-order topological information and local structural features; Step B3: Pool the node feature matrix H obtained in step B2 through a pooling operation (l+1) Aggregated into a global vector h, representing the structural information of the circuit, the calculation formula is as follows: h=pooling(H (l+1) ) Among them, pooling(·) represents the pooling operation function; Step B4: perform latent space mapping on the global vector, map the global vector h to the latent space through a fully connected layer, and generate distribution parameters; the calculation formula is as follows: μ=FC μ (h), logσ=FC σ (h) Where μ and logσ represent the mean and logarithmic variance of the latent space, respectively; FC μ (·) and FC σ (·) denote the fully connected operations of the latent space mean and log-variance, respectively; Step B5: Sample latent variables from the latent space distribution through reparameterization operation, and the calculation formula is as follows: z=μ+σ·ε Where z represents the latent variable obtained by sampling; ε~N(0,I) is the standard normal distribution; Step B6: Using the latent variables obtained through sampling, the decoder completes the node features and adjacency matrix. The calculation formula is as follows: A recon =σ(zz T ) in, Represented as the reconstructed node feature matrix; A recon Represented as the completed adjacency matrix; FC node (·) represents the fully connected operation to generate the node feature matrix; σ(·) represents the activation function; z T denoted as the transpose of the latent variable; Step B7: Construct the loss function of the completion model. In the training and reasoning of the incomplete circuit structure completion, the weighted addition of the reconstruction loss and the KL divergence loss is used as the total loss function L total ; Step B8: Through the back propagation mechanism, according to the total loss function L total The model parameters are optimized using the Adam optimizer. The training of the model is terminated when the loss value reaches the preset value in the iteration round. Step B9: input the circuit netlist with missing information to be tested into the trained inference completion model for completion, and output the completed directed graph; at the same time, create an initial netlist label in the completed directed graph to determine whether the node is a completed node; if it is a node generated by completion, the label value is set to 0, otherwise, it is set to 1.
6. The VLSI security design detection method for information loss according to claim 5, characterized in that: In step B7, the loss function calculation includes the following contents: The reconstruction loss is used to measure the difference between the completed graph and the input graph, including the reconstruction error of the node feature matrix and the adjacency matrix. Its calculation formula is as follows: in, Expressed as L2 norm; Expressed as the Frobenius norm; KL divergence is used to constrain the distribution of the latent space to be close to the standard normal distribution. Its calculation formula is as follows: Among them, μ i and σ i Represent the mean and variance of the latent variable of each dimension i respectively; Combining the reconstruction loss and KL divergence loss, the total loss function is: THE total =L recon +λL KL Among them, λ represents the weight of the loss function.
7. The VLSI security design detection method for information loss according to claim 2, characterized in that: The step C comprises the following steps: Step C1, using the circuit directed graph data completed in step B for feature selection, traversing each logic gate of the directed graph, collecting the number of input ports and output ports of each logic gate, which correspond to the fan-in number FI and the fan-out number FO of the logic gate respectively; Step C2: Use the breadth-first search algorithm to calculate the number of shortest paths from each logic gate node to other logic gate nodes in the directed graph, starting from the leaf node of the shortest path tree, back-propagate the dependencies between nodes, and record the betweenness centrality of all logic gates; the calculation formula is as follows: Among them, σ ij It is expressed as the total number of shortest paths from logic gate node i to logic gate node j; σ ij (v) represents the number of paths passing through logic gate node v in the shortest path from i to j; Step C3, traverse the logic gate nodes in the directed graph, count and record the type distribution of nodes in the neighborhood of each logic gate node; the type distribution characteristics of the neighborhood are represented by integer coding, where the length of the coding is the number of circuit gate types collected in step A; the initial value of each coding bit is 0, and the number of different types of nodes in the neighborhood is counted, thereby generating a complete neighborhood type distribution feature vector ND of the logic gate; Step C4: Obtain the logic gate node feature matrix obtained in step B Each node feature vector XN in the matrix is concatenated; the node feature vector XN, the fan-in number FI of the logic gate, the fan-out number FO of the logic gate, the betweenness centrality BC of the logic gate, and the neighborhood type distribution feature vector ND of the logic gate are concatenated, and finally merged into a feature matrix F for hardware Trojan classification.
8. The VLSI security design detection method for information loss according to claim 7, characterized in that: In step C4, the encoding lengths of XN and ND are both 32 bits, and FI, FO, and BC are all one bit, and the encoding length of the node feature matrix together constituted is 67 bits.
9. The VLSI security design detection method for information loss according to claim 2, characterized in that: The step D comprises the following steps: Step D1: construct a graph convolutional layer to extract and complete the feature representation of each node in the directed graph of the circuit. The calculation formula is as follows: h (l+1) =σ(A recon h (l) W (l) ) Among them, h (l) is the node feature matrix of the lth layer, and the initial input node feature matrix h (0) =F is the circuit gate feature F obtained in step C; W (l) is represented as the training weight matrix of the lth layer; σ(·) represents the activation function in the graph convolution operation; through multi-layer convolution operations, the features of the nodes gradually aggregate the information from their neighborhood to generate the final feature H for the input of the classification model; Step D2, construct a multi-layer perceptron; input the final feature H obtained in step D1 into the multi-layer perceptron, and infer the probability that the logic gate belongs to a hardware Trojan. The calculation formula is as follows: p=MLP(H) Where MLP(·) is a multi-layer perceptron, and p is a probability vector obtained by the multi-layer perceptron; Step D3: Input the hardware Trojan probability of the logic gate obtained in step D2 into the Softmax activation function for probability normalization to obtain the classification result of the node. The calculation formula is as follows: in, is the predicted type, which is divided into two categories: hardware Trojan gate node and normal gate node; Softmax(·) is the Softmax activation function; Step D4: For the nodes whose initial netlist labels are 1 in the completed directed graph, the cross entropy loss function is used to calculate the difference between the predicted result and the true label. The calculation formula is as follows: Where N represents the number of gate nodes that are not generated by completion in the completed directed graph; y i represents the true classification label of the i-th gate node in the Trojan label Y; Represents the classification result of the i-th node predicted by the model; Step D5: Optimize the model parameters according to the value of the loss function L through the back-propagation mechanism, wherein the optimization algorithm adopts the Adam optimizer; terminate the model training when the loss value iteration round reaches the preset value; Step D6: input the completed circuit netlist to be tested into the trained hardware Trojan gate classification model for detection, and output the classification results of the non-completed gate nodes in the completed directed graph.
10. A computer-readable storage medium having stored thereon computer program instructions that can be executed by a processor, and when the processor executes the computer program instructions, the method steps according to any one of claims 1 to 9 can be implemented.
Citation Information
Patent Citations
Integrated circuit Trojan horse detection method and system
CN115204078A
Hardware Trojan horse detection method and device based on gated recurrent neural network
CN118607018A
Structural information and semantic information fused hardware Trojan horse detection method
CN119249421A
Cited By
Circuit netlist optimization method and computer program product
CN121766231A