A method and system for building a deep representation model and extracting features of hardware Trojans
Through deep learning algorithms, the hardware Trojan deep representation model is constructed, combined with static and dynamic features, and the problem of insufficient feature extraction and model generalization capabilities in the existing technology is solved, and more accurate and reliable hardware Trojan detection is achieved.
Patent Information
- Application Number
- CN202510238301.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The existing hardware Trojan detection technology is difficult to effectively combine the static and dynamic characteristics of the hardware, and there is the problem of insufficient overfitting and generalization capabilities, making it difficult to achieve effective training and feature extraction under limited samples.
Deep learning algorithm is used to build a deep-level representation model of hardware Trojans. By combining convolutional neural networks and recurrent neural networks, static and dynamic features of the hardware are extracted, and regularization terms and cross-verification methods are introduced during model training to optimize the generalization ability and detection effect of the model.
A more comprehensive and accurate representation of hardware Trojans is achieved, the accuracy and reliability of detection is improved, the risk of overfitting is reduced, and the model training speed is accelerated under limited samples.
Smart Images

Figure CN119720060B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of model construction and extraction systems, and in particular to a method and system for constructing a deep-level representation model of a hardware Trojan and extracting features. Background Art
[0002] With the rapid development of computer technology and integrated circuit technology, hardware systems are increasingly used in various key fields, from military equipment, aerospace to finance, communications and other civilian fields. The security of hardware systems has become extremely important. However, hardware Trojans, as a serious hardware security threat, are gradually becoming an important factor threatening the security of hardware systems. Hardware Trojans are malicious hardware modifications that can be implanted into hardware such as integrated circuits and activated under certain conditions, thereby leaking confidential information, tampering with system functions or reducing system performance, bringing huge security risks to users and the entire system.
[0003] Traditional methods for detecting hardware Trojans mainly include methods based on circuit structure analysis, side channel analysis, and functional testing. Methods based on circuit structure analysis usually require an in-depth understanding of the hardware design documents and physical layout. For complex hardware systems, this method is often limited by the difficulty of obtaining information and the complexity of the design. Although the side channel analysis method can detect hardware Trojans through information such as power consumption and electromagnetic radiation, this method is easily interfered by environmental noise, and its detection accuracy and reliability are limited for new and complex hardware Trojans. Functional testing methods rely on the understanding of the normal functions of the hardware. It is difficult to detect Trojans hidden in complex functional logic through conventional functional testing.
[0004] In recent years, deep learning technology has achieved great success in image recognition, natural language processing and other fields, showing powerful feature extraction and pattern recognition capabilities. Therefore, introducing deep learning technology into the field of hardware Trojan detection has become a research hotspot. By building a deep representation model, deep learning can automatically learn the feature representation of hardware and overcome some limitations of traditional methods. However, there are still some problems with the existing hardware Trojan detection technology based on deep learning. On the one hand, in the process of building a deep representation model for hardware Trojans, there is a lack of a perfect method for how to effectively combine the static and dynamic characteristics of hardware. The static characteristics of the hardware system, such as circuit structure and logic gate type, and the dynamic characteristics, such as signal timing and power consumption characteristics, contain important hardware information, but it is still a challenge to integrate this information into the deep learning model and perform effective feature extraction and utilization. On the other hand, when building the model and extracting features, how to ensure the generalization ability of the model, avoid overfitting, and how to perform effective training under limited hardware Trojan samples are all urgent problems to be solved. In addition, there is also a lack of systematic solutions for how to effectively reduce the dimensionality and visualize the extracted features so that users can intuitively understand and analyze them. Therefore, developing a deep-level characterization model construction and feature extraction method and system for hardware Trojans that can comprehensively consider hardware characteristics, have powerful detection capabilities and good visualization has important practical significance and practical value for improving the security of hardware systems. Summary of the invention
[0005] The present invention proposes a method and system for building a deep-level representation model and extracting features of hardware Trojans to solve the problems mentioned in the above-mentioned prior art.
[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: a method for constructing a deep-level representation model and extracting features of a hardware Trojan, comprising the following steps:
[0007] Collect hardware Trojan sample data and normal hardware data without Trojans, and pre-process the data, including data cleaning and normalization. Data cleaning is achieved by removing outliers and noise data. A statistical analysis-based method is used to calculate the mean and standard deviation of the data. Data that deviates from the mean by more than a set multiple of the standard deviation is considered an outlier and is removed. When normalizing the data, the minimum-maximum normalization method is used to normalize the data to the range of [0,1]. The formula is: , where X is the original data, and are the minimum and maximum values in the data set, respectively;
[0008] A deep representation model of hardware Trojans is built based on a learning algorithm. The learning algorithm includes a convolutional neural network or a recurrent neural network. For the convolutional neural network, a convolutional layer, a pooling layer, and a fully connected layer are built. The convolutional layer uses convolution kernels of different sizes for feature extraction. The pooling layer uses maximum pooling or average pooling operations to downsample the feature map. The recurrent neural network uses a long short-term memory network LSTM or a gated recurrent unit GRU to process the sequence information of the hardware data.
[0009] The constructed model is trained with the data. The model parameters are adjusted to make the model distinguish between hardware Trojan data and normal hardware data. The stochastic gradient descent SGD algorithm is used to optimize the loss function during the training process. In addition, in order to prevent overfitting, regularization terms are introduced, including L1 or L2 regularization. The loss function is: or ,in is the original loss function, is the regularization coefficient, w is the weight of the model;
[0010] The hardware data to be detected is input into the trained model, and the feature vector of the hardware Trojan is extracted. The presence of a Trojan in the hardware is determined based on the feature vector and the type of Trojan is determined. When making the judgment, a threshold is set. When the output value corresponding to the feature vector exceeds the threshold, it is determined that a hardware Trojan exists. At the same time, based on the distribution of the feature vector values in different dimensions, the type of Trojan is determined through a predefined rule set, which is obtained through feature analysis of a large number of known hardware Trojan samples.
[0011] Furthermore, feature engineering is performed on the collected data to extract static and dynamic features of the hardware. Static features include circuit structure features and logic gate types of the hardware. The circuit structure is parsed using the hardware description language HDL to identify the types and connection relationships of logic gates, which are then converted into quantifiable feature representations. Dynamic features include signal timing features and power consumption features of the hardware during operation. Signal timing features are acquired at a specific sampling frequency by a hardware performance monitoring tool during hardware operation, and frequency domain features are extracted by Fourier transform spectrum analysis. Power consumption features are acquired by setting power sensors at nodes in the hardware circuit to obtain power consumption data under different working states. Features and preprocessed data are used as inputs to the model.
[0012] Furthermore, during the model training process, the cross-validation method is used to evaluate and optimize the performance of the model. The training data is divided into multiple subsets, and one subset is used as the validation set in turn, and the remaining subsets are used as the training set. K-fold cross-validation is used, and the choice of K value is determined according to the scale and complexity of the data set, and the value range is between 5 and 10. In each round of cross-validation, the model parameters are adjusted, including the learning rate, batch size, convolution kernel size, and number of network layers.
[0013] Furthermore, the model is optimized using transfer learning technology, and the model parameters trained on the hardware dataset are migrated to the current hardware Trojan detection model. According to the similarity between the source dataset and the target dataset, some layers of the pre-trained model are selectively frozen, and only some layers are fine-tuned. For convolutional neural networks, the front convolutional layers are frozen and the rear fully connected layers are trained. For recurrent neural networks, some LSTM or GRU units are frozen. At the same time, during the migration process, the migrated parameters are scaled or offset according to the characteristics of the target dataset.
[0014] Furthermore, in the feature extraction process, the principal component analysis PCA or linear discriminant analysis LDA dimensionality reduction algorithm is used to reduce the dimensionality of the extracted feature vectors. For PCA dimensionality reduction, the covariance matrix of the feature vector is calculated to solve its eigenvalues and eigenvectors. For LDA dimensionality reduction, the generalized eigenvalue problem is solved by calculating the between-class scatter matrix and the within-class scatter matrix.
[0015] Furthermore, the detection results are visualized, showing whether there are Trojans in the hardware, the types of Trojans, and the distribution of their characteristics in the form of charts or graphs.
[0016] A system using the hardware Trojan deep-level representation model construction and feature extraction method, comprising:
[0017] Data acquisition module: used to collect hardware Trojan sample data and normal hardware data without Trojans, and transmit the data to the data preprocessing module; the module data source interface is connected with the hardware test platform and simulation tool, receives hardware operation data through the protocol, and performs integrity and consistency checks on the received data;
[0018] Data preprocessing module: cleans, normalizes and performs feature engineering on the received data, and outputs the processed data to the model building and training module; the module includes an abnormal data detection submodule, which identifies abnormal values through the set data distribution range and statistical rules, and a normalization processing submodule, which uses the normalization formula to convert data, including a feature extraction submodule, which uses feature engineering methods to extract static and dynamic features;
[0019] Model building and training module: builds a deep representation model of hardware Trojans based on learning algorithms, trains the model through training data, and transmits the trained model to the feature extraction and detection module; the module includes a learning framework selection unit, which selects the learning framework according to the characteristics of different hardware data, and a training parameter adjustment unit;
[0020] Feature extraction and detection module: The hardware data to be detected is input into the trained model, the feature vector of the hardware Trojan is extracted, and the presence of the Trojan in the hardware is determined based on the feature vector and the type of Trojan is determined. The detection result is transmitted to the result visualization module; the feature extraction unit in the module uses the intermediate layer information output by the model as the final feature vector, and the detection unit determines the presence and type of the Trojan through threshold comparison and rule matching algorithm;
[0021] Result visualization module: Visualize the test results and present them to users in the form of charts or graphs.
[0022] Furthermore, the data collection module includes a data storage unit for storing the collected hardware Trojan sample data and normal hardware data. The storage unit adopts a distributed storage architecture to encrypt and store the data.
[0023] Furthermore, the model building and training module includes a model evaluation unit, which uses a cross-validation method to evaluate and optimize the performance of the model. The training data is divided into multiple subsets, and one subset is used as a validation set in turn, and the remaining subsets are used as training sets. K-fold cross-validation is used, and the selection of K value is determined according to the scale and complexity of the data set, and the value range is between 5 and 10. In each round of cross-validation, the model parameters are adjusted, including the learning rate, batch size, convolution kernel size, and number of network layers.
[0024] Furthermore, the feature extraction and detection module includes a dimensionality reduction processing unit. During the feature extraction process, the principal component analysis PCA or linear discriminant analysis LDA dimensionality reduction algorithm is used to perform dimensionality reduction processing on the extracted feature vector. For PCA dimensionality reduction, the covariance matrix of the feature vector is calculated to solve its eigenvalues and eigenvectors. For LDA dimensionality reduction, the generalized eigenvalue problem is solved by calculating the between-class scatter matrix and the within-class scatter matrix.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] First, the present invention provides an innovative deep-level characterization model construction and feature extraction method in hardware Trojan detection. By combining a deep learning algorithm, it can automatically learn the complex features of hardware and effectively overcome the limitations of traditional hardware Trojan detection methods. This method can make full use of the static and dynamic characteristics of hardware, including circuit structure, logic gate type, signal timing characteristics, and power consumption characteristics, so that the model's characterization of hardware Trojans is more comprehensive and accurate, improving the accuracy and reliability of hardware Trojan detection.
[0027] Secondly, by using a variety of technical means to optimize the model training process, such as cross-validation and transfer learning, we can not only improve the generalization ability of the model and prevent overfitting, but also speed up the model training speed and improve the training effect under limited hardware Trojan samples, providing more powerful technical support for hardware Trojan detection. Introducing regularization terms during model training helps adjust the complexity of the model and enhance the robustness of the model.
[0028] Furthermore, the use of dimensionality reduction algorithms, such as PCA or LDA, in the feature extraction process can remove redundant information, improve the effectiveness of feature vectors, reduce computing costs, and improve detection efficiency. In addition, the visualization of detection results uses data visualization tools to clearly present to users whether there are Trojans in the hardware, the type of Trojans, and the distribution of related features, making it easier for users to intuitively understand and analyze the detection results, helping users quickly understand the security status of the hardware and provide a strong basis for subsequent decision-making and processing.
[0029] Finally, the system of the present invention can realize the automation and intelligence of hardware Trojan detection, reduce the dependence on manual experience and professional knowledge, and provide an efficient and reliable solution for the research and practice in the field of hardware security, which is helpful to improve the security and stability of hardware systems. It has broad application prospects and huge social benefits in key fields such as military, finance, and communications that have extremely high requirements for hardware security. It will promote the advancement of hardware security protection technology, reduce the potential risks caused by hardware Trojans to the system, and ensure the normal operation of hardware systems and data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic block diagram of a hardware Trojan deep-level representation model construction and feature extraction system proposed by the present invention;
[0031] Figure 2 This is a schematic block diagram of a hardware Trojan deep characterization model construction and feature extraction method proposed in the present invention. DETAILED DESCRIPTION
[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0033] In the description of the present invention, it should be understood that the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the referred device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as limiting the present invention.
[0034] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined. In addition, the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, and it can be the internal connection of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances. The present invention will be further described in detail below in conjunction with the accompanying drawings.
[0035] Reference Figure 1-2 : A method for constructing a deep-level representation model and extracting features of a hardware Trojan includes the following steps:
[0036] Collect hardware Trojan sample data and normal hardware data without Trojans, and pre-process the data, including data cleaning and normalization. Data cleaning is achieved by removing outliers and noise data. Specifically, a statistical analysis-based method is used to calculate the mean and standard deviation of the data. Data that deviates from the mean by more than a set multiple of the standard deviation is considered an outlier and is removed. When normalizing the data, the minimum-maximum normalization method is used to normalize the data to the range of [0,1]. The formula is: , where X is the original data, and are the minimum and maximum values in the data set, respectively;
[0037] A deep-level representation model of hardware Trojans is constructed based on a deep learning algorithm. The deep learning algorithm includes a convolutional neural network or a recurrent neural network. For a convolutional neural network, multiple convolutional layers, pooling layers, and fully connected layers are constructed. The convolutional layer uses convolution kernels of different sizes for feature extraction. The pooling layer uses a maximum pooling or average pooling operation to downsample the feature map to reduce the amount of calculation and prevent overfitting. For a recurrent neural network, a long short-term memory network (LSTM) or a gated recurrent unit (GRU) is used to process the sequence information of hardware data. The gating mechanism is used to control the flow and forgetting of information, so that the model can better capture the timing characteristics of the hardware during operation.
[0038] The constructed model is trained using the training data. The model parameters are adjusted to enable the model to accurately distinguish hardware Trojan data from normal hardware data. The stochastic gradient descent (SGD) algorithm is used to optimize the loss function during the training process. In addition, regularization terms such as L1 or L2 regularization are introduced to prevent overfitting. The loss function is: or ,in is the original loss function, is the regularization coefficient, w is the weight of the model;
[0039] The hardware data to be detected is input into the trained model, the feature vector of the hardware Trojan is extracted, and the presence of a Trojan in the hardware and the type of Trojan are determined based on the feature vector. When making the judgment, a threshold is set. When the output value corresponding to the feature vector exceeds the threshold, it is determined that a hardware Trojan exists. At the same time, based on the distribution of the feature vector values in different dimensions, the type of Trojan is determined using a predefined rule set, which is obtained through feature analysis of a large number of known hardware Trojan samples.
[0040] The present invention also includes feature engineering processing of the collected data to extract static features and dynamic features of the hardware. The static features include circuit structure features and logic gate types of the hardware. Specifically, the circuit structure is parsed by hardware description language (HDL), the type and connection relationship of the logic gate are identified, and they are converted into quantifiable feature representations; the dynamic features include signal timing features and power consumption features of the hardware during operation, wherein the signal timing features are acquired by a hardware performance monitoring tool at a specific sampling frequency when the hardware is running, and spectrum analysis such as Fourier transform is performed on the data to extract frequency domain features; the power consumption features are acquired by setting power sensors at key nodes of the hardware circuit to obtain power consumption data under different working states, and these features are used together with the preprocessed data as input of the model.
[0041] In the present invention, during the model training process, the cross-validation method is used to evaluate and optimize the performance of the model. The training data is divided into multiple subsets, and one of the subsets is used as the validation set and the remaining subsets are used as the training set in turn. The division of these subsets is random and uniform to ensure that each subset can represent the characteristics and distribution of the entire training data set to a certain extent. For example, if we have a training data set containing N samples, we can divide it into subsets, each containing approximately samples. And K-fold cross validation is used. The selection of K value is determined by the size and complexity of the data set. The general value range is between 5 and 10. In each round of cross validation, by adjusting the model parameters, K-fold cross validation is used here, and the selection of K value is a key link. The determination of K value needs to comprehensively consider the size and complexity of the data set. Generally speaking, for data sets of moderate size and low complexity, K value can take a smaller value, such as 5; for data sets of large size and high complexity, K value can be appropriately increased, usually in the range of 5 to 10. For example, when we choose , it means that we divide the entire training data set into 5 subsets.
[0042] In the present invention, the model is optimized by using transfer learning technology, and the model parameters trained on other related hardware data sets are migrated to the current hardware Trojan detection model. Specifically, according to the similarity between the source data set and the target data set, some layers of the pre-trained model are selectively frozen, and only some layers are fine-tuned. For convolutional neural networks, the front convolutional layers can be frozen and only the rear fully connected layers are trained. For recurrent neural networks, some LSTM or GRU units can be frozen. At the same time, during the migration process, the migrated parameters are scaled or offset adjusted according to the characteristics of the target data set to speed up the training speed of the model and improve the accuracy of the model.
[0043] In the present invention, in the feature extraction process, a dimensionality reduction algorithm such as principal component analysis (PCA) or linear discriminant analysis (LDA) is used to perform dimensionality reduction processing on the extracted feature vector. For PCA dimensionality reduction, the covariance matrix of the feature vector is calculated to solve its eigenvalue and eigenvector, and the main components are selected according to the size of the eigenvalue, and the eigenvector whose cumulative contribution rate reaches a set threshold (such as 95%) is retained; for LDA dimensionality reduction, the generalized eigenvalue problem is solved by calculating the inter-class scatter matrix and the intra-class scatter matrix, and the eigenvalue problem is projected to the optimal discriminant vector space, so as to remove redundant information and improve the effectiveness of the feature vector and the detection efficiency.
[0044] In the present invention, the detection results are visualized to present whether there is a Trojan in the hardware, the type of Trojan, and the distribution of related features in the form of charts or graphs. Specifically, data visualization tools such as Matplotlib or Seaborn are used. Different types of Trojans are marked with different colors or shapes. For the distribution of features, the distribution of feature vectors in different dimensions is displayed by drawing histograms, scatter plots or heat maps, which is convenient for users to intuitively understand and analyze the detection results.
[0045] The present invention also discloses a system for constructing a deep-level representation model of hardware Trojans and a feature extraction method, including:
[0046] Data acquisition module: used to collect hardware Trojan sample data and normal hardware data without Trojans, and transmit the data to the data preprocessing module; this module contains multiple data source interfaces, which can be connected to hardware test platforms, simulation tools, etc., receive hardware operation data through specific protocols (such as TCP / IP, SPI, etc.), and perform integrity and consistency checks on the received data;
[0047] Data preprocessing module: cleans, normalizes and performs feature engineering on the received data, and outputs the processed data to the model building and training module; this module includes an abnormal data detection submodule, which identifies abnormal values through the set data distribution range and statistical rules, and a normalization processing submodule, which performs data conversion according to the above normalization formula, and also includes a feature extraction submodule, which extracts static and dynamic features according to the above feature engineering method;
[0048] Model building and training module: Build a deep-level representation model of hardware Trojans based on deep learning algorithms, train the model using training data, and transfer the trained model to the feature extraction and detection module; this module includes a deep learning framework selection unit, which can select a suitable deep learning framework such as TensorFlow or PyTorch according to the characteristics of different hardware data, and a training parameter adjustment unit, which can set parameters such as learning rate, training rounds, and batch size through the user interface;
[0049] Feature extraction and detection module: Input the hardware data to be detected into the trained model, extract the feature vector of the hardware Trojan, and determine whether the hardware has a Trojan and the type of Trojan based on the feature vector, and transmit the detection results to the result visualization module; in this module, the feature extraction unit uses the intermediate layer information output by the model as the final feature vector, and the detection unit determines the existence and type of the Trojan through threshold comparison and rule matching algorithm, and records the confidence of the detection;
[0050] Result visualization module: Visualize the test results and present them to users in the form of charts or graphs.
[0051] In the present invention, the data acquisition module also includes a data storage unit for storing the collected hardware Trojan sample data and normal hardware data. The storage unit adopts a distributed storage architecture, can store massive data, and encrypts the data to ensure data security. It also has data backup and recovery functions for subsequent data analysis and model training.
[0052] In the present invention, the model construction and training module also includes a model evaluation unit, which uses a cross-validation method to evaluate and optimize the performance of the model, divides the training data into multiple subsets, uses one of the subsets as a validation set in turn, and the remaining subsets as training sets, and uses K-fold cross-validation. The selection of the K value is determined according to the scale and complexity of the data set, and generally ranges from 5 to 10. In each round of cross-validation, by adjusting model parameters, including learning rate, batch size, convolution kernel size, number of network layers, etc., the performance indicators on the validation set, such as accuracy, recall rate, F1 value, etc., are used to evaluate the model performance, and the parameters are continuously adjusted according to the evaluation results to improve the generalization ability of the model; the unit also includes a performance evaluation indicator calculation submodule, which can calculate and display the change curves of different indicators in real time, so that users can monitor the model performance conveniently.
[0053] In the present invention, the feature extraction and detection module also includes a dimensionality reduction processing unit. In the feature extraction process, a dimensionality reduction algorithm such as principal component analysis (PCA) or linear discriminant analysis (LDA) is used to perform dimensionality reduction processing on the extracted feature vector. For PCA dimensionality reduction, the covariance matrix of the feature vector is calculated to solve its eigenvalue and eigenvector, and the main component is selected according to the size of the eigenvalue, and the eigenvector whose cumulative contribution rate reaches the set threshold (such as 95%) is retained; for LDA dimensionality reduction, the generalized eigenvalue problem is solved by calculating the inter-class scatter matrix and the intra-class scatter matrix, and the feature vector is projected to the optimal discriminant vector space to remove redundant information and improve the effectiveness and detection efficiency of the feature vector; in the feature extraction process, when principal component analysis (PCA) is used for dimensionality reduction, the covariance matrix of the feature vector needs to be calculated first. Suppose we have a data set X containing n samples, each sample has m features, and the calculation formula of its covariance matrix C is: By solving the eigenvalues of the covariance matrix C and the corresponding v i From the eigenvector, we can get a series of eigenvalue-eigenvector pairs. These eigenvalues reflect the amount of information contained in the corresponding eigenvector.
[0054] Then, the main components are selected according to the size of the eigenvalues. Usually, we sort the eigenvectors in descending order of eigenvalues. In order to retain the eigenvectors whose cumulative contribution rate reaches the set threshold (such as 95%), we calculate the cumulative contribution rate. Let the sum of the first k eigenvalues be , the sum of all eigenvalues is , then the calculation formula of the cumulative contribution rate CR is: .
[0055] The above are only preferred specific implementation modes of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical solutions and inventive concepts of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A method for constructing a deep representation model and extracting features of hardware Trojans, characterized in that: The following steps are involved: S1. Collect hardware Trojan sample data and normal hardware data without Trojans, and pre-process the data, including data cleaning and normalization. Data cleaning is achieved by removing outliers and noise data. The mean and standard deviation of the data are calculated using a statistical analysis method. Data that deviates from the mean by more than a set multiple of the standard deviation is considered an outlier and is removed. When normalizing the data, the minimum-maximum normalization method is used to normalize the data to the range of [0,1]. The formula is: , where X is the original data, and are the minimum and maximum values in the data set, respectively; S2. Build a deep representation model of hardware Trojans based on a learning algorithm, where the learning algorithm includes a convolutional neural network or a recurrent neural network; S3. The constructed model is trained with data. The model parameters are adjusted to make the model distinguish between hardware Trojan data and normal hardware data. During the training process, the stochastic gradient descent SGD algorithm is used to optimize the loss function to prevent overfitting. Regularization terms are introduced, including L1 or L2 regularization. The loss function is: or ,in is the original loss function, is the regularization coefficient, w is the weight of the model; S4. Input the hardware data to be detected into the trained model, extract the feature vector of the hardware Trojan, determine whether there is a Trojan in the hardware and determine the type of Trojan based on the feature vector, set a threshold when judging, and determine that a hardware Trojan exists when the output value corresponding to the feature vector exceeds the threshold. At the same time, determine the type of Trojan based on the distribution of the feature vector values in different dimensions through a predefined rule set, which is obtained by analyzing the features of a large number of known hardware Trojan samples; Also includes: S5. Perform feature engineering on the collected data to extract static and dynamic features of the hardware. Static features include circuit structure features and logic gate types of the hardware. The circuit structure is analyzed using the hardware description language HDL to identify the type and connection relationship of the logic gates and convert them into quantifiable feature representations. Dynamic features include signal timing features and power consumption features of the hardware during operation. Signal timing features are collected at a predetermined sampling frequency by hardware performance monitoring tools during hardware operation, and frequency domain features are extracted by Fourier transform spectrum analysis. Power consumption features are obtained by setting power sensors at nodes of the hardware circuit to obtain power consumption data under different working states. Features and pre-processed data are used as inputs to the model. In the feature extraction process, the principal component analysis PCA or linear discriminant analysis LDA dimensionality reduction algorithm is used to reduce the dimensionality of the extracted feature vectors. For PCA dimensionality reduction, the covariance matrix of the feature vector is calculated to solve its eigenvalues and eigenvectors. For LDA dimensionality reduction, the generalized eigenvalues are solved by calculating the between-class scatter matrix and the within-class scatter matrix.
2. A method for constructing a deep representation model and extracting features of a hardware Trojan according to claim 1, characterized in that: Also includes: S6. During the model training process, the cross-validation method is used to evaluate and optimize the performance of the model. The training data is divided into multiple subsets, and one subset is used as the validation set in turn, and the remaining subsets are used as the training set. K-fold cross-validation is used. The selection of K value is determined according to the scale and complexity of the data set, and the value range is between 5 and 10. In each round of cross-validation, the model parameters are adjusted, including learning rate, batch size, convolution kernel size, and number of network layers.
3. A method for constructing a deep representation model and extracting features of a hardware Trojan according to claim 1, characterized in that: Also includes: S7. Use transfer learning technology to optimize the model, migrate the model parameters trained on the hardware dataset to the current hardware Trojan detection model, and selectively freeze some layers of the pre-trained model based on the similarity between the source dataset and the target dataset, and only fine-tune some layers. For convolutional neural networks, freeze the front convolutional layers and train the back fully connected layers. For recurrent neural networks, freeze some LSTM or GRU units. At the same time, during the migration process, scale or offset the migrated parameters according to the characteristics of the target dataset.
4. A method for constructing a deep representation model and extracting features of a hardware Trojan according to claim 1, characterized in that: include: The detection results are displayed visually, showing whether there are Trojans in the hardware, the types of Trojans, and the distribution of their characteristics in the form of charts or graphs.
5. A system for implementing the method for constructing a deep-level representation model and extracting features of a hardware Trojan according to any one of claims 1 to 4, characterized in that: include: Data acquisition module: used to collect hardware Trojan sample data and normal hardware data without Trojans, and transmit the data to the data preprocessing module; The module data source interface is connected to the hardware test platform and simulation tool, receives hardware operation data through the protocol, and performs integrity and consistency checks on the received data; Data preprocessing module: cleans, normalizes and performs feature engineering on the received data, and outputs the processed data to the model building and training module; the module includes an abnormal data detection submodule, which identifies abnormal values through the set data distribution range and statistical rules, and a normalization processing submodule, which uses the normalization formula to convert data, including a feature extraction submodule, which uses feature engineering methods to extract static and dynamic features; Model building and training module: builds a deep representation model of hardware Trojans based on learning algorithms, trains the model through training data, and transmits the trained model to the feature extraction and detection module; the module includes a learning framework selection unit, which selects the learning framework according to the characteristics of different hardware data, and a training parameter adjustment unit; Feature extraction and detection module: The hardware data to be detected is input into the trained model, the feature vector of the hardware Trojan is extracted, and the presence of the Trojan in the hardware is determined based on the feature vector and the type of Trojan is determined. The detection result is transmitted to the result visualization module; the feature extraction unit in the module uses the intermediate layer information output by the model as the final feature vector, and the detection unit determines the presence and type of the Trojan through threshold comparison and rule matching algorithm; Result visualization module: Visualize the test results and present them to users in the form of charts or graphs.
6. A hardware Trojan deep representation model construction and feature extraction system according to claim 5, characterized in that: Also includes: The data collection module includes a data storage unit: used to store the collected hardware Trojan sample data and normal hardware data. The storage unit adopts a distributed storage architecture to encrypt and store the data.
7. A hardware Trojan deep representation model construction and feature extraction system according to claim 5, characterized in that: Also includes: The model building and training module includes a model evaluation unit: the cross-validation method is used to evaluate and optimize the performance of the model. The training data is divided into multiple subsets, and one subset is used as the validation set in turn, and the remaining subsets are used as the training set. K-fold cross-validation is used, and the K value is determined according to the scale and complexity of the data set, ranging from 5 to 10. In each round of cross-validation, the model parameters are adjusted, including the learning rate, batch size, convolution kernel size, and number of network layers.
8. A hardware Trojan deep representation model construction and feature extraction system according to claim 5, characterized in that: Also includes: The feature extraction and detection module includes a dimensionality reduction processing unit: in the feature extraction process, the principal component analysis PCA or linear discriminant analysis LDA dimensionality reduction algorithm is used to reduce the dimensionality of the extracted feature vector. For PCA dimensionality reduction, the covariance matrix of the feature vector is calculated to solve its eigenvalue and eigenvector. For LDA dimensionality reduction, the generalized eigenvalue problem is solved by calculating the between-class scatter matrix and the within-class scatter matrix.
Citation Information
Patent Citations
Hardware Trojan horse detection method and device based on deep neural network and multi-dimensional features
CN116956121A
Deep learning optimization control method for coprocessor
CN117456242A