An RTL-level Hardware Trojan Detection Method Based on Subgraph Isomorphism
By using sub-graph isomorphic matching algorithm and constraint pruning technology in RTL-level hardware design, the problem of difficulty and cost of RTL-level hardware Trojan detection is solved, and fast and accurate Trojan detection and result characterization are achieved.
Patent Information
- Application Number
- CN202211424243.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-11-14
AI Technical Summary
The prior art is difficult to efficiently detect hardware Trojans in the early stages of RTL-level hardware design, resulting in increased detection difficulty and increased time and cost.
The RTL-level hardware Trojan detection method based on subgraph isomorphism is adopted. By converting RTL code into feature directed graphs, and using the subgraph isomorphic matching algorithm, constraints and pruning conditions are added to optimize detection accuracy and efficiency.
It realizes the rapid and accurate detection of hardware Trojans in the early stages of RTL-level hardware design, reducing detection costs, and providing confidence and graphical representation of detection results.
Smart Images

Figure CN115688104B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of RTL-level hardware security analysis, and particularly relates to an RTL-level hardware Trojan detection method based on subgraph isomorphism. Background Art
[0002] With the increasing prominence of global cooperation in integrated circuit design and manufacturing, each key link in chip design and manufacturing may be scattered around the world and completed by different manufacturers, and a small number of large companies control a large amount of design and manufacturing resources. This globalized model creates excellent conditions for introducing malicious Trojans during the chip design and manufacturing process. Trojans may be implanted in all links of chip design and manufacturing. Once a chip is implanted with a Trojan, the Trojan can be activated or utilized in a certain way, resulting in the failure of the target chip or the leakage of confidential information, ultimately leading to serious consequences and posing a huge security risk to the modern electronic technology industry based on integrated circuit design and manufacturing. In recent years, China has vigorously developed the integrated circuit industry. However, due to the unprecedented complexity of integrated circuit design and manufacturing, China's integrated circuit industry will still rely on foreign supply chains and technologies to a certain extent in the short term. Due to technical barriers, various chip products and tools are extremely vulnerable to being implanted with Trojans and exploited by attackers, thus endangering the healthy development of China's integrated circuit industry.
[0003] Hardware Trojans are generally introduced intentionally or unintentionally during the integrated circuit design and manufacturing process by adding, deleting, or modifying the original design. They are difficult to detect when not triggered and have a certain degree of concealment. After being triggered under certain conditions, they may cause consequences such as circuit failure, damage, performance degradation, or information leakage. They may also facilitate the implementation of internal and external malicious behaviors and further expand the harm. For example Figure 1As shown, its structure usually includes two parts: trigger logic and payload logic. The trigger logic takes certain signals of a normal circuit as inputs. When the trigger condition is met, it will activate the payload logic, and the payload logic executes the pre-set malicious logic. According to different trigger conditions, the trigger conditions of hardware Trojans generally include combinational logic trigger, sequential logic trigger, and time trigger; according to the harmful behaviors of the payload logic, the harmful behaviors of hardware Trojans generally include function change, performance degradation, data leakage, and denial of service. In the detection of hardware Trojans, as the design and manufacturing process progresses, the detection difficulty, the required time, and the cost of the Trojans also gradually increase. Therefore, Trojans should be detected as early as possible in the design stage. The Register-Transfer Level (RTL) Hardware Description Language (HDL) is widely used in fields such as chip design, Field Programmable Gate Array (FPGA)-based design, and Intellectual Property (IP) core development. Therefore, the research on RTL-level circuit Trojan detection methods is of great significance.
[0004] Static analysis techniques have been relatively well studied in the field of software Trojan detection. Therefore, hardware Trojan detection techniques based on static feature analysis are receiving increasing attention. The static features of such methods can come from multiple aspects. For example, parsing the Abstract Syntax Tree (AST) of the code, tracking the information flow, or using control flow information can often obtain relatively accurate detection results at a relatively low cost and are applicable to situations where the magnitude of the design under test and the suspected Trojan differ significantly. Combining the characteristics of RTL-level code with the concept of directed graphs in graph theory, the Trojan code and the code of the project under test can be mapped to directed graphs, thereby converting RTL-level hardware Trojan detection into subgraph isomorphism matching, and designing a matching algorithm by adding constraints and pruning conditions to improve the detection accuracy and detection efficiency. Summary of the Invention
[0005] (1) Technical Problems to be Solved
[0006] The technical problem to be solved by the present invention is: how to design an RTL-level hardware Trojan detection method that can improve the Trojan detection effect.
[0007] (2) Technical Solutions
[0008] To solve the above technical problems, the present invention provides.
[0009] (3) Beneficial Effects
[0010] The RTL-level hardware Trojan detection method based on subgraph isomorphism provided by the present invention optimizes the subgraph isomorphism algorithm by adding constraints and pruning conditions, determines the location of the Trojan quickly and accurately at a low cost, and gives the confidence level and graphical representation of the detection result. It does not rely on a reference model and is convenient for use and popularization. Description of the Drawings
[0011] Figure 1 is a schematic diagram of the structure of a general hardware Trojan provided by the prior art;
[0012] Figure 2 is a framework diagram of the RTL-level hardware Trojan detection method based on subgraph isomorphism provided by the embodiment of the present invention;
[0013] Figure 3 is a schematic diagram of the steps of converting an RTL file into a data flow graph and a control flow graph provided by the embodiment of the present invention;
[0014] Figure 4 is a flow chart of the syntax analysis and lexical analysis of an RTL file provided by the embodiment of the present invention;
[0015] Figure 5 is a flow chart of the calculation of the branch probability of the control flow graph provided by the embodiment of the present invention;
[0016] Figure 6 is a schematic diagram of the obtained data flow graph (a) and control flow graph (b) provided by the embodiment of the present invention;
[0017] Figure 7 is a schematic diagram of the obtained feature directed graph provided by the embodiment of the present invention, where (a) is the data flow graph and (b) is the control flow graph;
[0018] Figure 8 is a flow chart of the exact matching algorithm provided by the embodiment of the present invention;
[0019] Figure 9 is a flow chart of the fuzzy matching algorithm provided by the embodiment of the present invention;
[0020] Figure 10 is a schematic diagram of the node pruning algorithm provided by the embodiment of the present invention;
[0021] Figure 11 is a schematic diagram of the source of confidence level 1 provided by the embodiment of the present invention;
[0022] Figure 12 is a schematic diagram of the source of confidence level 3 provided by the embodiment of the present invention;
[0023] Figure 13 is a schematic diagram of the source of confidence level 4 provided by the embodiment of the present invention;
[0024] Figure 14 It is a schematic diagram of Method 1 in the graphical representation of the detection results provided by the embodiments of the present invention;
[0025] Figure 15 It is a schematic diagram of Method 2 in the graphical representation of the detection results provided by the embodiments of the present invention. Detailed implementation manners
[0026] To make the objectives, content, and advantages of the present invention clearer, the following further describes in detail the specific implementation manners of the present invention with reference to the accompanying drawings and embodiments.
[0027] In the embodiments of the present invention, some contents are described by taking the AES carrier circuit as an example. It should be noted that the method of the present invention is not limited to the RTL-level hardware Trojan detection of the AES carrier circuit or the security algorithm IP core, but has the detection ability for the hardware Trojans of various RTL-level carrier circuits.
[0028] As Figure 2 shown, an RTL-level hardware Trojan detection method based on subgraph isomorphism includes the following steps:
[0029] Step 1: Analyze the known hardware Trojan file and convert it into a feature directed graph:
[0030] As Figure 3 shown, the conversion process includes lexical analysis and syntax analysis of the Trojan file (including one or more RTL files) to generate an abstract syntax tree, parsing the abstract syntax tree to obtain its data flow graph and control flow graph, calculating the control flow graph probability using static branch probability, and associating the two as a feature directed graph. Specifically, the conversion process includes the following steps:
[0031] Syntax tree generation: Read the RTL file (one or more RTL files) of the hardware Trojan, and construct an abstract syntax tree through lexical analysis and syntax analysis, as Figure 4 shown;
[0032] Data flow graph and control flow graph acquisition: By performing module parsing, port, variable, and code block parsing on the abstract syntax tree, obtain expressions and conditional statements, and then parse the expressions and conditional statements to obtain the edge weights of the data flow graph according to the variable relationships, forming the data flow graph and the control flow graph;
[0033] Among them, the data flow graph represents the process of data being processed in the system, reflecting the relationships between data, and its definition is as follows:
[0034] DFG(G)=(V,E,ρ e ,h,t)
[0035] Among them, V = {v0, v1, …, v n} represents a finite set of nodes, where a node is an abstract representation of a variable in the code; E is the set of edges connecting nodes. For example, (v0, v1) ∈ E indicates that there is an assignment, conditional, or other relationship between the variables represented by nodes v0 and v1; ρ e is the edge weight function, and ρ e (e0) represents the weight of edge e0; h and t represent the first and last nodes respectively.
[0036] The control flow graph represents the change and transfer of the system state, reflecting the relationship between system running states, and its definition is as follows:
[0037] CFG(G) = (V, E, ρ e , p, h, t)
[0038] Among them, V = {v0, v1, …, v n} represents a finite set of nodes, generally the abstract states of conditional statements; E is the set of edges connecting nodes. (v0, v1) ∈ E means that node (representing the state here) v0 enters node (representing the state here) v1 after meeting certain conditions; ρ e is the edge weight function, and ρ e (e0) represents the weight of edge e0; p is the probability function, and p(v0, v1) represents the probability of converting from node (representing the state here) v0 to node (representing the state here) v1; h and t represent the first and last nodes respectively.
[0039] To make full use of the neighborhood information in the RTL code, for the data flow graph, the relationship between nodes is represented by edge weights. There are the following three types of node relationships in the data flow graph: assignment relationship, conditional relationship, and conditional assignment relationship. The corresponding three edge weights are represented by 1, 2, and 3 respectively, and can be obtained through the weight expression parsing steps.
[0040] For the control flow graph, the branch probability is represented by edge weights.
[0041] Obtaining the branch probability: By performing static branch probability calculation on the abstract syntax tree and the control flow graph, the branch probability of the control flow graph is obtained; as Figure 5 shown, ports, variables, and conditional statements are obtained from the result of parsing the syntax tree, and then the statement to be calculated is further parsed and adjusted to a tree structure.
[0042] For example, for a control flow graph containing an if…else if…else… structure, the if statement is used as the upper-level node of the tree, and the node information includes the variable information and expressions in the if judgment condition. Similarly, the else if is used as the lower-level node of the tree, and the else is used as the bottom-level node.
[0043] After that, according to the satisfiability modulo theories, the node and expression information is converted into constraint information. When the node information of this layer is added, if there is an upper layer, the upper-layer node information should be added together.
[0044] For example, when calculating the else if, the conditions of the if need to be added together. Then, check the computability of the model. If it is computable, calculate the result and record it, and then calculate the next layer. If the result is not computable, directly output the control flow graph and set the branch probability to zero.
[0045] Feature directed graph acquisition: By associating the nodes of the data flow graph with the control flow graph (as Figure 6 shown), where one node is associated with one or more control flow graphs, a feature directed graph is obtained, as Figure 7 shown. There are multiple methods of association. For example, association is performed by node name. The index of the control flow graph is included in the node name of the data flow graph, and the corresponding control flow graph can be retrieved through the data flow graph.
[0046] Step 2: Construct a Trojan horse library based on the (Trojan horse) feature directed graph and Trojan horse-related information obtained in Step 1; the Trojan horse library contains the (Trojan horse) feature directed graph and Trojan horse library configuration information, and the Trojan horse library configuration information records relevant Trojan horse information for organizing its Trojan horse library.
[0047] For example, it is organized through a JSON file:
[0048]
[0049] "AES-T100" in the first line represents the Trojan horse name; the "is_netlist" field indicates whether the Trojan horse comes from a netlist file, and "True" means the Trojan horse comes from a netlist file, and this field is a necessary field; the "netlist type" records the type of the netlist file. If "is_netlist" in the Trojan horse description is "True", this field must be specified; the "structure name" represents the name of the effective structure participating in the matching in the feature directed graph, and this field is a necessary field and must be filled in according to the order of suspicious Trojan horse, trigger logic, payload logic, and other structures; the "description" represents other additional information, and this field is a non-necessary field. The descriptions and necessities of the fields are shown in the following table:
[0050]
[0051] Through the above configuration, the system only needs to read and parse the info.json file to obtain various information about the Trojans in the Trojan library. When using this file to organize the Trojan library, there is no need to frequently add or delete the feature directed graph in the Trojan library. Instead, you only need to update the configuration file. Only when adding a new Trojan does it need to operate the feature directed graph in the library, and when rolling back, deleting, and other operations, you only need to operate on this file, thereby improving efficiency.
[0052] Building a Trojan library includes the following steps:
[0053] Generate configuration file: Generate configuration file based on the information of Trojan file, including Trojan name, port list, trigger structure name, payload structure name, other valid information, etc.;
[0054] Adding Trojan features: According to the configuration file, updating the Trojan library configuration file, and adding the feature directed graph to the Trojan library.
[0055] Step 3: The project file to be tested is processed using the method in step 1 to obtain a feature directed graph of the project file to be tested.
[0056] Step 4: The feature directed graph of the project file to be tested is matched with the feature directed graph in the Trojan library by using a subgraph isomorphic matching method to obtain a matching result; the subgraph isomorphic matching method includes:
[0057] Exact match: Figure 8 As shown, the project file to be tested is matched one by one with the Trojan features in the Trojan library (i.e., the feature directed graph generated from the Trojan RTL code) using an exact matching algorithm. If the matching time exceeds the set time, the information of this match is recorded, and then the next match is performed; if the match is successful, it is marked as a Trojan detected, and a matching result is generated; if the match fails, it is marked as not detected, and a fuzzy match is performed; the above process is executed in a loop until all input project files to be tested are matched.
[0058] Fuzzy matching: Figure 9 As shown, first read the timeout record, extend the matching time for matching, and use the fuzzy matching algorithm to match the input project file to be tested recorded in the entry of the timeout record with all the valid structures of the Trojan (such as trigger logic and load logic). If the match still times out, record the entry and mark it as unable to be processed, and continue to process the next record; if the match passes, it is marked as the Trojan is detected, and the result is generated, and then the next record is processed; if the match fails, it is marked as the Trojan is not detected, and continue to process the next entry; the above process is executed in a loop until all entries are processed.
[0059] The exact matching algorithm includes:
[0060] Node pruning algorithm based on neighborhood information: For node matching, consider the constraints on the weights of the corresponding edges of the nodes, and only use the node sets that meet the constraints as candidate node sets. For example, use the relationship between nodes and edges to restrict the nodes to be matched during matching. Among all the edges of the corresponding nodes in the subgraph, the number of edges with the same attributes must be less than that of the graph to be matched, such as Figure 10 When the above rules are not executed, the traditional subgraph isomorphism algorithm will take the left node and the right node as the node candidate set. After executing the above rules, the node pair is excluded, thereby reducing the node candidate set.
[0061] Edge pruning algorithm based on neighborhood information: For edge matching, the weight constraints of the data flow graph edges are considered, and only the edge sets that meet the constraints are used as candidate edge sets. When matching, weight 1 (assignment relationship) is considered to be different from weight 2 (conditional relationship); weight 3 (conditional assignment relationship) is equivalent to weight 1 (assignment) or weight 2 (conditional relationship), as shown in the following table. When the return value is true, the current edge pair is added to the edge candidate set. If the return value is false, the edge pair is excluded, thereby reducing the edge candidate set.
[0062]
[0063]
[0064] In the fuzzy matching algorithm, the frequent subgraphs of all valid structures of the Trojan data flow graph are accurately matched with the data flow graph of the engineering file to be tested.
[0065] The subgraph isomorphism matching method further includes:
[0066] Forced fuzzy matching: Use the fuzzy matching algorithm to match the project file to be tested with the Trojan features in the Trojan library one by one. If the matching time exceeds the set time, the information of this match is recorded and the next match is performed; if the match is successful, it is marked as the Trojan is detected and a matching result is generated; if the match fails, it is marked as the Trojan is not detected and the next match is performed; the above process is repeated until all inputs are matched.
[0067] The above exact matching and forced fuzzy matching can be directly executed manually. They are two independent methods. The prerequisite for fuzzy matching is "exact matching fails".
[0068] Step 5: In step 5, if a Trojan (all or part of a Trojan) is detected, the detailed location information of the Trojan code is obtained and the confidence is calculated; if a Trojan (all or part of a Trojan) is not detected, the matching entry information is recorded; the calculation formula of the confidence is as follows:
[0069] confidence = k1c1 + k2c2 + k3c3 + k4c4
[0070] Among them, c1, c2, c3, and c4 are four parameters for calculating the confidence level, and k i is the weight of the parameter, and its specific value is determined by the Trojan horse characteristics.
[0071] c1 represents the ratio of the number of variables with the same usage between the data flow graph variables in the matching result and the key variables (such as counters) in the Trojan horse data flow graph to the total number of key variables in the Trojan horse data flow graph. The calculation formula is as follows:
[0072] Among them, v r is the data flow graph variable in the matching result, and v s is the key variable (such as a counter) in the Trojan horse data flow graph. s(v r , v s ) is the maximum number of similar variables (number of variables with the same usage) between v r and v s . n(v s ) is the number of v s . For example, the similarity is measured by the difference in node degrees. As shown in Figure 11 , v r = 5, v s = 4. When the difference in node degrees is within the allowable range, the two nodes are considered similar variables. For example, when nodes with exactly the same node degrees are considered similar, Figure 11 s(v r , v s ) = 3.
[0073] c2 represents whether the matching result contains clock and reset logic. The calculation formula is as follows:
[0074] c2 = b(reset, clk)
[0075] Among them, b(reset, clk) is 0 or 1. If there is an operation on either the clock or the reset, it is 1; if there is no operation on both, it is 0. For example, by checking whether there is an assignment path connecting the nodes in the matching result to the clock or reset signal of the top-level module, and whether there is an additional edge with a weight of 1 (assignment relationship) among the edges of the node. If both conditions are met, return 1; otherwise, return 0. That is, first search for the nodes in the matching result that are in the clock or reset tree, and then check whether these nodes are "operated".
[0076] c3 represents the average difference in edge probabilities between the matching result and the Trojan horse control flow graph. The calculation formula is as follows:
[0077] c3 = 1 - [|ρr (b1,b2)-ρ s (b1,b2)|+…|ρ r (b n ,b n+1 )-ρ s (b n ,b n+1 )|]
[0078] Among them, ρ r (b1,b2) is the probability that node b1 flows to node b2 in the matching result control flow graph, and ρ s (b1,b2) is the probability that node b1 flows to node b2 in the Trojan control flow graph. First, retrieve and parse the data flow graph of the matching result to obtain the corresponding control flow graph information; obtain the control flow graph information of the suspicious Trojan in the same way; then match the two and calculate the probability difference of the corresponding edges of the matching control flow graphs, as Figure 12 shown. Since the matching control flow graphs may not be unique, the arithmetic mean of all results is calculated for the parameter c3 used to calculate the final confidence level.
[0079] c4 is the dependence degree between the trigger and the payload in the matching result, and the calculation formula is as follows:
[0080]
[0081] Among them, v tirger is the variable of the trigger, v payload is the variable of the payload, s(v tirgger ,v payload ) is the number of shared variables between the trigger and the payload, and n(v payload ) is the total number of payload variables, as Figure 13 shown.
[0082] If a Trojan is detected in step 5, the output result is:
[0083] The detailed location, confidence level, and Trojan description information of the detected Trojan in the engineering file to be tested, as well as the graphical display of the matching structure: The detailed location includes information such as the file name, module name, and line number of the file where the Trojan is located; the Trojan description information includes information such as the name, source, and harm of the Trojan;
[0084] Two methods are used to give the graphical representation of the result: The first method is to display the result by marking the node names of the detected module (the code in the engineering file to be tested) in the matching suspicious vulnerability data flow graph, as Figure 14 shown; The second method is to display the result by marking the matching suspicious vulnerability nodes in different colors in the data flow graph of the detected module, as Figure 15as shown
[0085] If no Trojan horse is detected, the message "No Trojan horse detected" will be output.
[0086] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. An RTL-level hardware Trojan detection method based on subgraph isomorphism, characterized in that, It includes the following steps: Step 1: Analyze the known hardware Trojan file and convert it into a Trojan feature directed graph; Step 2: Construct a Trojan library according to the Trojan feature directed graph and Trojan-related information obtained in Step 1; Step 3: Process the engineering file to be tested by the method in Step 1 to obtain the feature directed graph of the engineering file to be tested; Step 4: Use the subgraph isomorphism matching method to match the feature directed graph of the engineering file to be tested with the Trojan feature directed graphs in the Trojan library to obtain a matching result; Step 5: Output the detection result based on the matching result; The conversion process in Step 1 includes performing lexical analysis and syntax analysis on the Trojan file to generate an abstract syntax tree, parsing the abstract syntax tree to obtain its data flow graph and control flow graph, calculating the control flow graph probability using static branch probability, and associating the data flow graph and control flow graph as a Trojan feature directed graph. The Trojan file includes one or more RTL files; The subgraph isomorphism matching method in Step 4 includes: Exact matching: Use the exact matching algorithm to match the engineering file to be tested with the Trojan features in the Trojan library, that is, the Trojan feature directed graph generated from the Trojan RTL code one by one. If the matching time exceeds the set time, record the information of this match, and then perform the next match; if the match passes, mark it as detecting a Trojan and generate a matching result; if the match fails, mark it as not detecting the Trojan and perform the following fuzzy matching; Fuzzy matching: First, read the timeout record, extend the matching time for matching, and use the fuzzy matching algorithm to match the input engineering file to be tested recorded in the entries of the timeout record with all valid structures of the Trojan. If the matching still times out, record this entry and mark it as unable to be processed, and continue to process the next record; if the match passes, mark it as detecting the Trojan and generate a result, and then process the next record; if the match fails, mark it as not detecting the Trojan and continue to process the next entry; The exact matching algorithm includes: Node pruning algorithm based on neighborhood information: For node matching, consider the edge weight constraints corresponding to the nodes, and only use the node set that meets the constraints as the candidate node set. Among them, use the relationship between nodes and edges to limit the nodes to be matched during matching. Among all the edges corresponding to the nodes in the subgraph, the number of edges with the same attribute must be less than the graph to be matched; Edge pruning algorithm based on neighborhood information: For edge matching, consider the weight constraints of the data flow graph edges, and only use the edge set that meets the constraints as the candidate edge set. When matching, it is considered that the weight 1, that is, the assignment relationship, is different from the weight 2, that is, the conditional relationship, and the weight 3, that is, the conditional assignment relationship, is equivalent to the weight 1 or the weight 2. When the return value is true, add the current edge pair to the edge candidate set. If the return value is false, the edge pair is excluded, thereby reducing the edge candidate set; In the fuzzy matching algorithm, the frequent subgraphs of all valid structures of the Trojan data flow graph are exactly matched with the data flow graph of the engineering file to be tested.
2. The method according to claim 1, wherein The conversion process in Step 1 specifically includes the following steps: Syntax tree generation: Read the RTL file of the hardware Trojan and construct an abstract syntax tree through lexical analysis and syntax analysis; Acquisition of data flow graph and control flow graph: by performing module parsing, port, variable and code block parsing on the abstract syntax tree, obtaining expressions and conditional statements, then parsing the expressions and conditional statements, obtaining data flow graph edge weights according to variable relationships, and forming data flow graph and control flow graph; The data flow diagram represents the process of data processing in the system and reflects the relationship between data. It is defined as follows: DF(G) = (V, E, ρ e , h, t) Among them, V = {v0, v1,..., v n} represents a finite set of nodes, and a node is an abstract representation of a variable in the code; E is a set of edges connecting nodes. For example, (v0, v1) ∈ E indicates that there is an assignment or conditional relationship between the variables represented by nodes v0 and v1; ρ e is an edge weight function, and ρ e (e0) represents the weight of edge e0; h and t represent the first and last nodes respectively; The control flow graph represents the changes and transfers of the system state, and reflects the relationship between the system operation states. It is defined as follows: CFG(G) = (V, E, ρ e , p, h, t) where V = {v0, v1,..., v n} represents a finite set of nodes, which is the abstract state of the conditional statement; is the set of edges connecting the nodes. (v0, v1) ∈ E means that node v0 enters node v1 after meeting certain conditions; ρ e is the edge weight function, and ρ e (e0) represents the weight of edge e0; p is the probability function, and p(v0, v1) represents the probability of converting from node v0 to node v1; h and t represent the first and last nodes respectively; For data flow graphs, the relationship between nodes is represented by edge weights. Data flow graphs have the following three types of node relationships: assignment relationship, conditional relationship, and conditional assignment relationship. The corresponding three edge weights are represented by 1, 2, and 3 respectively. For control flow graphs, the branch probability is represented by edge weights; Branch probability acquisition: by performing static branch probability calculation on the abstract syntax tree and the control flow graph, the branch probability of the control flow graph is obtained; ports, variables and conditional statements are obtained from the results of syntax tree parsing, and then further parsed to adjust the statements to be calculated into a tree structure; Acquisition of Trojan horse feature directed graph: The Trojan horse feature directed graph is obtained by associating the nodes of the data flow graph with the control flow graph.
3. The method according to claim 2, wherein The Trojan library constructed in step 2 contains a Trojan feature directed graph and Trojan library configuration information, and the Trojan library configuration information records relevant Trojan information for organizing its Trojan library; When organizing through JSON files, "AES-T100" is used to represent the Trojan name; the "is_netlist" field indicates whether the Trojan comes from the netlist file, "True" indicates that the Trojan comes from the netlist file, and the field "True" is a required field; the "netlist type" is used to record the type of the netlist file. If "is_netlist" is "True" in the Trojan description, this field must be specified; "structure name" indicates the name of the valid structure participating in the matching in the feature directed graph, and the field "structurename" is a required field and is filled in in the order of suspicious Trojans, trigger logic, payload logic, and other structures.
4. The method according to claim 3, characterized in that, Building the Trojan library in step 2 specifically includes the following steps: Generate configuration file: Generate configuration file based on the information of Trojan file, including Trojan name, port list, trigger structure name, and payload structure name; Adding Trojan features: According to the configuration file, updating the Trojan library configuration file, and adding the feature directed graph to the Trojan library.
5. The method according to claim 1, wherein In step 4, the subgraph isomorphism matching method further includes: Forced fuzzy matching: Use the fuzzy matching algorithm to match the project file to be tested with the Trojan features in the Trojan library one by one. If the matching time exceeds the set time, the information of this match is recorded and the next match is performed; if the match is successful, it is marked as the Trojan is detected and a matching result is generated; if the match fails, it is marked as the Trojan is not detected and the next match is performed.
6. The method according to claim 4, wherein In step 5, if a Trojan is detected, obtain the location information of the Trojan code and calculate the confidence level; if no Trojan is detected, record the matching entry information; the formula for calculating the confidence level is as follows: confidence=k1c1+k2c2+k3c3+k4c4 Among them, c1, c2, c3, and c4 are four parameters for calculating the confidence level, and k i is the weight of the parameter, and its specific value is determined by the Trojan horse characteristics, c1 represents the ratio of the number of variables with the same usage between the data flow diagram variables in the matching result and the key variables of the Trojan data flow diagram to the total number of key variables of the Trojan data flow diagram. The calculation formula is as follows: , Among them, v r is the data flow graph variable in the matching result, v s is the key variable in the Trojan data flow graph, s(v r , v s ) is the maximum number of similar variables of v r and v s , n(v s ) is the number of v s , for example, the similarity is measured by the difference in node degrees, v r = 5, v s = 4. When the difference in node degrees is within the allowable range, the two nodes are considered similar variables; c2 represents whether the matching result includes clock and reset logic, and the calculation formula is as follows: c2=b(reset,clk) where b(reset, clk) is 0 or 1, which is 1 if there is an operation on either the clock or the reset, and 0 if there is no operation on both; c3 represents the average difference in edge probabilities between the matching result and the Trojan control flow graph, and the calculation formula is as follows: c3 = 1 - [|ρ r (b1,b2) - ρ s (b1,b2)| + … |ρ r (b n ,b n+1 ) - ρ s (b n ,b n+1 )|] Among them, ρ r (b1, b2) is the probability that node b1 flows to node b2 in the matching result control flow graph, and ρ s (b1, b2) is the probability that node b1 flows to node b2 in the Trojan control flow graph; first, retrieve and parse the data flow graph of the matching result to obtain the corresponding control flow graph information; obtain the control flow graph information of the suspicious Trojan in the same way; then match the two and calculate the probability difference of the corresponding edges of the matching control flow graphs, and calculate the arithmetic mean of all results to be used as the parameter c3 for calculating the final confidence level; c4 is the degree of dependence between the trigger and the payload in the matching result, and the calculation formula is as follows: Among them, v tirgger is the triggered variable, v payload is the variable of the payload, s(v tirgger , v payload ) is the number of variables shared between the trigger and the payload, n(v payload ) is the total number of payload variables, If a Trojan is detected in step 5, the output results are: The location information, confidence level, and Trojan description information of the detected Trojan in the project file to be tested, as well as the graphical display of the matching structure: the location information includes the file name, module name, and line number information of the file where the Trojan is located; the Trojan description information includes the name, source, and hazard information of the Trojan; If no Trojan is detected, output the information "No Trojan detected".
7. The method according to claim 6, wherein In step 5, two methods are used to give a graphical representation of the results: the first method is to display the results by annotating the node names of the code in the project file to be tested in the matching suspicious vulnerability data flow graph; the second method is to display the results by marking the matching suspicious vulnerability nodes in different colors in the data flow graph of the code in the project file to be tested.
8. The method according to claim 2, wherein The method in step 1 for associating the nodes of the data flow graph with the control flow graph includes associating them by node name. The index of the control flow graph is included in the node name of the data flow graph, and the corresponding control flow graph can be retrieved through the data flow graph.
Citation Information
Patent Citations
RTL hardware Trojan horse detection method based on a gradient lifting algorithm
CN109657461A
FPGA software suspicious circuit detection method based on subgraph isomorphism
CN113051858A