Chip Trojan horse detection method and device, program product, medium and electronic equipment

By constructing a data flow diagram and using the graph attention neural network model, the malicious features in the chip are automatically identified, which solves the limitations of chip Trojan detection in the existing technology and achieves efficient and accurate chip Trojan detection.

CN120277671AActive Publication Date: 2025-07-08北京汤谷软件技术有限公司
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510766866.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-08
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The existing chip Trojan detection technology cannot fully cover potential trigger conditions, cannot handle unknown types of malicious behavior, is difficult to obtain based on trusted benchmarks, and has high computational complexity, resulting in limited detection capabilities.

Method used

By obtaining the target code module, a data flow diagram is constructed, and a pre-trained Trojan recognition model containing the graph attention neural network is used to automatically identify malicious features in the data flow diagram, including the graph attention neural network module and the alternating fully connected layer network module, for deep learning and feature analysis.

Benefits of technology

It improves the accuracy and robustness of chip Trojan detection, reduces the workload and false alarm and missed rate of manual detection, adapts to complex chip design, and can automatically detect various types of chip Trojans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277671A_ABST
    Figure CN120277671A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of chips and detection thereof, and provides a chip Trojan horse detection method and device, a program product, a medium and electronic equipment. The method comprises the steps that a target code module used for describing at least one corresponding operation function in a to-be-detected chip is obtained, and the operation function is achieved through a plurality of transistors integrated according to set logic; constructing a data flow diagram matched with the target code module, wherein the data flow diagram is used for representing data operation logic of the target code module in the operation process; and identifying a target sub-data flow diagram with malicious features in the data flow diagram through a pre-trained Trojan identification model, wherein the target sub-data flow diagram reflects that the chip Trojan exists in the to-be-detected chip. Through the technical scheme provided by the invention, the chip Trojan horse detection capability can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the technical field of chips and their detection, and in particular, to a chip Trojan detection method, device, program product, medium and electronic equipment. Background Art

[0002] With the continuous development of integrated circuit technology, the security of chips has received increasing attention. Chip Trojans (also known as hardware Trojans) are a potential security threat that may cause serious consequences such as chip malfunctions and information leakage. During the integrated circuit design stage, effective detection of chip Trojans is an important part of ensuring chip security.

[0003] At present, chip Trojan detection technologies mainly include test pattern generation, formal verification, code analysis, machine learning and graph matching. However, these technologies have some limitations, such as the inability to cover all potential trigger conditions, the inability to handle unknown types of malicious behaviors, the need for manual analysis and the easy omission of hidden Trojans, the reliance on known Trojan libraries and high computational complexity. These limitations have greatly limited the detection capabilities of chip Trojans. Based on this, how to improve the detection capabilities of chip Trojans is a technical problem that needs to be solved urgently. Summary of the invention

[0004] The embodiments of the present application provide a chip Trojan detection method, device, computer program product, computer-readable storage medium and electronic device, which can improve the chip Trojan detection capability to a certain extent.

[0005] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by the practice of the present application.

[0006] According to a first aspect of an embodiment of the present application, a chip Trojan detection method is provided, the method comprising: obtaining a target code module, the target code module being used to describe at least one corresponding operation function in a chip to be detected, the operation function being implemented by a plurality of transistors integrated according to a set logic; constructing a data flow graph matching the target code module, the data flow graph being used to characterize the data operation logic of the target code module during operation; identifying a target sub-data flow graph with malicious features in the data flow graph through a pre-trained Trojan recognition model, the target sub-data flow graph reflecting the presence of a chip Trojan in the chip to be detected, the Trojan recognition model comprising a graph attention neural network module, the graph attention neural network module being used to focus on the sub-data flow graph with malicious features in the data flow graph.

[0007] In some embodiments of the present application, based on the foregoing solution, the target code module acquisition includes: acquiring a hardware description program corresponding to the chip to be detected; determining at least one code module in the hardware description program, where each code module is used to describe the corresponding arithmetic function in the chip to be detected; performing flattening processing on the at least one code module to integrate the at least one code module into a target code module.

[0008] In some embodiments of the present application, based on the foregoing solution, the construction of the data flow graph matching the target code module includes: parsing the target code module to generate an abstract syntax tree of the target code module, the abstract syntax tree including multiple signals; extracting the dependency relationship between each adjacent signal from the abstract syntax tree, and constructing a dependency tree corresponding to each adjacent signal based on the dependency relationship, the dependency tree including two signal nodes and the pointing relationship between the two signal nodes; merging each dependency tree to obtain an initial data flow graph matching the target code module; removing redundant signal nodes in the initial data flow graph to obtain a data flow graph matching the target code module.

[0009] In some embodiments of the present application, based on the foregoing solution, the identification of the target sub-data flow graph with malicious features in the data flow graph by the pre-trained trojan horse identification model includes: converting each signal node in the data flow graph into a one-hot encoding, and converting the pointing relationship between each signal node in the data flow graph into an adjacency matrix; inputting the one-hot encoding corresponding to each signal node and the adjacency matrix into the trojan horse identification model to identify the target sub-data flow graph with malicious features in the data flow graph by the trojan horse identification model.

[0010] In some embodiments of the present application, based on the foregoing solution, the trojan horse identification model further includes a first alternating fully connected layer network module, and the first alternating fully connected layer network module is used to locate the positions of each sub-data flow graph in the data flow graph.

[0011] In some embodiments of the present application, based on the foregoing solution, the trojan horse identification model further includes a second alternating fully connected layer network module, and the second alternating fully connected layer network module is used to calculate the probability that each sub-data flow graph in the data flow graph has malicious features.

[0012] In some embodiments of the present application, based on the foregoing solution, the second alternating fully connected layer network module is used to calculate the probability that each sub-data flow graph in the data flow graph has malicious features in each attribute.

[0013] In some embodiments of the present application, based on the foregoing solution, the feature matrix input to the first alternating fully-connected layer network module includes the feature matrix output by the graph attention neural network module, and the feature matrix input to the second alternating fully-connected layer network module includes the feature matrix output by the graph attention neural network module and the feature matrix output by the first alternating fully-connected layer network module.

[0014] In some embodiments of the present application, based on the foregoing solution, the trojan horse recognition model is trained through the following steps: obtaining sample data flow graphs corresponding to a plurality of sample code modules, where sub-data flow graphs with malicious features in the sample data flow graphs are labeled with trojan horse labels; based on the sample data flow graphs corresponding to the plurality of sample code modules, performing supervised training on a pre-constructed trojan horse recognition model framework to obtain the trojan horse recognition model.

[0015] In some embodiments of the present application, based on the foregoing solution, the data flow graph is replaced with a control flow graph.

[0016] According to a second aspect of the embodiments of the present application, there is provided a chip trojan horse detection device, the device includes: an acquisition unit, configured to acquire a target code module, where the target code module is used to describe at least one arithmetic function corresponding to a chip to be detected, and the arithmetic function is implemented by a plurality of transistors integrated according to a set logic; a construction unit, configured to construct a data flow graph matching the target code module, where the data flow graph is used to characterize the data operation logic of the target code module during operation; an identification unit, configured to identify a target sub-data flow graph with malicious features in the data flow graph through a pre-trained trojan horse recognition model, where the target sub-data flow graph reflects the existence of a chip trojan horse in the chip to be detected, and the trojan horse recognition model includes a graph attention neural network module, and the graph attention neural network module is used to focus on the sub-data flow graphs with malicious features in the data flow graph.

[0017] According to a third aspect of the embodiments of the present application, there is provided a computer program product, the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium and are adapted to be read and executed by a processor so that a computer device having the processor performs operations implemented by the method described in the first aspect above.

[0018] According to a fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, where at least one computer program instruction is stored in the computer-readable storage medium, and the at least one computer program instruction is loaded and executed by a processor to implement the operations implemented by the method described in the first aspect above.

[0019] According to a fifth aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes one or more processors and one or more memories. At least one computer program instruction is stored in the one or more memories, and the at least one computer program instruction is loaded and executed by the one or more processors to implement the operations performed by the method described in the first aspect above.

[0020] Based on the technical solution proposed in the present application, by converting the target code module into a data flow graph and analyzing it using a pre-trained trojan horse recognition model containing a graph attention neural network module, the detection ability of chip trojans can be greatly improved. Specifically, the data flow graph can comprehensively reflect the data operation logic and control flow of the target code module during operation, and can more deeply reveal the potential abnormal behavior characteristics inside the chip. By analyzing these graph structures, the trojan horse recognition model can automatically focus on and identify sub-data flow graphs with malicious characteristics, thereby effectively detecting various types and forms of chip trojans, including hidden and complex trojans. This solution does not require a trusted benchmark and a trojan horse library as a reference for trojan horse detection. It can not only improve the accuracy and robustness of chip trojan detection, but also significantly improve the automation and efficiency of chip trojan detection, and can greatly reduce the workload of manual detection and the false alarm and missed detection rates. At the same time, with the increasing complexity of chip design, this solution can continuously adapt to and generalize the detection requirements of new chip trojans, and improve the chip security guarantee ability. Therefore, this technical solution can effectively improve the detection ability of chip trojans.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings: Figure 1 A flowchart of the chip trojan detection method in the embodiments of the present application is shown; Figure 2 A data flow graph in the embodiments of the present application is shown; Figure 3 A model framework diagram of the trojan horse recognition model in the embodiments of the present application is shown; Figure 4 A block diagram of the chip trojan detection device in the embodiments of the present application is shown; Figure 5The structural schematic diagram of the electronic device in the embodiment of the present application is shown. Detailed implementation manners

[0023] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0024] In addition, the described features, structures or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a full understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring aspects of the present application.

[0025] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices. It should also be noted that in the accompanying drawings, for the sake of simplicity of the drawings, some components that do not affect the explanation of the technical solutions of the present application are adaptively omitted.

[0026] The flowcharts shown in the accompanying drawings are only exemplary illustrations, not necessarily including all the contents and operations / steps, nor necessarily executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0027] In the description of the present application, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0028] To enable those skilled in the art to better understand the present application, first, a brief description of the technical concepts and application backgrounds involved in the present application is given.

[0029] Hardware Trojan: A hardware trojan (HT), also known as a hardware trojan horse, refers to malicious circuits or modules deliberately implanted in an integrated circuit (IC). These malicious circuits are usually introduced during the chip design or manufacturing process, and their purposes may include stealing information, disrupting chip functions, causing system failures, or providing attackers with illegal access rights, etc. Hardware trojans can have various forms and characteristics. They can be very small circuit structures that are difficult to detect through conventional detection methods. These trojans may be triggered under specific conditions, such as a specific input signal sequence, time interval, or environmental conditions. Once triggered, they may perform various malicious operations, such as leaking sensitive information, tampering with data, disabling the chip, or bypassing security mechanisms.

[0030] In today's integrated circuit field, the security of chips is of utmost importance. As a potential security threat, hardware trojans pose a great risk to the normal operation of chips and information security. With the rapid development of integrated circuit technology, the complexity of chips continues to increase, making the detection of hardware trojans even more difficult.

[0031] Currently, hardware trojan detection techniques mainly include methods such as test pattern generation, formal verification, code analysis, machine learning, and graph matching. However, these methods all have certain limitations. Among them, the test pattern generation method relies on specific test vectors to trigger trojans, but often fails to cover all potential triggering conditions, resulting in some trojans being difficult to detect. The formal verification method detects trojans through predefined security attributes, but is often powerless against unknown types of malicious behaviors. The code analysis method marks suspicious signals based on coverage metrics, but requires manual analysis, which is not only inefficient but also prone to missing hidden trojans. The machine learning and graph matching methods use graph structures to represent hardware designs, but they rely on known trojan libraries and have a high computational complexity, facing many challenges in practical applications.

[0032] In addition, there are some other problems in the existing technologies. For example, many methods rely on a golden reference model and require a trusted design without trojans as a benchmark. However, in actual scenarios, it is very difficult to obtain a completely trusted trojan-free design. At the same time, the existing technologies have insufficient detection capabilities for unknown trojans and cannot effectively identify new types of trojans not included in the training library. These problems seriously affect the accuracy and efficiency of hardware trojan detection, and there is an urgent need for a more effective detection method to improve the detection ability of hardware trojans. In this context, this application proposes a hardware trojan detection scheme to improve the detection ability of hardware trojans.

[0033] The implementation details of the technical solutions of the embodiments of this application are elaborated as follows: Refer to Figure 1, which shows the flowchart of the chip Trojan detection method in the embodiments of the present application. This chip Trojan detection method can be executed by a device with computing and processing capabilities. Referring to Figure 1 As shown, this chip Trojan detection method at least includes steps 110 to 130, which are introduced in detail as follows: In step 110, a target code module is obtained. The target code module is used to describe at least one corresponding arithmetic function in the chip to be detected, and the arithmetic function is implemented by multiple transistors integrated according to a set logic.

[0034] In the present application, the target code module can be a code snippet, unit, or function used to describe at least one arithmetic function in the chip to be detected. Generally, these code modules can be written in a hardware description language (such as Verilog, VHDL, SystemVerilog, etc.), and can clearly express the structure, data flow, signal interaction, and control logic of each arithmetic function inside the chip. Each arithmetic function can be composed of several basic circuit modules such as logic gates, flip-flops, and registers, and finally can be implemented by multiple transistors integrated according to a set logic in the physical implementation. It should be noted that the set logic for integrating multiple transistors to implement the arithmetic function can be defined by the target code module.

[0035] In the present application, when detecting a chip, for example, in a chip containing the AES encryption function, the Verilog module describing the AES algorithm can be located and extracted as the target code module for subsequent data flow and control flow analysis; also, for example, in the detection of a RISC-V processor chip, the Verilog code of the arithmetic logic unit (ALU) module that implements arithmetic functions such as addition, subtraction, AND, OR, and shift can be identified and extracted as the target code module to provide a basis for subsequent graph structure construction and Trojan detection; also, for example, for a complex SoC chip, multiple function modules related to the data path (such as register files, multipliers, accumulators, etc.) can be flattened and integrated as a whole target code module for analysis to discover cross-module Trojan logic.

[0036] Specifically, in the present application, the obtaining of the target code module can be executed according to the following steps 111 to 113: Step 111, obtain the hardware description program corresponding to the chip to be detected.

[0037] Step 112, determine at least one code module in the hardware description program, where each code module is used to describe the corresponding arithmetic function in the chip to be detected.

[0038] Step 113, perform flattening processing on the at least one code module to integrate the at least one code module into a target code module.

[0039] In this application, the hardware description program can be the source code used to describe the chip circuit structure and functions, usually written in hardware description languages (HDLs) such as Verilog, VHDL, SystemVerilog, etc. This program contains the detailed definitions of each module, signal, data path, and control logic of the chip. In actual application scenarios, the complete HDL source code file of the chip to be tested can be retrieved from the chip design team or the design database; if the chip to be tested is a third-party IP core, the hardware description program can also be obtained through the source code or intermediate representation file provided by the supplier.

[0040] In this application, a code module refers to a unit, entity, process, or function in the HDL program that can independently describe a specific function. Each code module corresponds to an arithmetic function in the chip to be tested, such as an adder, multiplier, state machine, register bank, etc.

[0041] In this application, flattening refers to integrating multiple code modules with a hierarchical structure into a target code module without a hierarchical structure. Its purpose is to eliminate the hierarchical boundaries between modules and unify the signals, operations, and control logic originally scattered in different code modules into a whole, facilitating subsequent data flow and control flow analysis. The specific implementation methods can include: First, expand module instantiation and inline all the called sub-code modules into the top-level module; then, merge the signal definitions and logic descriptions of each code module, unify the naming space, and eliminate duplicates and conflicts; finally, retain the dependencies and control relationships between the original signals to ensure that the functional logic remains unchanged.

[0042] To help those skilled in the art better understand this application, the following uses a specific embodiment to illustrate the flattening process: Suppose there is the following original hierarchical Verilog code: / / Sub-module module adder(input [3:0] a, input [3:0]b, output [4:0] sum); assign sum = a + b; endmodule / / Top-level module module top(input [3:0] x, input [3:0]y, output [4:0] z); wire [4:0] temp; adder u1(.a(x),.b(y),.sum(temp)); assign z = temp; endmodule By flattening the above sub-module and top module (i.e., the two code modules), the function of the adder module is directly expanded inside the top module, and module instantiation is no longer retained, obtaining the following target code module 1: module top(input [3:0] x, input [3:0]y, output [4:0] z); wire [4:0] temp; / / Directly expand the function of adder assign temp = x + y; assign z = temp; endmodule In this application, the flattening process can break the original hierarchical structure limit, enabling the subsequent data flow graph to cover signals and logical relationships across modules and levels, thereby enhancing the comprehensiveness and accuracy of chip Trojan detection and helping to discover hidden Trojans between modules.

[0043] In this application, based on the above steps 111 to 113, the target code module can be systematically and accurately extracted and integrated from the hardware description program of the chip to be detected. This not only ensures the integrity and clear structure of the input data for subsequent Trojan detection and analysis but also provides a solid foundation for detecting complex Trojan behaviors across modules, thereby significantly enhancing the ability and engineering practical value of chip Trojan detection.

[0044] Continue to refer to Figure 1 , in step 120, construct a data flow graph that matches the target code module, and the data flow graph is used to represent the data operation logic of the target code module during operation.

[0045] In this application, a data flow graph (DFG) is a graphical representation method used to describe the data flow and processing process in a system. In hardware design, the data flow graph is mainly used to show the data transfer and operation process between various signals. Specifically, the nodes in the data flow graph represent signals or operations, and the edges represent the data flow direction. Through the data flow graph, the entire process of data from input to output after a series of processing operations can be clearly seen. It emphasizes the data flow and changes, as well as the data dependency relationships between various operations. In this application, after the input target code module undergoes steps such as preprocessing, parsing, data flow analysis, merging, and pruning, a data flow graph is generated. This data flow graph is used for subsequent hardware Trojan detection, and through the analysis and learning of the graph structure, automated detection of hardware Trojans is achieved.

[0046] In this application, it should be noted that it is also possible to only construct a control flow graph that matches the target code module and replace the data flow graph described above with the control flow graph. Specifically, a control flow graph (CFG) is a graphical representation method used to represent the control flow of a program. In a program, the control flow determines the execution order and branching situation of the program. The nodes of the control flow graph represent the basic blocks of the program, and a basic block is a group of sequentially executed statements with only one entry and one exit. The edges of the control flow graph represent the possible flow directions of program execution, reflecting control structures such as conditional judgments and loop structures in the program. Through the control flow graph, the execution path and logical structure of the program can be intuitively understood, which helps to analyze the behavior of the program, discover potential errors, and perform optimizations. For example, in software testing, test cases can be generated according to the control flow graph to cover different execution paths of the program; in program analysis and optimization, loop structures, conditional judgments, etc. can be identified by analyzing the control flow graph, and corresponding optimizations can be carried out.

[0047] Specifically, in step 120, constructing a data flow graph and / or a control flow graph that matches the target code module can be executed according to the following steps 121 to 124: Step 121, parse the target code module to generate an abstract syntax tree of the target code module, and the abstract syntax tree includes multiple signals.

[0048] Step 122, extract the dependency relationship between each adjacent signal from the abstract syntax tree, and construct a dependency tree corresponding to each adjacent signal based on the dependency relationship. The dependency tree includes two signal nodes and the pointing relationship between the two signal nodes.

[0049] Step 123, merge the various dependency trees to obtain an initial data flow graph that matches the target code module.

[0050] Step 124, remove redundant signal nodes in the initial data flow graph to obtain a data flow graph that matches the target code module.

[0051] In this application, first of all, it needs to be explained that an Abstract Syntax Tree (AST) is a tree-like data structure used to represent the structure of program code, that is, an abstract representation of source code, which transforms the syntax structure of source code into a tree structure. Each signal represents a syntax element in the source code, such as variables, functions, registers, function calls, statements, etc., and the relationships between signals represent the syntax relationships between these syntax elements. By performing syntax analysis on the source code, an abstract syntax tree can be constructed. In the abstract syntax tree, leaf signals usually represent basic syntax units, such as identifiers, constants, etc., while internal signals represent more complex syntax structures, such as expressions, statement blocks, function definitions, etc. The structure of the abstract syntax tree reflects the syntax structure of the source code, making the analysis and processing of the source code more convenient and efficient.

[0052] In this application, it also needs to be explained that a dependency tree is a tree-like structure constructed with signals as root nodes, used to represent the dependency relationships between signals. In hardware design, the generation of a signal may depend on the input or operations of other signals. By constructing a dependency tree, the dependency situation between each signal and other signals can be clearly shown. Specifically, the construction of the dependency tree is to extract the data flow information of each signal from the abstract syntax tree, and then use the signal as the root node and the signals it depends on as child nodes to form a tree-like structure. In this way, through the dependency tree, it can be intuitively understood how the generation of a signal depends on other signals, as well as the transmission and influence relationships between signals, which helps to analyze the structure and behavior of the hardware design and discover potential security issues.

[0053] In this application, a suitable programming language parser (such as tools based on Pyverilog, ANTLR, Clang, Babel, etc.) can be used to perform lexical analysis and syntax analysis on the target code module. Then, according to the parsing results, an abstract syntax tree of the target code module is generated, and multiple signals (variables, registers, function calls, etc.) are identified and marked in the abstract syntax tree. These signals will serve as the basis for subsequent dependency relationship analysis.

[0054] After marking multiple signals in the abstract syntax tree, the abstract syntax tree can be traversed to identify the dependencies between adjacent signals, which may appear in the same expression, statement, or basic block. For example, for the statement SUM = A + B, it indicates that SUM depends on A and B. Dependencies are mainly divided into data dependencies and control dependencies. Among them, data dependencies are reflected in the influence of the right value on the left value in assignment statements, while control dependencies involve the influence of conditional judgments on the execution of subsequent statements.

[0055] For each pair of adjacent signals, a dependency tree can be constructed. The nodes of the tree represent signals, and the edges represent dependencies. The direction of the edges indicates the direction of the dependency (e.g., from the dependent signal to the dependent signal). The dependency tree usually presents a directed tree structure, reflecting the chain of signal transmission and influence. For example, the arrow from A to SUM indicates that data flows from A to SUM. When constructing the dependency tree, it is necessary to distinguish different types of dependencies and ensure the accuracy and integrity of the dependencies when dealing with complex dependencies.

[0056] After constructing the dependency tree, the local dependency trees can be integrated into a global dependency graph, that is, the initial data flow graph matching the target code module, to reflect the signal flow and / or control flow of the entire target code module. In specific operations, each dependency tree is merged according to the shared signal nodes to form a unified graph structure. During the merging process, duplicate nodes need to be avoided, and the node representations of the same signal need to be unified. The generated graph is the initial data flow graph, where the initial data flow graph reflects the data transmission and dependencies between signals. During the merging process, the coherence and correctness of the graph structure must be maintained.

[0057] In this application, a redundant signal node refers to a signal node that has no actual arithmetic or control significance in the initial data flow graph or makes no contribution to the recognition result, such as constant signals, unused signals, and duplicate nodes. During the process of removing redundant signal nodes, isolated nodes, useless nodes, and duplicate nodes can be identified through graph traversal and analysis. These redundant signal nodes are filtered and deleted according to preset rules (such as node degree, signal type, whether to participate in the critical path, etc.), and the edge connection relationships are adjusted to ensure the coherence and correctness of the graph structure. The advantage of removing redundant signal nodes is that it can simplify the graph structure, reduce the computational complexity, and thus improve the processing efficiency and accuracy of the subsequent Trojan recognition model. For example, in an adder graph, if there are unused signal nodes or constant nodes (such as a fixed value "0" signal), removing them can obtain a more concise final data flow graph.

[0058] In this application, to make it easier for those skilled in the art to understand, the following presents a data flow graph corresponding to a target code module in conjunction with an embodiment of a target code module.

[0059] In this application, specifically refer to the following target code module 2: “module top( input rst, / / Reset signal, active high input clock, / / Clock signal, active high input io_sfence, / / Input decision condition signal input [32:0] flag, / / 128-bit wide input signal state output reg top.v_17 / / Output signal top.v_17 ); always @(posedge clock or posedge rst) begin if (io_sfence==0) begin top.v_17 <= 1'h0; / / When io_sfence is non-negative, assign the output signal top.v_17 to 0 end else if (6'h11 == io_sfence) begin top.v_17 <= 1'h1; / / When io_sfence is 11, assign the output signal top.v_17 to 1 end else begin top.v_17 <= 1'h2; / / Otherwise, assign the output signal top.v_17 to 2 end end endmodule”。

[0060] Based on the target code module 2 exemplified above, it can be understood that the main function of the module top defined by this Verilog code module is as follows: to determine the value of the output signal top.v_17 based on the input reset signal rst, clock signal clock, and status signal io_sfence. When reset and clock are at high level and io_sfence is 0, top.v_17 is set to 0; when rst and clock are at high level and io_sfence is 11, top.v_17 is set to 1; when rst and clock are at high level and io_sfence is neither 0 nor 11, top.v_17 is set to 2. In this module, the always block is used to perform corresponding operations according to the changes of input signals, and the sensitive signal list of the always block is io_sfence1.

[0061] For the above target code module 2, construct a corresponding data flow diagram. For details, please refer to Figure 2 , which shows the data flow diagram in the embodiment of the present application.

[0062] For the data flow diagram as Figure 2 shown, the description is as follows: First, the top-level node "top.v_17" is the starting point of the data flow, representing a certain trigger signal or master control signal.

[0063] Second, the "Branch1" node represents the judgment of a certain condition, determining which subsequent branch the data will flow to. Among them, "COND" represents the judgment condition branch, "TRUE" represents that when the judgment condition branch is true, the data flow goes to the TRUE branch; "FALSE" represents that when the judgment condition branch is false, the data flow goes to the FALSE branch.

[0064] Third, the "Eq1" node represents an equality judgment; the "top_io_sfence" node represents the judgment signal; the "1'h0" node represents the number 0; this branch is used to judge whether top_io_sfence is equal to 0. If so, top.v_17 is assigned the value 0.

[0065] Fourth, the "Branch2" node represents that when the condition in the "Branch1" node is false, the data flow goes to the branch under the "Branch2" node.

[0066] Fifth, the "Eq2" node represents an equality judgment, the "top_io_sfence" node represents the status signal; the "6'h11" node represents the number 6'h11; the "1'h1" node represents the number 1; this branch is used to judge whether top_io_sfence is equal to 6'h11. If so, top.v_17 is assigned the value 1.

[0067] Sixth, the "Branch3" node indicates that when the condition in the "Branch2" node is false, the data flow goes to the branch under the "Branch3" node.

[0068] Seventh, the "1'h2" node represents the number 2; this branch is used to assign the value 2 to top.v_17.

[0069] It can be seen that as Figure 2 shown in the data flow diagram, after a signal (top.v_17) enters, it sequentially passes through a series of conditional judgments (branches), and equality judgments are made through Eq nodes on each branch. The judgment results of each branch will determine how the data continues to flow, and may ultimately affect a certain output or control signal.

[0070] In step 130, a pre-trained Trojan recognition model is used to recognize the target sub-data flow diagram with malicious features in the data flow diagram, and the target sub-data flow diagram reflects the existence of a chip Trojan in the chip to be detected.

[0071] In this application, a malicious feature refers to a program feature that can reflect Trojan behavior. Further, malicious features have different attributes, such as malicious features of abnormal data dependence, malicious features of unconventional control flow jumps, malicious feature columns of suspicious API call sequences, malicious features of hidden communication, etc.

[0072] In this application, a sub-data flow diagram can be a "subset" or "local area" in the overall data flow diagram. It contains a part of the nodes and edges in the overall data flow diagram and represents the local data flow relationship. For example, the nodes and edges related to the assignment and use of a certain variable form a sub-data flow diagram.

[0073] In this application, if the data flow diagram is replaced by a control flow diagram, then the sub-control flow diagram is part of the nodes and edges in the entire control flow diagram and represents the structure of a certain local control flow in the program. For example, the control flow within a certain function body, or a certain branch structure, loop structure, etc. is a sub-control flow diagram.

[0074] In this application, the Trojan recognition model may include a graph attention neural network module, and the graph attention neural network module is used to focus on the sub-data flow diagram with malicious features in the data flow diagram. The Graph Attention Network (GAT) is a deep learning model that can process graph-structured data. Its core idea is to dynamically allocate weights to different node neighbors through a self-attention mechanism, so as to focus on the key features related to the Trojan in the data flow diagram.

[0075] Specifically, first, the data flow graph and the characteristics of its nodes and edges can be used as inputs to construct graph-structured data. Then, the graph attention layer calculates the attention scores of the adjacent nodes of each node, where the nodes of the subgraph with malicious characteristics (i.e., the sub-data flow graph) are given higher weights in the attention mechanism, thus having a greater influence in the feature aggregation process. Next, each layer of the graph attention network updates the representation of the current node by weighted aggregation of the features of neighboring nodes. After multiple layers are stacked, the model can capture more complex structural and behavioral patterns. In addition, the model automatically identifies and focuses on the sub-data flow graphs highly relevant to malicious behaviors through the attention mechanism, thereby strengthening the features of key regions. Finally, the output graph-level or node-level features are used for Trojan horse identification. By introducing the graph attention neural network module, the key sub-structures in the data flow graph that may carry malicious behavior features can be automatically learned and focused on, thus significantly improving the detection accuracy of complex or hidden Trojan horses and enhancing the accuracy and robustness of chip Trojan horse identification. In addition, the graph attention neural network module can adapt to various graph structure inputs, can process not only data flow graphs but also control flow graphs, and even can process the fusion graph of the two, showing stronger generalization ability.

[0076] In this application, the Trojan horse identification model may further include a first alternating fully connected layer network module, and the first alternating fully connected layer network module is used to locate the positions of each sub-data flow graph in the data flow graph. That is, the first alternating fully connected layer network module can accurately locate the specific positions of these subgraphs in the overall data flow graph. For example, Figure 2 the subgraph 201 shown, and also for example, Figure 2 the subgraph 202 shown. By locating the positions of the sub-data flow graphs in the data flow graph through the first alternating fully connected layer network module, basic data support can be provided for subsequent malicious behavior analysis, evidence collection, and program repair.

[0077] In this application, the Trojan horse identification model further includes a second alternating fully connected layer network module, and the second alternating fully connected layer network module is used to calculate the probability that each sub-data flow graph in the data flow graph has malicious characteristics.

[0078] In this application, the feature matrix input to the first alternating fully connected layer network module may include the feature matrix output by the graph attention neural network module, and the feature matrix input to the second alternating fully connected layer network module may include the feature matrix output by the graph attention neural network module and the feature matrix output by the first alternating fully connected layer network module.

[0079] In this way, the second alternating fully-connected layer network module can combine the feature information output by the graph attention neural network module and the sub-graph position information output by the first alternating fully-connected layer network module to accurately analyze whether there are malicious features in each sub-graph of the data flow graph. Specifically, the second alternating fully-connected layer network module can classify each sub-data flow graph in the data flow graph and output the probability value that each sub-graph has malicious features. This probability value can reflect the confidence of the model in the trojan behavior of this sub-structure. Through this probability value, it can be determined whether the corresponding sub-graph is a trojan with malicious features.

[0080] Furthermore, the second alternating fully-connected layer network module can also be used to calculate the probability that each sub-data flow graph in the data flow graph has malicious features in each attribute. For example, for a certain sub-data flow graph in the data flow graph, the second alternating fully-connected layer network module can calculate the probability values that it belongs to malicious features of different attributes such as abnormal data dependence, unconventional control flow jump, suspicious API call sequence, and hidden communication. In this way, a more refined trojan detection result can be output.

[0081] To enable those skilled in the art to better understand the trojan recognition model, the following will be combined with Figure 3 a specific embodiment for illustration.

[0082] See Figure 3 , which shows the model framework diagram of the trojan recognition model in the embodiment of the present application.

[0083] As Figure 3 shown, first, each signal node in the data flow graph is characterized by one-hot encoding, and at the same time, the adjacency matrix is used to describe the connection relationship between signal nodes, so as to completely express the structural information of the graph. Subsequently, these node features and the adjacency matrix are input into the graph attention neural network module. Since the graph attention neural network module introduces an attention mechanism, it can adaptively learn the high-dimensional expression of nodes according to the importance between nodes and its neighbors, further improving the modeling ability of the graph structure and semantics. The node features output by the graph attention neural network module are then respectively input into two alternating fully-connected layer network modules. One branch is used to locate and predict the specific position information of the sub-data flow graph, and the other branch is used to judge the probability that these sub-graphs contain malicious features, realizing the recognition of malicious behaviors. Through this structured multi-task modeling method, not only can the complex structure and potential anomalies in the data flow graph be accurately captured and analyzed, but also the sub-graph position and the malicious probability results of each sub-graph can be output simultaneously, so as to identify the target sub-data flow graph with malicious features in the data flow graph, greatly improving the accuracy and practicality of detection.

[0084] In addition, the alternating fully-connected layer can be used to perform further non-linear transformation and deep information fusion on the high-dimensional node features output by the graph attention neural network module. By stacking multiple layers to extract more abstract features, combining the features extracted from the previous layer, and separately optimizing the position information prediction and malicious feature discrimination tasks of the sub-data flow graph in the multi-task branch, it can effectively prevent the loss of important feature information during network transmission and improve the accuracy of feature information extraction. At the same time, it can also prevent overfitting and improve the generalization ability of the model. Therefore, by continuously reorganizing and screening features, the discrimination ability of the model is enhanced, thereby providing richer and more efficient feature support for malicious feature detection and improving the overall detection performance.

[0085] In this application, identifying the target sub-data flow graph with malicious features in the data flow graph by using the pre-trained trojan recognition model can be performed according to the following steps 131 to 132: Step 131: Convert each signal node in the data flow graph into a one-hot encoding, and convert the pointing relationship between each signal node in the data flow graph into an adjacency matrix.

[0086] Step 132: Input the one-hot encoding corresponding to each signal node and the adjacency matrix into the trojan recognition model, so that the trojan recognition model can identify the target sub-data flow graph with malicious features in the data flow graph.

[0087] In this application, one-hot encoding is a commonly used method for vectorizing discrete features. Suppose there are N different types of signal nodes (such as assignment, addition, conditional judgment, etc.), and each type is assigned a unique number. For each signal node, it is represented by an N-dimensional vector, where only the position corresponding to the corresponding type is 1, and the rest are 0. For example, for a signal node with the node type "addition", if "addition" is in the 3rd position, its one-hot encoding is [0, 0, 1, 0,..., 0]. In this way, the originally symbolic and categorical node information is converted into numerical input that can be directly processed by models such as neural networks.

[0088] In this application, the pointing relationship between each signal node in the data flow graph is the direction of data flow or control flow transfer between the edges of the signal nodes (that is, who points to whom).

[0089] The edges between signal nodes in the data flow graph represent the data flow or control flow transfer between signals, that is, the pointing relationship. If signal node m points to signal node n, the structure of the graph is represented by a two-dimensional matrix A, where A[m][n] = 1, and vice versa it is 0. The following adjacency matrix is as follows Figure 2The adjacency matrix corresponding to the shown data flow diagram. For example, the signal node "Top.v_17" points to the signal node "Branch1".

[0090] In this application, the adjacency matrix can be used as the standard input format for models such as graph neural networks to process graph structures, that is, to convert the original graph structure information into a matrix form readable by the model, which is convenient for subsequent deep learning processing of graph structures.

[0091] In this application, the input of the Trojan horse recognition model includes the one-hot encoded vector of each signal node and the adjacency matrix of the pointing relationships between signal nodes. During the model recognition process, the feature information of each node and its neighbor nodes is aggregated through a multi-layer neural network, and high-order structural features can be extracted. The model can detect these patterns automatically in the input data flow diagram by learning graph structure patterns representing "malicious features", such as specific data flow patterns or abnormal control flow structures. The model outputs the "malicious probability" or classification label of each subgraph or node / edge, and can further output the specific index and node set of the sub-data flow diagram determined to have malicious features. Combining the model output results with the original adjacency matrix, the specific location of the malicious subgraph in the overall data flow diagram can be quickly located, providing data support for subsequent traceability analysis.

[0092] In this way, based on the technical solutions of the above steps 131 to 132, the graph structure information and node attributes can be fully utilized to achieve more accurate malicious behavior recognition, and support automated and batch Trojan horse detection and positioning.

[0093] In this application, the Trojan horse recognition model can be trained through the following steps 101 to 102: Step 101, obtain sample data flow diagrams corresponding to multiple sample code modules, and the sub-data flow diagrams with malicious features in the sample data flow diagrams are labeled with Trojan horse labels.

[0094] Step 102, based on the sample data flow diagrams corresponding to the multiple sample code modules, perform supervised training on a pre-constructed Trojan horse recognition model framework to obtain the Trojan horse recognition model.

[0095] In this application, a sample code module refers to a batch of data samples for training. Each sample is usually a piece of source code, a function, a module, or a complete program. The samples should include normal code and malicious code containing Trojan horses to ensure that the pre-constructed Trojan horse recognition model framework of the model can learn the characteristics of malicious code.

[0096] In this application, the process of obtaining the sample data flow diagrams corresponding to multiple sample code modules can refer to the process of constructing the data flow diagram matching the target code module in the above steps 121 to 124, which will not be elaborated herein.

[0097] In this application, the sub-data flow diagrams with malicious features in the sample data flow diagrams are labeled with Trojan tags. Among them, the sub-data flow diagrams labeled with Trojan tags are known Trojan implementation methods, suspicious data flow patterns, or malicious behavior segments determined by experts after analysis. The labeling method can be manual labeling, rule-based automatic labeling, or assisted labeling combined with static / dynamic analysis tools.

[0098] In this application, the pre-constructed Trojan recognition model framework can also adopt model types such as graph neural network (GNN) and graph convolutional network (GCN), which can process node features and graph structures simultaneously. The input of the model includes the node feature matrix (such as one-hot encoding) and the adjacency matrix, and the output is the classification result of each subgraph or node / edge, such as whether it is a malicious subgraph. In the supervised training process, the prepared labeled data is used for training, and loss functions such as cross-entropy loss can be used for backpropagation to optimize the model parameters. The training objective is to let the model learn to distinguish normal subgraphs and malicious subgraphs, so as to automatically identify Trojan features.

[0099] In terms of training details, means such as mini-batch, data augmentation, and regularization can be used to enable the model to learn the features of subgraphs with malicious features, so as to have the ability to identify unknown types of malicious subgraphs and improve the generalization ability of the model. At the same time, the training samples can also be expanded by means of subgraph sampling and sliding windows. Finally, the trained model can input new data flow diagrams and automatically identify the malicious sub-data flow diagrams therein, thus realizing Trojan detection.

[0100] It can be seen that in this application, through the labeled real samples, the model can learn complex Trojan implementation methods and the structural features of hidden malicious behaviors, improving the accuracy and automation level of Trojan detection.

[0101] Based on the technical solution proposed in this application, by converting the target code module into a data flow graph and using a pre-trained Trojan horse recognition model containing a graph attention neural network module for analysis, the detection ability of chip Trojan horses can be greatly improved. Specifically, the data flow graph can comprehensively reflect the data operation logic and control flow of the target code module during operation, and can more deeply reveal the potential abnormal behavior characteristics inside the chip. Through the analysis of these graph structures, the Trojan horse recognition model can automatically focus on and identify sub-data flow graphs with malicious characteristics, thereby effectively detecting various types and forms of chip Trojan horses, including hidden and complex Trojan horses. This solution does not require a trusted benchmark and a Trojan horse library as a reference for Trojan horse detection. It can not only improve the accuracy and robustness of chip Trojan horse detection, but also significantly improve the automation and efficiency of chip Trojan horse detection, and can greatly reduce the workload of manual detection and the false alarm and missed detection rates. At the same time, as chip design becomes increasingly complex, this solution can continuously adapt to and generalize the detection requirements of new chip Trojan horses, and improve the chip security guarantee ability. Therefore, this technical solution can effectively improve the detection ability of chip Trojan horses.

[0102] The technical solution proposed in this application realizes the detection of chip Trojan horses based on the data flow graph, and can be widely applied to the field of integrated circuit design. In particular, it has important application prospects in the following scenarios: 1. Chip design stage: During the chip design process, perform Trojan horse detection on the designed code to timely discover and eliminate potential security hazards, and ensure the security and reliability of the chip.

[0103] 2. Third-party IP core verification: For the IP cores provided by third parties, use this method for Trojan horse detection to verify their security and avoid introducing potential risks into the chip.

[0104] 3. Quality control during chip production: During the chip production process, conduct spot checks on the chips and use this method to detect whether there are Trojan horses in the chips to ensure the quality and security of the chips.

[0105] 4. Chip security assessment: Conduct a security assessment on the chips that have been put into use, use this method to detect whether there are Trojan horses in the chips, timely discover and handle security problems, and ensure the normal operation and information security of the chips.

[0106] The following introduces the device embodiments of this application, which can be used to execute the chip Trojan horse detection method in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the embodiments of the chip Trojan horse detection method above in this application.

[0107] See Figure 4 which shows the block diagram of the chip Trojan horse detection device in the embodiments of this application.

[0108] As shown Figure 4 in FIG. 1, the chip trojan detection device 400 according to an embodiment of the present application includes: an acquisition unit 401, a construction unit 402, and an identification unit 403.

[0109] Among them, the acquisition unit 401 is configured to acquire a target code module, where the target code module is used to describe at least one corresponding arithmetic function in the chip to be detected, and the arithmetic function is implemented by a plurality of transistors integrated according to a set logic; the construction unit 402 is configured to construct a data flow graph matching the target code module, where the data flow graph is used to represent the data arithmetic logic during the operation of the target code module; the identification unit 403 is configured to identify a target sub-data flow graph with malicious features in the data flow graph through a pre-trained trojan identification model, where the target sub-data flow graph reflects that there is a chip trojan in the chip to be detected, and the trojan identification model includes a graph attention neural network module, and the graph attention neural network module is used to focus on the sub-data flow graph with malicious features in the data flow graph.

[0110] In some embodiments of the present application, based on the foregoing solution, the acquisition unit 401 is configured to: acquire a hardware description program corresponding to the chip to be detected; determine at least one code module in the hardware description program, where each code module is used to describe the corresponding arithmetic function in the chip to be detected; perform flattening processing on the at least one code module to integrate the at least one code module into a target code module.

[0111] In some embodiments of the present application, based on the foregoing solution, the construction unit 402 is configured to: parse the target code module to generate an abstract syntax tree of the target code module, where the abstract syntax tree includes a plurality of signals; extract the dependency relationship between each adjacent signal from the abstract syntax tree, and construct a dependency tree corresponding to each adjacent signal based on the dependency relationship, where the dependency tree includes two signal nodes and the pointing relationship between the two signal nodes; merge the respective dependency trees to obtain an initial data flow graph matching the target code module; remove redundant signal nodes in the initial data flow graph to obtain a data flow graph matching the target code module.

[0112] In some embodiments of the present application, based on the foregoing solution, the identification unit 403 is configured to: convert each signal node in the data flow graph into a one-hot encoding, and convert the pointing relationship between each signal node in the data flow graph into an adjacency matrix; input the one-hot encoding corresponding to each signal node and the adjacency matrix into the trojan identification model, so that the trojan identification model identifies a target sub-data flow graph with malicious features in the data flow graph.

[0113] In some embodiments of the present application, based on the foregoing solution, the Trojan horse recognition model further includes a first alternating fully connected layer network module, and the first alternating fully connected layer network module is used to locate the positions of each sub-data flow graph in the data flow graph.

[0114] In some embodiments of the present application, based on the foregoing solution, the Trojan horse recognition model further includes a second alternating fully connected layer network module, and the second alternating fully connected layer network module is used to calculate the probability that each sub-data flow graph in the data flow graph has malicious features.

[0115] In some embodiments of the present application, based on the foregoing solution, the second alternating fully connected layer network module is used to calculate the probability that each sub-data flow graph in the data flow graph has malicious features on each attribute.

[0116] In some embodiments of the present application, based on the foregoing solution, the feature matrix input to the first alternating fully connected layer network module includes the feature matrix output by the graph attention neural network module, and the feature matrix input to the second alternating fully connected layer network module includes the feature matrix output by the graph attention neural network module and the feature matrix output by the first alternating fully connected layer network module.

[0117] In some embodiments of the present application, based on the foregoing solution, the device further includes: a training unit, configured to obtain sample data flow graphs corresponding to a plurality of sample code modules, and sub-data flow graphs with malicious features in the sample data flow graphs are labeled with Trojan horse labels; based on the sample data flow graphs corresponding to the plurality of sample code modules, perform supervised training on a pre-constructed Trojan horse recognition model framework to obtain the Trojan horse recognition model.

[0118] In some embodiments of the present application, based on the foregoing solution, the data flow graph is replaced with a control flow graph.

[0119] Based on the same inventive concept, an embodiment of the present application provides a computer program product, where the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium and are adapted to be read and executed by a processor, so that a computer device having the processor executes to implement the operations performed by the chip Trojan horse detection method as described above.

[0120] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, where at least one computer program instruction is stored in the computer-readable storage medium, and the at least one computer program instruction is loaded and executed by a processor to implement the operations performed by the chip Trojan horse detection method as described above.

[0121] Based on the same inventive concept, an embodiment of the present application further provides an electronic device, refer to Figure 5, which shows a schematic structural diagram of an electronic device in an embodiment of the present application. The electronic device includes one or more memories 504, one or more processors 502, and at least one computer program (computer program instructions) stored in the memory 504 and executable on the processor 502. When the processor 502 executes the computer program, the chip trojan detection method described above is implemented.

[0122] Among them, in Figure 5 , the bus architecture (represented by bus 500), bus 500 may include any number of interconnected buses and bridges. Bus 500 links together various circuits including one or more processors represented by processor 502 and a memory represented by memory 504. Bus 500 may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and thus will not be further described herein. The bus interface 505 provides an interface between the bus 500 and the receiver 501 and the transmitter 503. The receiver 501 and the transmitter 503 may be the same element, i.e., a transceiver, providing a unit for communicating with various other devices over a transmission medium. The processor 502 is responsible for managing the bus 500 and general processing, while the memory 504 may be used to store data used by the processor 502 when performing operations.

[0123] The functions described herein may be implemented in hardware, software executed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored on a computer-readable medium or transmitted via a computer-readable medium as one or more instructions or codes. Other examples and implementations are within the scope and spirit of the present application and the appended claims. For example, due to the nature of software, the functions described above may be implemented using software executed by a processor, hardware, firmware, hardwiring, or any combination of these. In addition, each functional unit may be integrated in one processing unit, may exist separately physically as individual units, or two or more units may be integrated in one unit.

[0124] In several embodiments provided in the present application, it should be understood that the disclosed technical content may be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units may be a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other may be through some interfaces, and the indirect coupling or communication connection of units or modules may be in an electrical or other form.

[0125] The unit described as a separating component may or may not be physically separated. The component as a control device may or may not be a physical unit, that is, it may be located in one place or may be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that makes a contribution to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store computer program instructions.

[0127] The foregoing are only the embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the scope of the claims of this application.

Claims

1. A method for detecting chip trojans, characterized in that The method includes: Obtain a target code module, where the target code module is used to describe at least one corresponding arithmetic function in the chip to be detected, and the arithmetic function is implemented by multiple transistors integrated according to a set logic; Construct a data flow graph matching the target code module, where the data flow graph is used to characterize the data arithmetic logic during the operation of the target code module; Identify a target sub-data flow graph with malicious features in the data flow graph through a pre-trained trojan recognition model. The target sub-data flow graph reflects the existence of a chip trojan in the chip to be detected. The trojan recognition model includes a graph attention neural network module, and the graph attention neural network module is used to focus on the sub-data flow graph with malicious features in the data flow graph.

2. The method according to claim 1, characterized in that, The obtaining of the target code module includes: Obtain a hardware description program corresponding to the chip to be detected; Determine at least one code module in the hardware description program, where each code module is used to describe the corresponding arithmetic function in the chip to be detected; Perform flattening processing on the at least one code module to integrate the at least one code module into a target code module.

3. The method according to claim 1, wherein The constructing of the data flow graph matching the target code module includes: Parse the target code module to generate an abstract syntax tree of the target code module, and the abstract syntax tree includes multiple signals; Extract the dependency relationship between each adjacent signal from the abstract syntax tree, and construct a dependency tree corresponding to each adjacent signal based on the dependency relationship. The dependency tree includes two signal nodes and the pointing relationship between the two signal nodes; Merge each dependency tree to obtain an initial data flow graph matching the target code module; Remove redundant signal nodes in the initial data flow graph to obtain a data flow graph matching the target code module.

4. The method according to claim 1, characterized in that, The identifying of the target sub-data flow graph with malicious features in the data flow graph through a pre-trained trojan recognition model includes: Convert each signal node in the data flow graph into a one-hot encoding, and convert the pointing relationship between each signal node in the data flow graph into an adjacency matrix; Input the one-hot encoding corresponding to each signal node and the adjacency matrix into the trojan recognition model, so that the trojan recognition model identifies the target sub-data flow graph with malicious features in the data flow graph.

5. The method according to claim 1, wherein The trojan recognition model further includes a first alternating fully-connected layer network module, and the first alternating fully-connected layer network module is used to locate the position of each sub-data flow graph in the data flow graph.

6. The method according to claim 5, wherein The trojan recognition model further includes a second alternating fully-connected layer network module, and the second alternating fully-connected layer network module is used to calculate the probability that each sub-data flow graph in the data flow graph has malicious features.

7. The method according to claim 6, wherein The second alternating fully-connected layer network module is used to calculate the probability that each sub-data flow graph in the data flow graph has malicious features in each attribute.

8. The method according to claim 6, characterized in that, The feature matrix input to the first alternating fully-connected layer network module includes the feature matrix output by the graph attention neural network module, and the feature matrix input to the second alternating fully-connected layer network module includes the feature matrix output by the graph attention neural network module and the feature matrix output by the first alternating fully-connected layer network module.

9. The method according to any one of claims 1 to 8, characterized in that The trojan horse recognition model is trained through the following steps: Obtain sample data flow graphs corresponding to multiple sample code modules, and the sub-data flow graphs with malicious features in the sample data flow graphs are labeled with trojan horse tags; Based on the sample data flow graphs corresponding to the multiple sample code modules, perform supervised training on a pre-constructed trojan horse recognition model framework to obtain the trojan horse recognition model.

10. The method according to claim 1, wherein Replace the data flow graph with a control flow graph.

11. A chip trojan detection device, characterized in that, The device includes: An acquisition unit, configured to acquire a target code module, where the target code module is used to describe at least one arithmetic function corresponding to a chip to be detected, and the arithmetic function is implemented by a plurality of transistors integrated according to a set logic; A construction unit, configured to construct a data flow graph matching the target code module, where the data flow graph is used to represent the data arithmetic logic during the operation of the target code module; An identification unit, configured to identify a target sub-data flow graph with malicious features in the data flow graph through a pre-trained trojan horse recognition model, where the target sub-data flow graph reflects the existence of a chip trojan horse in the chip to be detected, and the trojan horse recognition model includes a graph attention neural network module, and the graph attention neural network module is used to focus on the sub-data flow graphs with malicious features in the data flow graph.

12. A computer program product, characterized in that, The computer program product includes computer instructions, which are stored in a computer-readable storage medium and are adapted to be read and executed by a processor, so that a computer device having the processor executes the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, At least one program code is stored in the computer-readable storage medium, and the at least one program code is loaded and executed by a processor to implement the operations performed by the method according to any one of claims 1 to 10.

14. An electronic device, characterized in that, The electronic device includes one or more processors and one or more memories, and at least one program code is stored in the one or more memories, and the at least one program code is loaded and executed by the one or more processors to implement the method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Hardware Trojan horse detection method and system based on bidirectional graph convolutional neural network

    CN114065307A

  • Intelligent contract vulnerability detection method based on graph neural network and electronic equipment

    CN116467720A

  • RTL-level hardware Trojan horse detection method based on graph neural network and storage medium

    CN116522334A

  • Malicious code classification method and device, storage medium and electronic equipment

    CN117521067A

  • Hardware Trojan horse detection method based on automatic semantic feature extraction

    CN118797640A