The invention discloses a multi-task interpretable
hardware Trojan horse detection method and
system based on a graph attention mechanism, and belongs to the field of
integrated circuit hardware safety. The method comprises the following steps: analyzing a gate-level
netlist, abstracting the gate-level
netlist into a
directed graph, and extracting node features; the multi-head graph
attention network is used for learning nodes and embedding representation of the graph; node-level Trojan classification and graph-level
risk classification are executed at the same time through a multi-
task learning framework, and joint optimization training is carried out; based on probability distribution output by the model, calculating a
hazard severity risk value and a triggering
score in combination with a preset
risk assessment function, and fusing to obtain a comprehensive risk
score; and for the high-risk
netlist, the feature contribution degree is quantified by adopting GNNExplainer, an evidence chain is constructed, and a large
language model is driven to generate an interpretable report containing risk composition, key evidence and an abnormal mode. According to the method, high-precision
Trojan horse detection and fine-grained
risk assessment are realized, an intuitive and credible decision basis is provided, and the practicability and
interpretability of a detection result are remarkably improved.