The invention discloses a fine-grained evaluation method for logic security of a large
language model, which comprises the following steps of: aiming at an obtained dynamic
test case, analyzing variable interaction influence, injecting a private
instruction sequence by embedding a risk module, and determining a
test case version of which the difficulty is adjusted; selecting an intermediate step sequence from the determined
test case version, and calculating a deviation value with a standard reasoning path by adopting a
sequence comparison algorithm to obtain a fine-grained deviation index set; according to the obtained deviation index set, the accuracy data derived in multiple steps are fused, if the deviation value exceeds a preset threshold value, the deviation value is marked as a logic
bottleneck point, and a potential safety
crash position is judged; quantifying a step-by-step reliability index of the reasoning process through the judged safety
crash position to obtain a numerical
score of the logic capability; and a safety performance subset is extracted from the obtained numerical
score, clustering analysis is adopted to conclude a performance mode in a high-risk scene, and a final comprehensive evaluation report is determined.