The present application provides a
large model endogenous security defense method based on interlayer contrast decoding, belonging to the technical field of
artificial intelligence and
network security, which can at least partially solve the problems of high fine-tuning cost, lagging defense and easy illusion of the existing large
language model security defense means. The present application comprises: offline construction of field risk
data set and input of model, utilization of JS
divergence to analyze the difference between each layer output and final layer output, positioning of high-risk sensitive layer set and extraction of harmful feature
fingerprint; real-time monitoring of the activation mode of generated word units in the sensitive layer in the
inference stage, triggering of security intervention when the similarity with the harmful feature
fingerprint exceeds the threshold; execution of negative contrast decoding, stripping of the sensitive layer output from the final layer output through weighted subtraction operation; introduction of an adaptive
rollback mechanism, dynamic adjustment of the inhibition coefficient or
rollback decoding mode according to the
perplexity. The present application can dynamically block hidden semantic
attack and harmful
content generation at the decoding end without fine-tuning
model parameters.