The invention provides an adaptive
large model hint
attack and
security assessment method based on feedback learning, and belongs to the technical field of
artificial intelligence security. The method comprises the following steps: constructing an
attack library containing a plurality of text transformation
attack strategies, and obtaining
expression data of the attack strategies on a target model; the success rate, the average
score and the R value and the Q value of the
attack strategy are calculated through the data, and the optimal
attack strategy is selected in a descending order according to the Q value; applying the strategy to an original malicious prompt to generate adversarial input, submitting the adversarial input to a target model, judging whether an attack is successful or not through a pre-
training evaluation model
score, if yes, re-selecting the strategy, if yes, updating statistical data, re-calculating the
score, and iteratively trying the strategy according to a Q value until a preset condition is met; and finally recording a test process and generating a report. According to the method, the security evaluation of the
large model is realized, the success rate and efficiency of attack testing are improved, an empirical basis is provided for the design of a
large model defense mechanism, and the
application security of the large model is ensured.