The application discloses an
open source dialogue model-oriented automatic jailbreaking prompt
word generation and
attack method and
system. The application first screens a prompt word set from an original prompt word set; secondly, based on an
attack question, the prompt word with the optimal
attack efficiency is screened out from the original
open source prompt word set, multi-path parallel testing is carried out by using a proxy model, and the attack success rate and the prompt
word length of different prompt words in the prompt word set are evaluated in real time through a dynamic evaluation
mechanism based on a
greedy selection strategy, the final attack prompt word is output to a target model, and after being spliced with the attack question, the attack prompt word is returned and an analysis report is output; finally, the prompt word in the analysis report is adjusted by using an automatic
mutation and expert modification strategy, and the prompt word is reconstructed and iterated according to the feedback result of each round of attack, so that a new prompt word is generated. The application combines offline evolution and online decision-making to automatically generate a high-success-rate jailbreaking prompt word, controls the length and overhead, and optimizes
mutation by using semantic
rewriting and logic skeleton extraction.