Jailbreak vulnerability testing method, device, storage medium and program product
By generating adversarial prompts through multiple non-semantic perturbations of malicious commands, the problem of insufficient attack surface coverage in jailbreak vulnerability testing by black-box optimization methods is solved, thereby improving the accuracy of testing and the defense capabilities of large models.
Patent Information
- Application Number
- CN202610673941.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-07-24
- Estimated Expiration
- 2046-05-15
AI Technical Summary
Existing black-box optimization methods lack sufficient attack surface coverage in jailbreak vulnerability testing, are easily blocked by targeted defenses, and result in low testing effectiveness.
By performing multiple non-semantic perturbations on malicious commands, multiple adversarial prompts are generated in the first generation. The current generation of adversarial prompts that can successfully jailbreak the target large model are then selected until the evolution termination condition is met, thereby expanding the search space for jailbreak attacks and increasing the diversity and exploration capabilities of adversarial prompts.
It improves the accuracy of jailbreak vulnerability testing based on black-box optimization, reduces the risk of missing detection of adversarial prompts, and enhances the defense capabilities of large models.