Jailbreak vulnerability testing method, device, storage medium and program product

By generating adversarial prompts through multiple non-semantic perturbations of malicious commands, the problem of insufficient attack surface coverage in jailbreak vulnerability testing by black-box optimization methods is solved, thereby improving the accuracy of testing and the defense capabilities of large models.

CN122197036BActive Publication Date: 2026-07-24IFLYTEK CO LTD
2 Cites 0 Cited by

Patent Information

Application Number
CN202610673941.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-15
Publication Date
2026-07-24
Estimated Expiration
2046-05-15

AI Technical Summary

Technical Problem

Existing black-box optimization methods lack sufficient attack surface coverage in jailbreak vulnerability testing, are easily blocked by targeted defenses, and result in low testing effectiveness.

Method used

By performing multiple non-semantic perturbations on malicious commands, multiple adversarial prompts are generated in the first generation. The current generation of adversarial prompts that can successfully jailbreak the target large model are then selected until the evolution termination condition is met, thereby expanding the search space for jailbreak attacks and increasing the diversity and exploration capabilities of adversarial prompts.

Benefits of technology

It improves the accuracy of jailbreak vulnerability testing based on black-box optimization, reduces the risk of missing detection of adversarial prompts, and enhances the defense capabilities of large models.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application discloses a jailbreak vulnerability testing method and device, a storage medium and a program product, relates to the technical field of artificial intelligence, and comprises the following steps: performing multiple non-semantic perturbations on a malicious instruction to obtain multiple adversarial prompt words; each non-semantic perturbation comprises at least one of the following: mapping the malicious instruction into structured content conforming to the syntax of a programming language, and encoding and converting a key sensitive word in the malicious instruction; screening a target current generation adversarial prompt word capable of successfully jailbreaking a target large model; if a target current generation adversarial prompt word is screened out, determining the screened target current generation adversarial prompt word as a jailbreak vulnerability of the target large model; if no target current generation adversarial prompt word is screened out, updating the adversarial prompt word based on the current generation adversarial prompt word and the malicious instruction, returning to the step of screening the target current generation adversarial prompt word capable of successfully jailbreaking the target large model, and continuing until an evolution end condition is met. The application improves the accuracy of jailbreak vulnerability testing based on black box optimization.
Need to check novelty before this filing date? Find Prior Art