Intelligent Evolution Generation Method and System for Large-Scale Jailbreak Attack Evaluation Corpus

By using an intelligent evolution generation system and a multimodal feature analysis engine and a dynamic evolution engine, the problems of difficulty in obtaining and uniformity of large-scale jailbreak attack evaluation data are solved, generating diverse evaluation data and improving the evaluation effect of large-scale models on jailbreak attacks.

CN121051739BActive Publication Date: 2026-01-30HANGZHOU ANQUAN DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511596279.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-01-30
Estimated Expiration
2045-11-04

AI Technical Summary

Technical Problem

The existing technology suffers from the difficulty and limited variety of data for evaluating large-scale jailbreak attacks, resulting in incomplete evaluation results that fail to accurately reflect the resistance level of large-scale models to various jailbreak attacks.

Method used

By acquiring initial jailbreak attack evaluation corpus, multi-dimensional features are extracted using a pre-accessed jailbreak attack multimodal feature analysis engine. These features are then automatically evolved using a dynamic evolution engine based on a dual-Q network, generative adversarial network, and large language model to generate new jailbreak attack evaluation corpus. Finally, combined with contextual rationality verification, a hierarchical evolution and dynamic optimization intelligent algorithm system is constructed.

Benefits of technology

It achieves efficient generation of diverse jailbreak attack evaluation corpora, can deeply mine and expand the corpus in a short time, accurately adapt to new attack modes, comprehensively evaluate the defense capabilities of large models, and improve the comprehensiveness and accuracy of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121051739B_ABST
    Figure CN121051739B_ABST
Patent Text Reader

Abstract

This specification discloses an intelligent evolutionary generation method and system for large-scale jailbreak attack evaluation corpora. The method includes: extracting multi-dimensional features from the initial jailbreak attack evaluation corpus based on a jailbreak attack multimodal feature analysis engine; integrating a dynamic evolutionary engine comprising a dual-Q network, a generative adversarial network, and a large language model; the dynamic evolutionary engine automating the evolutionary processing of the corpus's multi-dimensional features; determining optimal evolutionary rules corresponding to the corpus's multi-dimensional features from a corpus evolutionary rule base using the dual-Q network; generating a second derivative corpus corresponding to the first derivative corpus using the generative adversarial network; and performing contextual rationality verification on the second derivative corpus using the large language model. This specification embodiment efficiently generates new jailbreak attack evaluation corpora, improving the large model's resistance to various jailbreak attacks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of large-scale security technology, specifically to an intelligent evolution generation method and system for large-scale jailbreak attack evaluation corpora. Background Technology

[0002] In the field of large-scale model security, jailbreak attack evaluation is crucial. Currently, the main method for obtaining corpora for large-scale model jailbreak attack evaluation is manual collection and organization. Manual collection typically relies on security experts manually filtering text content related to large-scale model jailbreak attacks from various public channels as evaluation corpora. This method has several drawbacks: First, corpus acquisition is difficult. Information related to large-scale model jailbreak attacks is scattered across numerous different platforms and channels, and manually searching and collecting them one by one consumes a significant amount of time and effort, resulting in extremely low efficiency. Moreover, with the rapid development of large-scale model technology, new jailbreak attack techniques are constantly emerging, making it difficult for humans to keep up and obtain the latest attack corpora in a timely manner. Second, there is the problem of corpus uniformity. Manually collected corpora are often limited by the collector's knowledge domain and focus, resulting in a relatively homogeneous corpus type. For example, it may concentrate on a few common jailbreak attack types, such as hint-based injection attacks, while ignoring other relatively niche but equally threatening attack methods. This makes the evaluation results incomplete and unable to accurately reflect the large-scale model's resistance level to various jailbreak attacks. Summary of the Invention

[0003] This specification provides an intelligent evolutionary generation method and system for large-scale jailbreak attack evaluation corpora, the technical solution of which is as follows:

[0004] Firstly, the embodiments of this specification provide an intelligent evolution generation method for large-scale jailbreak attack evaluation corpora, including: acquiring initial jailbreak attack evaluation corpora and an evolution rule base; extracting multi-dimensional features corresponding to the initial jailbreak attack evaluation corpora based on a pre-accessed jailbreak attack multimodal feature analysis engine; accessing a dynamic evolution engine including a dual-Q network, a generative adversarial network, and a large language model, wherein the dynamic evolution engine is used to automatically perform evolution processing on the multi-dimensional features to generate new jailbreak attack evaluation corpora, and the automated evolution processing flow includes: determining the preferred evolution rules corresponding to the multi-dimensional features from the evolution rule base through the dual-Q network, and outputting the first derived corpora corresponding to the multi-dimensional features according to the preferred evolution rules; generating the second derived corpora corresponding to the first derived corpora through the generative adversarial network; and performing contextual rationality verification on the second derived corpora through the large language model to obtain new jailbreak attack evaluation corpora.

[0005] Secondly, the embodiments of this specification provide an intelligent evolution generation system for large-scale jailbreak attack evaluation corpora, including: a corpus acquisition module for acquiring initial jailbreak attack evaluation corpora and an evolution rule base; a feature extraction module for extracting multi-dimensional features corresponding to the initial jailbreak attack evaluation corpora based on a pre-accessed jailbreak attack multimodal feature analysis engine; and an evolution processing module for accessing a dynamic evolution engine including a dual-Q network, a generative adversarial network, and a large language model. The dynamic evolution engine is used to automatically perform evolution processing on the multi-dimensional features to generate new jailbreak attack evaluation corpora. The automated evolution processing flow includes: a first derivation module for determining the preferred evolution rules corresponding to the multi-dimensional features from the evolution rule base through a dual-Q network and outputting the first derived corpora corresponding to the multi-dimensional features according to the preferred evolution rules; a second derivation module for generating a second derived corpora corresponding to the first derived corpora through a generative adversarial network; and a verification module for performing contextual rationality verification on the second derived corpora through a large language model to obtain new jailbreak attack evaluation corpora.

[0006] Thirdly, embodiments of this specification provide an electronic device, including a processor and a memory;

[0007] The processor is connected to the memory; the memory is used to store executable program code; the processor runs the program corresponding to the executable program code by reading the executable program code stored in the memory, so as to perform steps in intelligent evolution generation methods such as large-model jailbreak attack evaluation corpora.

[0008] Fourthly, embodiments of this specification provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the intelligent evolution generation method for large-scale jailbreak attack evaluation corpus.

[0009] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:

[0010] This embodiment of the specification can efficiently generate large-scale jailbreak attack evaluation corpora. By using a pre-accessed jailbreak attack multimodal feature analysis engine, it extracts multidimensional features corresponding to the initial jailbreak attack evaluation corpora, thereby enabling in-depth mining and expansion of existing corpora in a short time. Furthermore, this embodiment can also integrate a dynamic evolution engine including a dual-Q network, a generative adversarial network, and a large language model. The dynamic evolution engine is used to automatically evolve multidimensional features, generating new jailbreak attack evaluation corpora. This embodiment, through the jailbreak attack multimodal feature analysis engine and the dynamic evolution engine, can solve the problems of difficult and singular corpora acquisition in large-scale jailbreak attack evaluation. This embodiment utilizes a dual-Q network to generate a first derived corpus corresponding to multidimensional features, then uses a generative adversarial network to generate a second derived corpus corresponding to the first derived corpus. Next, a large language model is used to perform contextual rationality verification on the second derived corpus, thereby constructing a hierarchical evolution and dynamic optimization intelligent algorithm system. This further achieves efficient generation, accurate adaptation, and dynamic evolution of corpus evolution generation, comprehensively and deeply enhancing the large model's resistance to various jailbreak attacks. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this specification, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a schematic diagram illustrating an application scenario of the intelligent evolution generation method for a large-scale jailbreak attack evaluation corpus provided in this manual.

[0013] Figure 2 This is a flowchart illustrating an intelligent evolutionary generation method for a large-scale jailbreak attack evaluation corpus provided in this manual.

[0014] Figure 3 This is a flowchart illustrating the process of determining the optimal evolution rules corresponding to multidimensional features, as provided in this specification.

[0015] Figure 4 This is a flowchart illustrating the process of generating the second derived corpus provided in this manual.

[0016] Figure 5 This is a schematic diagram of the structure of an intelligent evolution generation system for a large-scale jailbreak attack evaluation corpus provided in this manual.

[0017] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this specification. Detailed Implementation

[0018] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings.

[0019] The terms "first," "second," etc., in the description, claims, and accompanying drawings are used to distinguish different objects and not to describe a particular order. Furthermore, the term "comprising" and any variations thereof are intended to cover a non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to such processes, methods, products, or apparatus.

[0020] The intelligent evolution generation method for large-model jailbreak attack evaluation corpus provided in several embodiments of this specification can be implemented by the intelligent evolution generation system for large-model jailbreak attack evaluation corpus provided in the embodiments of this invention.

[0021] Before this manual elaborates on the intelligent evolution generation method for large-scale jailbreak attack evaluation corpora in conjunction with one or more embodiments, it first introduces the application scenarios of this intelligent evolution generation method for large-scale jailbreak attack evaluation corpora.

[0022] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the intelligent evolution generation method for large-scale jailbreak attack evaluation corpora provided in this embodiment of the invention. In this embodiment, the intelligent evolution generation system 100 for large-scale jailbreak attack evaluation corpora can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC); the server can be a single server or a server cluster composed of multiple servers.

[0023] In some embodiments, the intelligent evolution generation system 100 can also be integrated into multiple electronic devices. For example, the intelligent evolution generation system 100 can be integrated into multiple servers, and multiple servers can implement the intelligent evolution generation method of the large-scale jailbreak attack evaluation corpus of this application.

[0024] In some embodiments, the server may also be implemented as a terminal. The terminal may be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer (PC), etc. The terminal includes a central processing unit (CPU), a graphics processing unit (GPU), memory, storage devices, a network communication module, sensors, a display screen, a battery and power management module, etc.

[0025] For example, refer to Figure 1 The electronic device may include a server 110, a storage terminal 120, etc. The storage terminal 120 stores the initial jailbreak attack evaluation corpus and the derived rule base, etc. The server 110 and the storage terminal 120 communicate with each other, which will not be described in detail here.

[0026] Server 110 may include a processor and memory. Server 110 can acquire initial jailbreak attack evaluation corpus and a derivation rule base; based on a pre-accessed jailbreak attack multimodal feature analysis engine, it extracts multidimensional features corresponding to the initial jailbreak attack evaluation corpus; it accesses a dynamic derivation engine including a dual-Q network, a generative adversarial network, and a large language model. The dynamic derivation engine is used to automatically derivate the multidimensional features to generate new jailbreak attack evaluation corpus. The automated derivation process includes: determining the preferred derivation rules corresponding to the multidimensional features from the derivation rule base using the dual-Q network, and outputting the first derived corpus corresponding to the multidimensional features according to the preferred derivation rules; generating the second derived corpus corresponding to the first derived corpus using the generative adversarial network; and performing contextual rationality verification on the second derived corpus using the large language model to obtain new jailbreak attack evaluation corpus.

[0027] It should be noted that, Figure 1 The schematic diagram of the intelligent evolution generation system for large-scale jailbreak attack evaluation corpus shown is merely an example. The intelligent evolution generation system and scenario for large-scale jailbreak attack evaluation corpus described in this embodiment are for the purpose of more clearly illustrating the technical solutions of this embodiment and do not constitute a limitation on the technical solutions provided by this embodiment. As those skilled in the art will know, with the evolution of the intelligent evolution generation system for large-scale jailbreak attack evaluation corpus and the emergence of new scenarios, the technical solutions provided by this embodiment are also applicable to similar technical problems.

[0028] Please see Figure 2 , Figure 2 This is a flowchart illustrating an intelligent evolutionary generation method for a large-scale jailbreak attack evaluation corpus provided in an embodiment of the present invention. This intelligent evolutionary generation method for the large-scale jailbreak attack evaluation corpus can be achieved by… Figure 1The intelligent evolution generation system 100 shown is executed; the intelligent evolution generation system 100 can be a server, etc. The intelligent evolution generation method for this large-scale jailbreak attack evaluation corpus may include at least the following steps:

[0029] 200. Obtain the initial jailbreak attack evaluation corpus and the derived rule library;

[0030] 210. Based on the pre-accessed jailbreak attack multimodal feature analysis engine, extract multi-dimensional features corresponding to the initial jailbreak attack evaluation corpus; 220. Access a dynamic evolution engine including a dual-Q network, a generative adversarial network, and a large language model. The dynamic evolution engine is used to automatically evolve multi-dimensional features to generate new jailbreak attack evaluation corpus. The automated evolution process includes:

[0031] 2200. Determine the preferred evolution rules corresponding to the multidimensional features from the evolution rule base using a dual-Q network, and output the first derived corpus corresponding to the multidimensional features based on the preferred evolution rules;

[0032] 2210. Generate a second derived corpus corresponding to the first derived corpus using an adversarial generative network;

[0033] 2220. By using a large language model to verify the contextual rationality of the second derived corpus, a new jailbreak attack evaluation corpus is obtained.

[0034] This embodiment presents an intelligent evolution and generation method for large-scale jailbreak attack evaluation corpus. By employing an intelligent algorithm system to perform in-depth analysis and feature extraction on existing jailbreak attack corpus, and then performing automated and intelligent evolution and generation on this basis, a large number of new evaluation corpus with diversity and targeting are generated.

[0035] In this embodiment, the initial jailbreak attack evaluation corpus is a dataset used to evaluate the ability of AI models (such as large language models, content moderation systems, etc.) to resist malicious jailbreak attacks. The initial jailbreak attack evaluation corpus contains adversarial prompts and can simulate illegal requests in real-world scenarios.

[0036] The evolution rule base includes several corpus evolution rules, which can be used as prompts to describe the generation of new corpus based on the features of existing corpus. These corpus evolution rules include synonym substitution, sentence structure transformation, attack type combination, intensity adjustment, scene expansion, and language style conversion.

[0037] For example, synonym substitution rules can replace keywords and phrases in existing corpora based on semantic similarity, while maintaining the malicious intent. Sentence transformation rules can change the syntactic structure of the corpus, such as changing active voice to passive voice, splitting long sentences, and merging short sentences. Attack type combination rules can combine different types of jailbreak attack techniques to generate new composite attack corpora. Intensity adjustment rules can adjust the stealth, complexity, or directness of the attack, such as adding obfuscated information, nesting multiple layers of instructions, and changing the encoding method. Scene expansion rules can expand malicious scenarios based on existing corpora to other related or unrelated scenarios to increase the diversity of the corpus. Language style conversion rules can change the language style of the corpus, such as from formal to informal.

[0038] In some embodiments, based on the pre-accessed jailbreak attack multimodal feature analysis engine, multidimensional features corresponding to the initial jailbreak attack evaluation corpus are extracted, including: based on a natural language processing model, identifying the jailbreak attack type to which the initial jailbreak attack evaluation corpus belongs, obtaining the attack type features corresponding to the initial jailbreak attack evaluation corpus; and extracting the semantic features, syntactic and structural features, context dependency features, and target model attack features corresponding to the initial jailbreak attack evaluation corpus.

[0039] In this embodiment, the jailbreak attack multimodal feature analysis engine may include multiple natural language processing models, such as multi-label BiLSTM-CRF models, BERT models, dual semantic encoders, Stanford CoreNLP, DialogXL, LSTM-Attention classifiers, etc.

[0040] In some embodiments, attack type features include, but are not limited to: instruction injection, obfuscation techniques, chained hints, role-playing, encoding attacks, instruction interference, format anomalies, and fictional scenarios; semantic features include, but are not limited to: malicious intent, sensitive words, themes, and sentiment characteristics; syntactic and structural features include, but are not limited to: syntactic structure, grammatical patterns, rhetorical devices, special symbols, and encoding methods; context-dependent features include, but are not limited to: contextual relationships between corpora, guidance logic, and attack paths; and target model attack features include, but are not limited to: known vulnerabilities of the target model and weaknesses in its defense mechanisms.

[0041] This embodiment accurately adapts to the corpus scenario, ensuring a strong correlation between the corpus and the attack scenario to effectively expose potential model vulnerabilities. During corpus generation, this embodiment utilizes a jailbreak attack multimodal feature analysis engine to meticulously classify and extract features from different attack scenarios, thereby establishing a precise mapping relationship between attack scenarios and corpus features. For each specific attack scenario, such as fraud-induced attacks targeting large models in the financial field or privacy-stealing attacks targeting large models in the medical field, the system will customize and generate highly matched corpus based on the characteristics and needs of the scenario through multidimensional feature analysis. These corpora closely match the actual attack scenarios in terms of semantics, context, and attack intent, accurately simulating the attacker's behavior and methods in real-world scenarios. By using these precisely adapted corpora for evaluation, the potential vulnerabilities of large models facing different attack scenarios can be more deeply explored, providing targeted basis for model security optimization.

[0042] In some embodiments, please refer to Figure 3 , Figure 3 This is a flowchart illustrating the process of determining the preferred evolution rules corresponding to multidimensional features according to an embodiment of the present invention. The process involves determining the preferred evolution rules corresponding to multidimensional features from a evolution rule base using a dual-Q network, including:

[0043] 300. Concatenate the multidimensional features into a first feature vector, which is the state input of the dual-Q network;

[0044] 310. Dynamic Q-value evaluation of several corpus evolution rules in the evolution rule base is performed using a double Q network. The reward function of the double Q network is the BLEU score between the first feature vector and the derived corpus generated by the corpus evolution rules.

[0045] 320. The optimal evolution rules corresponding to the multidimensional features are output through the dual-Q network.

[0046] In this embodiment, the dual-Q network is an improved Q-learning algorithm based on reinforcement learning, which can solve the problem of overestimation caused by maximizing bias in traditional deep Q-networks. The dual-Q network reduces error by decoupling action selection and action value evaluation, using two independent Q-networks.

[0047] In this embodiment, multidimensional features are concatenated into a first feature vector, which is then used as the state input of the dual-Q network. The first feature vector is processed by corpus evolution rules to generate derived corpus. In the dual-Q network set in this embodiment, the reward function is the BLEU score between the first feature vector and the derived corpus generated by the corpus evolution rules. The BLEU score calculation formula is as follows: BLEU is used to calculate the matching ratio between the derived corpus and the first feature vector. The proportion of n-grams matching the first feature vector in the derived corpus generated by the corpus evolution rules. For the weighted weights of different n-grams, It can be set to 1 / N, where N is the order of the maximum n-gram, and n-gram is a sequence of consecutive words; for example, the text "natural language processing" has the corresponding 2-gram as ["natural language", "language processing"]. The feature vector "0101 1110 1011" has the corresponding 2-gram as ["0101 1110 ", "1110 1011"].

[0048] In some embodiments, please refer to Figure 4 , Figure 4 This is a schematic diagram of the process for generating a second derived corpus provided in an embodiment of the present invention. The adversarial generative network includes a generator and an evaluator. The second derived corpus corresponding to the first derived corpus is generated through the adversarial generative network, including:

[0049] 400. Determine the second feature vector corresponding to the first derived corpus;

[0050] 410. Input the second feature vector into the generator to obtain the first derived sample corresponding to the second feature vector;

[0051] 420. Input the first derived sample into the evaluator to obtain the similarity probability between the first derived sample and the second feature vector;

[0052] 430. With the goal of increasing the similarity probability, update the parameters of the evaluator, then fix the parameters of the evaluator and update the parameters of the generator to obtain the trained generator;

[0053] 440. The trained generator is used to generate target derived samples, which are the second derived samples corresponding to the first derived corpus.

[0054] In this embodiment, the Generative Adversarial Network (GAN) is a deep learning model consisting of a generator and an evaluator. This embodiment learns the data distribution through dynamic game theory between the generator and evaluator, thereby generating realistic new samples, such as target-derived samples. The generator and evaluator of this GAN can be recurrent neural networks (such as RNN, LSTM, GRU), BERT models, etc.

[0055] In this embodiment, the generator generates the first derived sample corresponding to the second feature vector, i.e., generates fake data that is as similar as possible to the real data. Then, the first derived sample is input into the evaluator, which distinguishes whether the input data comes from the real dataset or is fabricated by the generator. That is, with the goal of increasing the similarity probability, the evaluator's parameters are updated. Afterwards, the evaluator's parameters are fixed, and the generator's parameters are updated to obtain the trained generator. This embodiment uses adversarial optimization to train the generator and evaluator, thereby obtaining a trained generator. The trained generator then generates target derived samples, i.e., the second derived corpus, that are as similar as possible to the first derived corpus.

[0056] In some embodiments, the second derived corpus is subjected to contextual rationality verification using a large language model to obtain a new jailbreak attack evaluation corpus, including: determining the semantic similarity between the initial jailbreak attack evaluation corpus and the second derived corpus using a large language model; when the semantic similarity is greater than a similarity threshold, the second derived corpus is the new jailbreak attack evaluation corpus.

[0057] In this embodiment, the semantic similarity can be calculated by using the [CLS] vector of a pre-trained language model such as BERT to determine the semantic similarity between the second derived corpus and the initial jailbreak attack evaluation corpus. When the semantic similarity is greater than the similarity threshold, the second derived corpus is a new jailbreak attack evaluation corpus.

[0058] The embodiments described in this specification can be configured with real-time optimization generation strategies to adapt to new attack patterns. With the continuous development of large-scale model technology and the constant innovation of attackers' methods, new jailbreak attack patterns are constantly emerging. The system in this embodiment possesses real-time monitoring and analysis capabilities, enabling it to promptly capture newly emerging attack patterns and characteristics. Once a new attack pattern is detected, the system will quickly adjust and optimize the existing generation strategy. By introducing machine learning and deep learning algorithms, the system learns and models new attack patterns, updating the rules and parameters for corpus generation. For example, when an attack pattern based on a new semantic understanding vulnerability emerges, the system can automatically adjust the semantic feature extraction and transformation rules in the corpus generation process to generate test corpus targeting that vulnerability. This dynamic evolution capability ensures that the corpus generated by this invention can always keep up with changes in attack patterns, providing continuous and effective support for the security evaluation of large-scale models.

[0059] This embodiment of the specification can efficiently generate large-scale jailbreak attack evaluation corpora. By using a pre-accessible jailbreak attack multimodal feature analysis engine, it extracts multidimensional features corresponding to the initial jailbreak attack evaluation corpus, thereby enabling in-depth mining and expansion of existing corpora in a short time. Furthermore, this embodiment can also integrate a dynamic evolution engine including dual-Q networks, generative adversarial networks, and large language models. The dynamic evolution engine is used to automatically perform evolutionary processing on multidimensional features, generating new jailbreak attack evaluation corpora. This embodiment, through the jailbreak attack multimodal feature analysis engine and the dynamic evolution engine, can solve the problems of difficult and singular corpus acquisition in large-scale jailbreak attack evaluation. This specification's embodiments utilize a dual-Q network to generate a first derived corpus corresponding to multi-dimensional features, then use an adversarial generative network to generate a second derived corpus corresponding to the first derived corpus. Next, a large language model is used to verify the contextual rationality of the second derived corpus, thereby constructing a hierarchical evolution and dynamic optimization intelligent algorithm system. This further enables efficient generation, accurate adaptation, and dynamic evolution of corpus evolution, comprehensively and deeply assisting in accurately evaluating the large model's resistance level to various jailbreak attacks.

[0060] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0061] Please see Figure 5 , Figure 5 This is a schematic diagram of the structure of an intelligent evolution generation system for evaluating large-scale jailbreak attacks, provided in an embodiment of this specification.

[0062] like Figure 5 As shown, the intelligent evolution generation system for the large-scale jailbreak attack evaluation corpus can include at least a corpus acquisition module 500, a feature extraction module 510, an evolution processing module 520, a first derivation module 5200, a second derivation module 5210, and a verification module 5220, wherein:

[0063] The corpus acquisition module 500 is used to acquire the initial jailbreak attack evaluation corpus and the derived rule library;

[0064] The feature extraction module 510 is used to extract multi-dimensional features corresponding to the initial jailbreak attack evaluation corpus based on the pre-accessed jailbreak attack multimodal feature analysis engine; the evolution processing module 520 is used to access a dynamic evolution engine including a double-Q network, a generative adversarial network, and a large language model. The dynamic evolution engine is used to automatically perform evolution processing on the multi-dimensional features to generate new jailbreak attack evaluation corpus. The automated evolution processing flow includes:

[0065] The first derivative module 5200 is used to determine the preferred derivative rules corresponding to the multidimensional features from the derivative rule base through the double Q network, and output the first derivative corpus corresponding to the multidimensional features according to the preferred derivative rules;

[0066] The second derivative module 5210 is used to generate a second derivative corpus corresponding to the first derivative corpus through an adversarial generative network.

[0067] The verification module 5220 is used to perform contextual rationality verification on the second derived corpus through a large language model to obtain a new jailbreak attack evaluation corpus.

[0068] In some embodiments, the first derivation module 5200 includes a rule selection module, which is used to: concatenate multidimensional features into a first feature vector, the first feature vector being the state input of a double-Q network; dynamically evaluate the Q-values ​​of several corpus derivation rules in the derivation rule base through the double-Q network, the reward function of the double-Q network being the BLEU score between the first feature vector and the derived corpus generated by the corpus derivation rules of the first feature vector; and output the preferred derivation rule corresponding to the multidimensional features through the double-Q network.

[0069] In some embodiments, the adversarial generative network includes a generator and an evaluator. The second derivation module 5210 includes an adversarial generation module, which is configured to: determine a second feature vector corresponding to a first derived corpus; input the second feature vector into the generator to obtain a first derived sample corresponding to the second feature vector; input the first derived sample into the evaluator to obtain a similarity probability between the first derived sample and the second feature vector; update the parameters of the evaluator with the goal of increasing the similarity probability, and then fix the parameters of the evaluator and update the parameters of the generator to obtain a trained generator; the trained generator is used to generate a target derived sample, which is the second derived corpus corresponding to the first derived corpus.

[0070] In some embodiments, the verification module 5220 includes a semantic verification module, which is used to: determine the semantic similarity between the initial jailbreak attack evaluation corpus and the second derived corpus through a large language model; when the semantic similarity is greater than the similarity threshold, the second derived corpus is a new jailbreak attack evaluation corpus.

[0071] In some embodiments, the feature extraction module 510 includes a feature extraction submodule, which is used to: identify the jailbreak attack type to which the initial jailbreak attack evaluation corpus belongs based on a natural language processing model, obtain the attack type features corresponding to the initial jailbreak attack evaluation corpus; and extract the semantic features, syntactic and structural features, context dependency features, and target model attack features corresponding to the initial jailbreak attack evaluation corpus.

[0072] In some embodiments, attack type features include, but are not limited to: instruction injection, obfuscation techniques, chained hints, role-playing, encoding attacks, instruction interference, format anomalies, and fictional scenarios; semantic features include, but are not limited to: malicious intent, sensitive words, themes, and sentiment characteristics; syntactic and structural features include, but are not limited to: syntactic structure, grammatical patterns, rhetorical devices, special symbols, and encoding methods; context-dependent features include, but are not limited to: contextual relationships between corpora, guidance logic, and attack paths; and target model attack features include, but are not limited to: known vulnerabilities of the target model and weaknesses in its defense mechanisms.

[0073] In some embodiments, the evolution rule base includes several corpus evolution rules, which include synonym substitution, sentence transformation, attack type combination, intensity adjustment, scenario expansion, and language style conversion.

[0074] Based on the intelligent evolution generation system for large-model jailbreak attack evaluation corpora in several embodiments of this specification, it can be seen that the embodiments of this specification can efficiently generate large-model jailbreak attack evaluation corpora. By using a pre-accessed jailbreak attack multimodal feature analysis engine, multidimensional features corresponding to the initial jailbreak attack evaluation corpus are extracted, thereby enabling in-depth mining and expansion of existing corpora in a short time. Furthermore, the embodiments of this specification can also access a dynamic evolution engine including dual-Q networks, generative adversarial networks, and large language models. The dynamic evolution engine is used to automatically evolve multidimensional features to generate new jailbreak attack evaluation corpora. Through the jailbreak attack multimodal feature analysis engine and the dynamic evolution engine, the embodiments of this specification can solve the problems of difficulty in obtaining and uniformity of corpora in large-model jailbreak attack evaluation. This specification's embodiments utilize a dual-Q network to generate a first derived corpus corresponding to multi-dimensional features, then use an adversarial generative network to generate a second derived corpus corresponding to the first derived corpus. Next, a large language model is used to verify the contextual rationality of the second derived corpus, thereby constructing a hierarchical evolution and dynamic optimization intelligent algorithm system. This further enables efficient generation, accurate adaptation, and dynamic evolution of corpus evolution, comprehensively and deeply assisting in accurately evaluating the large model's resistance level to various jailbreak attacks.

[0075] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the embodiment of the intelligent evolution generation system for large-model jailbreak attack evaluation corpus is relatively simple in description because it is fundamentally similar to the embodiment of the intelligent evolution generation method for large-model jailbreak attack evaluation corpus; relevant parts can be referred to the descriptions in the method embodiment.

[0076] Please see Figure 6 The diagram shown is a structural schematic of an electronic device based on a start-up control device provided in an embodiment of this specification.

[0077] like Figure 6 As shown, the electronic device 600 may include at least one processor 610, at least one network interface 640, a user interface 630, a memory 650, and at least one communication bus 620.

[0078] The communication bus 620 can be used to realize the connection and communication of the above components.

[0079] The user interface 630 may include buttons, and the optional user interface may also include a standard wired interface or a wireless interface.

[0080] The network interface 640 may include, but is not limited to, Bluetooth modules, NFC modules, ZigBee modules, and UWB modules.

[0081] The processor 610 may include one or more processing cores. The processor 610 connects to various parts within the electronic device 600 using various interfaces and lines. It performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 650, and by calling data stored in the memory 650. Optionally, the processor 610 may be implemented using at least one hardware form selected from DSP, FPGA, and PLA. The processor 610 may integrate one or more combinations of CPU and GPU. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content required for display on the screen.

[0082] The memory 650 may include RAM or ROM. Optionally, the memory 650 may include a non-transitory computer-readable medium. The memory 650 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 650 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 650 may also be at least one storage device located remotely from the aforementioned processor 610. As a computer storage medium, the memory 650 may include an operating system, a communication module, a user interface module, and an intelligent evolution generation application for large-model jailbreak attack evaluation corpus. The processor 610 may be used to call the intelligent evolution generation application for large-model jailbreak attack evaluation corpus stored in the memory 650 and execute the steps of intelligent evolution generation of the large-model jailbreak attack evaluation corpus mentioned in the foregoing embodiments.

[0083] This specification also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform the above-described instructions. Figures 2-4 One or more steps in the illustrated embodiment. If the constituent modules of the above-described electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.

[0084] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this specification is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).

[0085] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks. Unless otherwise specified, the technical features of this embodiment and its implementation can be combined arbitrarily.

[0086] The above embodiments are merely preferred embodiments described in this specification and are not intended to limit the scope of this specification. Any modifications and improvements made by those skilled in the art to the technical solutions of this specification without departing from the spirit of this specification should fall within the protection scope defined by the claims of this specification.

Claims

1. An intelligent derivation generation method for large model jailbreaking attack evaluation corpus, characterized in that, The application relates to a method for generating a new jailbreak attack evaluation corpus, comprising the following steps: an initial jailbreak attack evaluation corpus and a derivation rule library are acquired; a pre-accessed jailbreak attack multi-modal feature analysis engine is used to extract multi-dimensional features corresponding to the initial jailbreak attack evaluation corpus; a dynamic derivation engine comprising a double Q network, a generative adversarial network and a large language model is accessed, the dynamic derivation engine is used for automatic derivation processing of the multi-dimensional features, and a new jailbreak attack evaluation corpus is generated, the automatic derivation processing process comprises the following steps: a preferred derivation rule corresponding to the multi-dimensional features is determined from the derivation rule library through the double Q network, and a first derived corpus corresponding to the multi-dimensional features is output according to the preferred derivation rule; a second derived corpus corresponding to the first derived corpus is generated through the generative adversarial network; context rationality verification of the second derived corpus is performed through the large language model, and a new jailbreak attack evaluation corpus is obtained.

2. The method of claim 1, wherein, The preferred derivation rule corresponding to the multi-dimensional features is determined from the derivation rule library through the double Q network, and the method comprises the following steps: the multi-dimensional features are spliced into a first feature vector, and the first feature vector is a state input of the double Q network; dynamic Q value evaluation of a plurality of corpus derivation rules in the derivation rule library is performed through the double Q network, a reward function of the double Q network is a BLEU score between the first feature vector and a derived corpus generated by the first feature vector through the corpus derivation rule; the preferred derivation rule corresponding to the multi-dimensional features is output through the double Q network.

3. The method of claim 1, wherein, The generative adversarial network comprises a generator and an evaluator, and the second derived corpus corresponding to the first derived corpus is generated through the generative adversarial network, which comprises the following steps: a second feature vector corresponding to the first derived corpus is determined; the second feature vector is input into the generator to obtain a first derivation sample corresponding to the second feature vector; the first derivation sample is input into the evaluator to obtain a similarity probability between the first derivation sample and the second feature vector; parameters of the evaluator are updated with the aim of increasing the similarity probability, and then the parameters of the evaluator are fixed, parameters of the generator are updated, and a trained generator is obtained; the trained generator is used for generating a target derivation sample, and the target derivation sample is the second derived corpus corresponding to the first derived corpus.

4. The method of claim 1, wherein, The context rationality verification of the second derived corpus through the large language model comprises the following steps: the semantic similarity between the initial jailbreak attack evaluation corpus and the second derived corpus is determined through the large language model; when the semantic similarity is greater than a similarity threshold, the second derived corpus is a new jailbreak attack evaluation corpus.

5. The method of claim 1, wherein, The pre-accessed jailbreak attack multi-modal feature analysis engine is used to extract the multi-dimensional features corresponding to the initial jailbreak attack evaluation corpus, and the method comprises the following steps: Based on a natural language processing model, a jailbreak attack type to which the initial jailbreak attack evaluation corpus belongs is identified, and attack type features corresponding to the initial jailbreak attack evaluation corpus are obtained; and semantic features, syntax and structure features, context dependency features, and target model attack features corresponding to the initial jailbreak attack evaluation corpus are extracted.

6. The method of claim 5, wherein, The attack type features include but are not limited to instruction injection, obfuscation technology, chain prompt, role playing, encoding attack, instruction interference, format anomaly, and fictitious scenario; the semantic features include but are not limited to malicious intent, sensitive vocabulary, theme, and sentiment tendency feature; The syntax and structure features include but are not limited to syntax structure, grammar pattern, rhetorical device, special symbol, and encoding mode; the context dependency features include but are not limited to context correlation between corpora, guiding logic, and attack path; and the target model attack features include but are not limited to known vulnerabilities of a target model and defense mechanism weakness features.

7. The method of claim 1, wherein, The derivation rule library includes a plurality of corpus derivation rules, including synonym substitution, sentence transformation, attack type combination, intensity adjustment, scenario expansion, and language style conversion.

8. An intelligent derivation generation system for large model jailbreaking attack evaluation corpus, characterized in that, It comprises: A corpus acquisition module configured to acquire initial jailbreak attack evaluation corpus and a derivation rule library; A feature extraction module configured to extract multi-dimensional features corresponding to the initial jailbreak attack evaluation corpus based on a pre-accessed jailbreak attack multi-modal feature analysis engine; A derivation processing module configured to access a dynamic derivation engine including a double Q network, a generative adversarial network, and a large language model, the dynamic derivation engine being configured to perform automatic derivation processing on the multi-dimensional features to generate new jailbreak attack evaluation corpus, the automatic derivation processing procedure including: A first derivation module configured to determine, by the double Q network, preferred derivation rules corresponding to the multi-dimensional features from the derivation rule library, and output first derived corpus corresponding to the multi-dimensional features according to the preferred derivation rules; A second derivation module configured to generate second derived corpus corresponding to the first derived corpus by the generative adversarial network; A verification module configured to perform context rationality verification on the second derived corpus by the large language model to obtain new jailbreak attack evaluation corpus.

9. An electronic device, comprising: It comprises a processor and a memory; The processor is connected to the memory; The memory is configured to store executable program code; The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, to execute the method of any one of claims 1-7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the method of any one of claims 1-7. The computer program, when executed by the processor, implements the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Prison break attack instruction data generation method and device, medium and equipment

    CN117131513A

  • Multi-modal large model confrontation safety detection method and system

    CN120639526A