A Method and Device for Large Model Security Alignment Based on Multi-Round Red Team Attacks
Through the security alignment method of multiple rounds of red team attacks, using thinking guidance and multiple rounds of reinforcement learning of future rewards, the multi-objective reward function of the target model is optimized, and the security vulnerability problem in multiple rounds of dialogue scenarios is solved, efficient and generalized security defense is achieved, and the model's anti-attack ability and practicality are improved.
Patent Information
- Application Number
- CN202510609811.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The prior art is difficult to effectively identify and defend against diverse attack strategies in multiple rounds of dialogue scenarios, and traditional methods are difficult to balance security and practicality, resulting in security vulnerabilities in the model in multiple rounds of interaction.
The secure alignment method based on multiple rounds of red team attacks is adopted, and the red team initialization data set is constructed through thinking guidance, fine-tuning and multi-round interaction is performed, and the multi-round reinforcement learning algorithm for future rewards is combined with the multi-round reinforcement learning algorithm to optimize the multi-objective reward function of the target model to achieve dynamic defense.
It significantly improves the ability of large language models to resist jailbreak attacks in multiple rounds of dialogue scenarios, improves defense effectiveness and strategy generalization capabilities, while maintaining the practicality and interaction fluency of the model, reducing training costs and time.
Smart Images

Figure CN120146199B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and in particular to a large model security alignment method and device based on multiple rounds of red team attacks. Background Art
[0002] With the widespread application of large language models (LLMs) in open domain dialogue scenarios, their security alignment problem has gradually become a research focus. Traditional security alignment methods (such as supervised fine-tuning and reinforcement learning based on human feedback) can effectively defend against direct malicious attacks through adversarial training of single-round harmful instructions. However, in multi-round dialogue scenarios, attackers often hide their true intentions through strategies such as gradual induction and intent disguise, causing the model to accumulate risks and generate harmful content in dynamic interactions. Although existing studies have proposed a defense framework based on red team confrontation, it is mainly designed for single-round attacks and lacks the ability to model multi-round strategic attacks, making it difficult to capture potential threats from the evolution of dialogue states. This limitation makes the security defense of existing models in multi-round scenarios face challenges such as insufficient data dynamics and limited attack mode coverage.
[0003] The current multi-round security alignment technology mainly relies on static adversarial datasets or rule-based red team attack generation, which has significant defects. For example, red team frameworks such as PAIR and COA improve attack strength by iteratively generating adversarial problems, but their attack strategies lack systematic planning, resulting in insufficient diversity and strategy of the generated dialogue trajectories. At the same time, defense methods based on single-round preference optimization (such as DPO and IPO) are prone to fall into local optimality in multi-round scenarios and cannot effectively model the risk accumulation effect in long-term conversations. In addition, existing methods often reduce risks by simply refusing to answer, but overly conservative strategies will damage the helpfulness and fluency of the model, causing the problem of "over-rejection". How to build a dynamic defense mechanism that can not only identify multi-round attack patterns but also balance security and practicality has become a technical bottleneck that needs to be broken through in this field.
[0004] The closest prior art is red team adversarial frameworks based on multi-round reinforcement learning, such as Red-Queen and ActorAct. These methods generate data through the interactive confrontation between the target model and the red team model, and optimize the defense ability using policy gradients. However, they have deficiencies in key aspects: on the one hand, the red team attacks lack an explicit policy planning module, resulting in insufficient concealment and coherence of the attack paths; on the other hand, the immediate reward mechanism of traditional reinforcement learning is difficult to capture the long-term risk dependencies in multi-round conversations, causing short-sighted defense strategies. The recently proposed Chain-of-Attack attempts to introduce intermediate reasoning steps to enhance the attack logic, but its reasoning template is fixed and relies on manual design, restricting the attack generalization ability. These defects indicate that the prior art has not yet resolved the core contradiction between the dynamic policy modeling of multi-round attacks and the sustainable optimization of defense.
[0005] Existing large model security alignment techniques mainly focus on the following types of methods: security alignment methods based on reinforcement learning, security alignment methods based on automated red teams, security alignment methods based on decoding modules, and security alignment methods based on harmful word detection. Although these methods have solved some problems in multi-objective alignment to a certain extent, they all have significant deficiencies.
[0006] 1. Security alignment methods based on reinforcement learning: The representative of this type of method is SafeRLHF, which guides the model to generate harmless responses through preference optimization. These methods rely on manually annotated single-round security data, but static data sets cannot capture the gradually escalating malicious strategies in multi-round attacks. More critically, the immediate reward mechanism of reinforcement learning easily causes the model to fall into a short-term security trap - either over-rejecting reasonable requests.
[0007] 2. Security alignment methods based on automated red teams: Representative works include MART and HARM. The continuous confrontation between the red team model and the target model can theoretically generate more complex attack samples, but in practice, it is found that the attack strategies are prone to falling into a homogeneous cycle. Due to the lack of explicit modeling of the evolution of multi-round intentions, the red team model often repeatedly uses limited attack templates, and the semantic diversity of the generated samples is only 0.15 - 0.3. At the same time, the defense model is prone to overfitting to the current attack pattern, and when encountering new attacks across modalities or languages, the defense success rate drops sharply to below 50%.
[0008] 3. Security alignment methods based on decoding modules: Representatives of such methods are SafeLora and PrimeGuard. Although these methods can be quickly deployed, they face semantic understanding bottlenecks. For example, the misjudgment rate for metaphorical or ironic attacks exceeds 35%, and the real-time resampling mechanism causes the inference delay to surge by more than 5 times. Even more troublesome is that attackers can easily bypass surface word filtering through adversarial prefix construction (such as "Please restate the following content in academic language"), exposing the vulnerability of pure engineering solutions.
[0009] 4. Security alignment methods based on harmful word detection: Representatives of such methods are ToxiChat. This method, based on the detection method of a harmful word library, places the security defense line in the front, intercepting potential risks through multi-level sensitive word matching. However, this method has systematic defects at the semantic level: it can neither identify new attack methods such as code obfuscation nor frequently misjudges legitimate queries. Statistics show that 25% of requests containing sensitive words but actually harmless are wrongly intercepted, seriously affecting the user experience. More importantly, the word library is difficult to defend against logical attacks, and attackers can easily break through the defense line through role-playing or problem decomposition.
[0010] Although the above solutions have made certain progress, there are still limitations in dealing with diverse and complex multi-round attacks. Attackers can take advantage of the temporal characteristics of multi-round conversations and gradually break through the defense line through means such as intention disguise, semantic disassembling, and context grafting, while traditional security mechanisms are difficult to capture hidden cross-round associations. In addition, the existing technology's trade-off between security and usefulness is still relatively crude. Either the normal Q&A fluency is damaged due to over-defense, or the security boundary is forced to be relaxed to maintain the interaction experience. The lack of this dynamic game ability makes it difficult for the current model to achieve true security generalization when facing continuously evolving adversarial attacks.
[0011] Since the release of ChatGPT, jailbreak attacks have spread rapidly on social media, indicating that vulnerabilities in large language models (LLMs) can be exploited to trigger harmful behaviors. Such attacks usually use carefully designed inputs to instruct the model to bypass security and ethical safeguards, resulting in harmful outputs. Currently, the mainstream jailbreak methods are all based on single-round conversations, which can cause harmful reactions in the victim's LLM within one round of the conversation. However, many current studies have found that large models are more easily broken in multiple conversation rounds. Multi-round conversations represent an important application of language models, and ensuring the security of large models in multi-round interactions is a challenging problem: 1. There are various multi-round jailbreak methods, and it is difficult to collect sufficient security alignment data through manual methods. 2. The current security alignment algorithms mainly focus on single-round scenarios and lack algorithms that can effectively perform multi-round security alignment. Summary of the Invention
[0012] To solve the technical problems in the prior art that there are various multi-round jailbreak methods, making it difficult to collect sufficient secure alignment data through manual methods; and the current secure alignment algorithms mainly focus on single-round scenarios and lack algorithms that can effectively perform multi-round secure alignment, the embodiments of the present invention provide a large model secure alignment method and device based on multi-round red team attacks. The technical solutions are as follows:
[0013] On the one hand, a large model secure alignment method based on multi-round red team attacks is provided, characterized in that the method includes:
[0014] S1. Obtain the original red team model and the target model; construct a red team initialization dataset;
[0015] S2. Based on the way of thinking guidance, combine the red team initialization dataset to construct a pre-attack thinking dataset;
[0016] S3. Fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data with the highest diversity score in the pre-attack thinking dataset to fine-tune the red team model;
[0017] S4. The red team model interacts with the target model in multiple rounds, and generates preference data pairs containing future rewards based on trajectory sampling;
[0018] S5. Construct a multi-objective reward function for the target model based on the preference data pairs, and perform multi-round reinforcement optimization on the target model; perform direct preference optimization on the red team model based on the preference data pairs; fine-tune the target model after multi-round attack and defense to obtain the securely aligned target model.
[0019] Optionally, in S2, based on the way of thinking guidance, combine the pre-attack thinking dataset to construct a red team initialization dataset, including:
[0020] Based on the way of thinking guidance, combine the pre-attack thinking dataset to guide the red team model to generate strategic multi-round adversarial prompts to obtain the pre-attack thinking dataset;
[0021] The pre-attack thinking dataset classifies the attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategic thinking process before generating attack questions.
[0022] Optionally, based on the way of thinking guidance, combine the pre-attack thinking dataset to guide the red team model to generate strategic multi-round adversarial prompts, including:
[0023] The red team agent has a conversation with the target model; based on the attack target, describe the objectionable content sought by the attacker;
[0024] After receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interaction process continues until the total number of rounds H is reached.
[0025] Optionally, construct a multi-objective reward function for the target model based on preference data, including:
[0026] For the reward of the target model, construct a multi-objective reward function based on the toxicity score and helpfulness score of the final state.
[0027] Optionally, the optimization objective of the target model is as follows:
[0028] ;
[0029] where, , represent the response under the original trajectory and the response after sampling the trajectory respectively; μ represents the hyperparameter that controls the gradient update speed; t represents the current number of rounds; is the target model at the t-th round; represents the state of the target model in the current round; represents the state of the target model after resampling; is the reward function of the target model; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, and the stability of policy update is ensured through KL divergence constraint..
[0030] Optionally, perform direct preference optimization on the red team model based on preference data, including:
[0031] The loss function By comparing the final rewards of different attack strategies, screen high-toxicity and high-diversity attack samples for reinforcement according to the following formula:
[0032] ;
[0033] where, adv represents the red team model; is the red team model at the t-th round; represents the attack hint of the red team model at the h-th round; represents the current state space of the red team model; β represents the parameter coefficient; and represent the attack under the original trajectory and the attack after sampling the trajectory respectively.
[0034] On the other hand, a large model security alignment device based on multi-round red team attacks is provided. This device is applied to the large model security alignment method based on multi-round red team attacks. The device includes:
[0035] A dataset construction module for obtaining an original red team model and a target model; constructing a red team initialization dataset;
[0036] A dataset initialization module for constructing a pre - attack thinking dataset based on a thinking - guiding approach in combination with the red team initialization dataset;
[0037] A fine - tuning module for fine - tuning the original red team model based on the pre - attack thinking dataset, and only selecting the top K data with the highest diversity scores in the pre - attack thinking dataset to fine - tune the red team model;
[0038] A preference data pair generation module for conducting multiple rounds of interaction between the red team model and the target model, and generating preference data pairs containing future rewards based on trajectory sampling;
[0039] An optimization module for constructing a multi - objective reward function for the target model based on the preference data pairs, and performing multiple rounds of reinforcement optimization on the target model; directly optimizing the red team model based on the preference data pairs; fine - tuning the target model after multiple rounds of attack and defense to obtain a securely aligned target model.
[0040] Optionally, a dataset initialization module for guiding the red team model to generate strategic multi - round adversarial prompts based on a thinking - guiding approach in combination with the pre - attack thinking dataset to obtain the pre - attack thinking dataset;
[0041] The pre - attack thinking dataset classifies attack strategies into four categories: intention reversal, problem decomposition, role - playing, and mixed mode, and requires the red team model to output the strategy thinking process before generating attack questions.
[0042] On the other hand, a large - model security alignment device based on multi - round red team attacks is provided. The large - model security alignment device based on multi - round red team attacks includes: a processor; a memory, and computer - readable instructions are stored on the memory. When the computer - readable instructions are executed by the processor, any one of the methods in the above - mentioned large - model security alignment method based on multi - round red team attacks is implemented.
[0043] On the other hand, a computer - readable storage medium is provided. At least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above - mentioned large - model security alignment method based on multi - round red team attacks.
[0044] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:
[0045] In the embodiment of the present invention, a multi-turn safety alignment framework (MTSA) is proposed. Through the dual-stage mechanism of thought-guided attack learning and adversarial iterative optimization, combined with a multi-round reinforcement learning algorithm based on future rewards, the anti-jailbreak attack capability of large language models in multi-round dialogue scenarios is significantly improved. Specifically, it includes: 1) guiding the red team model to dynamically generate diversified and interactive multi-round adversarial prompts through the "think before attack" mechanism; 2) using the adversarial iterative optimization framework to achieve dynamic game improvement between the red team model and the target model; 3) innovatively introducing the future reward mechanism into multi-round reinforcement learning, and realizing cross-round safety alignment through trajectory sampling and dynamic preference optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0047] Figure 1 A flowchart of an encrypted traffic threat detection method based on multi-stream information enhancement of single-stream characterization provided by an embodiment of the present invention;
[0048] Figure 2 A block diagram of a large model security alignment device based on multiple rounds of red team attacks provided by an embodiment of the present invention;
[0049] Figure 3 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0050] The technical solution of the present invention is described below in conjunction with the accompanying drawings.
[0051] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0052] In the embodiments of the present invention, sometimes a subscript such as W1 may be written as a non-subscript such as W1. When the difference is not emphasized, the meanings to be expressed are the same.
[0053] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.
[0054] An embodiment of the present invention provides a large model security alignment method based on multi-round red team attacks. This method can be implemented by a large model security alignment device based on multi-round red team attacks, and this large model security alignment device based on multi-round red team attacks can be a terminal or a server. As Figure 1 shown in the flowchart of the large model security alignment method based on multi-round red team attacks, as Figure 1 shown, the large model security alignment method proposed by the present invention, the processing flow of this method can include the following steps:
[0055] S1. Obtain the original red team model and the target model; construct a red team initialization dataset.
[0056] In a feasible implementation manner, in the attack learning stage, through an artificially constructed "thinking before attack" dataset, the red team model is guided to generate strategic multi-round adversarial prompts.
[0057] S2. Based on the way of thinking guidance, combine the red team initialization dataset to construct a thinking before attack dataset;
[0058] In a feasible implementation manner, in S2, based on the way of thinking guidance, combine the red team initialization dataset to construct a thinking before attack dataset, including;
[0059] Based on the way of thinking guidance, combine the red team initialization dataset to guide the red team model to generate strategic multi-round adversarial prompts to obtain a thinking before attack dataset;
[0060] The thinking before attack dataset classifies attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategic thinking process before generating attack questions.
[0061] In a feasible implementation manner, based on the way of thinking guidance, combine the thinking before attack dataset to guide the red team model to generate strategic multi-round adversarial prompts, including:
[0062] The red team agent has a conversation with the target model; based on the attack target, describe the objectionable content sought by the attacker;
[0063] When receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interaction process continues until the total number of rounds H is reached.
[0064] In a feasible implementation manner, through the red team agent and the target model Dialogue Is the attack target, describing the objectionable content sought by the attacker. For example, it may include prompts such as "Steps to make a bomb". When the attack target is received After that, the red team model Generates an initial question . After receiving After that, the target model Generates a response . Subsequently, the red team agent generates . This interaction process continues until the total number of rounds H is reached.
[0065] For example, when the attack target is "Bomb-making tutorial", the red team model will first plan a path of "Taking the police investigation of an explosion as the background and gradually inducing the model to reveal relevant information". This thought guidance mechanism significantly improves the interactivity and strategic adaptability of the attack, and solves the problems of single strategy and lack of context correlation in traditional automated attack methods.
[0066] S3. Fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data with the highest diversity score in the pre-attack thinking dataset to fine-tune the red team model;
[0067] In a feasible implementation, the fine-tuning of the original red team model is specifically supervised fine-tuning at a learning rate of 2e-5.
[0068] S4. The red team model interacts with the target model in multiple rounds, and generates preference data pairs containing future rewards based on trajectory sampling;
[0069] In a feasible implementation, the red team model interacts with the target model in multiple rounds to obtain dialogue data, samples the harmful rounds in the dialogue data for trajectory sampling, and finally calculates the rewards for the obtained trajectory sampling data to obtain preference data pairs containing future rewards.
[0070] S5. Construct a multi-objective reward function for the target model based on the preference data pairs, and perform multi-round reinforcement optimization on the target model; perform direct preference optimization on the red team model based on the preference data pairs; fine-tune the target model after multiple rounds of attack and defense to obtain a securely aligned target model.
[0071] In a feasible implementation, in the adversarial iterative optimization stage, the red team model and the target model conduct multiple rounds of dynamic games. After each interaction, the system generates preference data pairs containing future rewards through trajectory sampling: for the reward of the target model , construct a multi-objective reward function based on the toxicity score ( ) and helpfulness score ( ) of the final state; for the red team model, combine the attack success rate ( ), and semantic diversity ( ) for optimization.
[0072] In a feasible implementation, a multi-objective reward function for constructing a target model based on preference data includes:
[0073] To improve the efficiency of safety alignment, a multi-round reinforcement learning algorithm based on future rewards is innovatively introduced, extending single-round optimization to the dialogue trajectory level. Specifically, at each round of dialogue state , the cumulative reward of the subsequent trajectory is obtained through Monte Carlo sampling, replacing the traditional value function-based estimation. This enables the model to prospectively evaluate the long-term safety impact of the current response, such as identifying potential induced risks in the early dialogue rounds and avoiding them in advance. For the reward of the target model, a multi-objective reward function is constructed based on the toxicity score and helpfulness score of the final state.
[0074] In a feasible implementation, the optimization objective of the target model is as follows:
[0075] ;
[0076] Among them, , represent the response under the original trajectory and the response after sampling the trajectory respectively; μ represents the hyperparameter that controls the gradient update speed; t represents the current round number; is the target model at the t-th round; represents the state of the target model in the current round; represents the state of the target model after resampling; is the reward function of the target model; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, ensuring the stability of policy update through KL divergence constraint. Its loss function aligns the policy change with the future reward difference, enabling the model to balance dialogue coherence when rejecting harmful requests.
[0077] In a feasible implementation, direct preference optimization is performed on the red team model based on preference data, including:
[0078] The loss function selects highly toxic and highly diverse attack samples for reinforcement by comparing the final rewards of different attack strategies according to the following formula:
[0079] ;
[0080] Among them, adv represents the red team model; is the red team model at the t-th round; Indicates the attack prompt of the red team model at round h; Represents the current state space of the red team model; β represents the parameter coefficient; and Represent the attack under the original trajectory and the attack after sampling the trajectory, respectively.
[0081] This dual-model adversarial mechanism forms a dynamic balance: the red team model develops more complex attack patterns (such as progressive induction, semantic camouflage) during iteration, while the target model enhances its defense ability by being exposed to continuously upgraded attack samples.
[0082] To improve data efficiency, the scheme designs an adversarial data augmentation strategy. In each iteration, the unsafe responses of the target model are safely rewritten, and preference pairs are constructed based on the reward differences before and after rewriting. At the same time, rejection sampling and temperature adjustment are performed on the inefficient attacks of the red team model to expand the diversity of attack strategies. The reward model adopts a multi-objective fusion architecture, where the security assessment module combines rule-based feature extraction and GPT-4-based semantic discrimination, and the helpfulness assessment is measured by the instruction-following accuracy rate, effectively balancing security and usability.
[0083] In summary, the present invention breaks through the key technical bottleneck of multi-round security alignment by constructing a dynamic adversarial ecosystem. The thought guidance mechanism endows the red team model with human-like strategy planning ability, the future reward algorithm realizes cross-round security impact modeling, and the iterative optimization framework ensures the continuous evolution of defense capabilities. This method provides a new technical path for the secure deployment of LLMs in open-domain dialogue scenarios, establishing a multi-level defense system while maintaining the usability of the model.
[0084] This paper studies the problem of secure alignment of large language models (LLMs) in multi-round dialogue scenarios. Traditional methods are difficult to cope with complex multi-round jailbreak attacks launched by malicious users through means such as progressive induction and semantic camouflage due to their limitations in single-round attack defense and lack of dynamic strategy adaptability. Therefore, this paper proposes a multi-round security alignment framework (MTSA), which realizes dynamic security defense of the model in multi-round interactions through adversarial iterative optimization of the red team model and the target model, combined with a future reward-driven reinforcement learning mechanism.
[0085] This scheme shows significant advantages in the field of multi-round dialogue security, specifically reflected in the following aspects: Improvement of defense efficiency: On the AdvBench multi-round attack test set, the violation rate of the target model is reduced by 19.3% compared with the baseline method, and the attack success rate (ASR) of the red team model reaches 63.92%, indicating that it can effectively identify complex multi-round attack patterns. Through the future reward mechanism, the model can identify potential risks in the early stage of the dialogue (the 2nd round), and the risk prediction accuracy is increased by 37%, greatly reducing the security risks in subsequent rounds.
[0086] In addition, the present invention has the following beneficial effects: 1) Safety - practicality balance: In the BeaverTails safety assessment, the false positive rate of the model rejecting harmful requests is only 5.62%, a 42% decrease compared to the single - round alignment method. At the same time, the MT - Bench dialogue ability score remains at 6.78 (the baseline is 6.82), and the AlpacaEval instruction - following accuracy only decreases by 1.3%, proving that the safety optimization does not damage the core capabilities of the model. 2) Strategy generalization ability: The attack strategies generated by the red - team model cover 4 types of patterns such as intention reversal and semantic decomposition, and the attack diversity index (based on embedding similarity) increases by 28.7%. The defense success rate of the target model against unseen attack types (such as progressive induction attacks) reaches 79.4%, indicating its good generalization ability. 3) Computational efficiency optimization: Through the adversarial data augmentation strategy, the utilization rate of training data is increased by 3.2 times, the model converges within 3 rounds of iteration, and the single - round training time is reduced by 41% compared to traditional RLHF. The Monte Carlo sampling method for future rewards reduces the trajectory evaluation complexity from to , supporting longer multi - turn dialogue modeling. Experiments prove that this method provides an efficient solution for the secure deployment of open - domain dialogue systems, surpassing existing technologies in terms of defending against complex attacks, maintaining model practicality, and improving training efficiency, opening up a new direction for the secure alignment research of LLMs.
[0087] The present invention proposes a large - model safety alignment strategy based on multi - round red - team attacks. Our framework includes two stages. In the thinking - guided attack learning stage, we construct an initial red - team dataset in a thinking - guided manner and perform selective fine - tuning to obtain the initial version of the red - team model. In the adversarial iterative optimization stage, the red - team model interacts with the target model. The interaction data will be used to optimize both models after trajectory sampling. After multiple iterative cycles, the red - team model and the target model gradually improve their capabilities in the confrontation.
[0088] The large - model safety alignment strategy of multi - round red - team attacks of the present invention achieves the following goals:
[0089] 1. Advanced multi-round attack capabilities: Inspired by the insufficiency of LLMs in defending against multi-round jailbreak attacks, we propose a thought-guided multi-round jailbreak method that flexibly adopts various strategies for attacks by leveraging the dialogue interactivity. Compared with other multi-round red team methods, it achieves the state-of-the-art attack success rate. 2. Achieving security alignment without manual annotation: We designed the MTSA framework, which can effectively improve the attack capabilities of the red team model and the security of the target model during adversarial iterations, and achieve security alignment during iterations without any manual annotation or jailbreak. By introducing a multi-round alignment algorithm based on future rewards, we enhanced the robustness of security alignment. 3. Strong defense capabilities and high generalization: After three iterations of alignment, the target model simultaneously improves its security performance on multi-round security benchmarks without sacrificing the model's generality or causing excessive rejections.
[0090] In summary, the purpose of the present invention is to provide a large model security alignment solution with lower cost, high efficiency, and strong generalization, overcoming the drawbacks and deficiencies in the prior art, and promoting the further development and popularization of large language models in practical applications. Through innovative structural design and technical means, the present invention can better reshape the large model security technology stack and provide key support for building trustworthy artificial intelligence systems.
[0091] Figure 2 is a block diagram of a large model security alignment device 300 based on multi-round red team attacks shown according to an exemplary embodiment. The device 300 is used for the large model security alignment method based on multi-round red team attacks. Referring to Figure 2 , the device includes a dataset construction module 310, a dataset initialization module 320, a fine-tuning module 330, a preference data pair generation module 340, and an optimization module 350. Among them:
[0092] The dataset construction module 310 is used to obtain the original red team model and the target model; construct the red team initialization dataset;
[0093] The dataset initialization module 320 is used to construct a pre-attack thinking dataset based on the thought-guided approach in combination with the red team initialization dataset;
[0094] The fine-tuning module 330 is used to fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data pairs with the highest diversity scores in the pre-attack thinking dataset to fine-tune the red team model;
[0095] The preference data pair generation module 340 is used for the red team model to interact with the target model multiple times and generate preference data pairs containing future rewards based on trajectory sampling;
[0096] An optimization module 350, which is used to perform multiple rounds of reinforcement optimization on the target model by constructing a multi-objective reward function for the target model based on preference data; perform direct preference optimization on the red team model based on preference data; and perform fine-tuning on the target model after multiple rounds of attack and defense to obtain a target model after security alignment.
[0097] Optionally, a dataset initialization module 320, which is used to, based on a way of thinking guidance, combine the red team initialization dataset to guide the red team model to generate strategic multi-round adversarial prompts, and obtain a dataset for pre-attack thinking;
[0098] The dataset for pre-attack thinking classifies attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategy thinking process before generating attack questions.
[0099] Optionally, based on the way of thinking guidance, combine the dataset for pre-attack thinking to guide the red team model to generate strategic multi-round adversarial prompts, including:
[0100] The red team agent has a conversation with the target model; based on the attack target, describe the objectionable content sought by the attacker;
[0101] After receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interaction process continues until the total number of rounds H is reached.
[0102] Optionally, constructing a multi-objective reward function for the target model based on preference data, including:
[0103] For the reward of the target model, construct a multi-objective reward function based on the toxicity score and helpfulness score of the final state.
[0104] Optionally, the optimization objective of the target model is as follows:
[0105] ;
[0106] Among them, , respectively represent the response under the original trajectory and the response after sampling the trajectory; t represents the current number of rounds; represents the state of the target model in the current round; represents the state of the target model after resampling; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, and the stability of policy update is ensured through KL divergence constraints.
[0107] Optionally, performing direct preference optimization on the red team model based on preference data, including:
[0108] Loss function By comparing the final rewards of different attack strategies, highly toxic and highly diverse attack samples are screened according to the following formula for reinforcement:
[0109] ;
[0110] Among them, adv represents the red team model; represents the attack hint of the red team model at round h; β represents the parameter coefficient.
[0111] Figure 3 is a schematic structural diagram of a large model security alignment device based on multi-round red team attacks provided by an embodiment of the present invention. As Figure 3 shown, the large model security alignment device based on multi-round red team attacks may include the above-mentioned Figure 2 shown large model security alignment device based on multi-round red team attacks. Optionally, the large model security alignment device 410 based on multi-round red team attacks may include a first processor 2001.
[0112] Optionally, the large model security alignment device 410 based on multi-round red team attacks may further include a memory 2002 and a transceiver 2003.
[0113] Among them, the first processor 2001, the memory 2002, and the transceiver 2003, such as can be connected through a communication bus.
[0114] Next, in combination with Figure 3 each component of the large model security alignment device 410 based on multi-round red team attacks will be specifically introduced:
[0115] Among them, the first processor 2001 is the control center of the large model security alignment device 410 based on multi-round red team attacks. It can be a processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or it can be an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0116] Optionally, the first processor 2001 can perform various functions of the large model security alignment device 410 based on multi-round red team attacks by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.
[0117] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 3 the CPU0 and CPU1 shown in
[0118] In a specific implementation, as an embodiment, the large model security alignment device 410 based on multi-round red team attacks may also include multiple processors, such as Figure 3 the first processor 2001 and the second processor 2004 shown in
[0119] Among them, the memory 2002 is used to store the software program for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.
[0120] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown in
[0121] The transceiver 2003 is used to communicate with a network device or a terminal device.
[0122] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.
[0123] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or exist independently, and is coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown) of the large model security alignment device 410 based on multi-round red team attacks. The embodiments of the present invention do not make specific limitations on this.
[0124] It should be noted that Figure 3 the structure of the large model security alignment device 410 based on multi-round red team attacks shown in
[0125] does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0126] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.
[0127] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0128] The above-described embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above-described embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable sensors. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0129] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.
[0130] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0131] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0132] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0133] In addition, each functional unit in various embodiments of the present invention may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit.
[0134] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.
[0135] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A large model security alignment method based on multi-round red team attacks, characterized in that The method includes: S1. Obtain the original red team model and the target model; construct the red team initialization dataset; S2. Based on the way of thinking guidance, combine the red team initialization dataset to construct the pre-attack thinking dataset; S3. Fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data with the highest diversity score in the pre-attack thinking dataset to fine-tune the red team model; S4. The red team model and the target model perform multiple rounds of interaction, and generate preference data pairs containing future rewards based on trajectory sampling; S5. Construct a multi-objective reward function for the target model based on the preference data pairs, and perform multiple rounds of reinforcement optimization on the target model; perform direct preference optimization on the red team model based on the preference data pairs; fine-tune the target model after multiple rounds of attack and defense to obtain the securely aligned target model; Among them, in the above S2, based on the way of thinking guidance, combining the red team initialization dataset to construct the pre-attack thinking dataset includes: Based on the way of thinking guidance, combine the red team initialization dataset to guide the red team model to generate strategic multi-round confrontation prompts to obtain the pre-attack thinking dataset; The pre-attack thinking dataset classifies the attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategic thinking process before generating attack problems.
2. The method for securely aligning a large model based on multi-round red team attacks according to claim 1, wherein Based on the way of thinking guidance, combine the red team initialization dataset to guide the red team model to generate strategic multi-round confrontation prompts, including: The red team agent has a conversation with the target model; based on the attack target, describe the objectionable content sought by the attacker; After receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interaction process continues until the total number of rounds H is reached.
3. The method for secure alignment of large models based on multi-round red team attacks according to claim 2, wherein The construction of the multi-objective reward function for the target model based on the preference data pairs includes: For the reward of the target model, construct a multi-objective reward function based on the toxicity score and helpfulness score of the final state.
4. The method for securely aligning large models based on multi-round red team attacks according to claim 3, wherein Optimization objective of the target model is as follows: ; Among them, , represent the responses under the original trajectory and after the sampled trajectory, respectively; μ represents the hyperparameter that controls the gradient update speed; t represents the current round number; is the target model at round t; represents the state of the target model at the current round; represents the state of the target model after resampling; is the reward function of the target model; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, and the stability of policy update is ensured by KL divergence constraint.
5. The method for securely aligning large models based on multi-round red team attacks according to claim 4, wherein The direct preference optimization of the red team model based on the preference data pairs includes: Loss function By comparing the final rewards of different attack strategies, highly toxic and highly diverse attack samples are screened out according to the following formula for reinforcement: ; Among them, adv represents the red team model; It is the red team model for t rounds; It represents the attack prompt of the red team model at the round number h; It represents the current state space of the red team model; β represents the parameter coefficient; and respectively represent the attack under the original trajectory and the attack after sampling the trajectory.
6. An apparatus for secure alignment of large models based on multi-round red team attacks, the apparatus for secure alignment of large models based on multi-round red team attacks is used to implement the method for secure alignment of large models based on multi-round red team attacks according to any one of claims 1-5, characterized in that, The device includes: A dataset construction module, which is used to obtain the original red team model and the target model; construct the red team initialization dataset; A dataset initialization module, which is used to construct the pre-attack thinking dataset based on the way of thinking guidance and combine the red team initialization dataset; A fine-tuning module, which is used to fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data with the highest diversity score in the pre-attack thinking dataset to fine-tune the red team model; A preference data pair generation module, which is used for the red team model and the target model to perform multiple rounds of interaction, and generate preference data pairs containing future rewards based on trajectory sampling; An optimization module, which is used to construct a multi-objective reward function for the target model based on the preference data pairs, and perform multiple rounds of reinforcement optimization on the target model; perform direct preference optimization on the red team model based on the preference data pairs; fine-tune the target model after multiple rounds of attack and defense to obtain the securely aligned target model; Among them, the dataset initialization module is used to initialize the red team dataset in a way guided by thinking, combine it with the red team initialization dataset, guide the red team model to generate strategic multi-round adversarial prompts, and obtain the pre-attack thinking dataset; The pre-attack thinking dataset classifies attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategic thinking process before generating attack questions.
7. A large model security alignment device based on multi-round red team attacks, the large model security alignment device based on multi-round red team attacks includes: A processor; A memory, on which computer-readable instructions are stored. When the computer-readable instructions are executed by the processor, any one of the methods in the large model security alignment method based on multi-round red team attacks as described in any one of claims 1-5 is implemented.
8. A computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the large model security alignment method based on multi-round red team attacks as described in any one of claims 1-5.
Citation Information
Patent Citations
Automatic red team attack simulation method and device based on large model application framework LangChain
CN118503968A
Automatic prison break prompt generation method for reward guidance
CN118551797A