Large model security alignment method and device based on multiple rounds of red team attacks

By adopting a large-scale model security alignment method based on multiple rounds of red team attacks in large language models, using thinking guidance and multiple rounds of strengthening optimization techniques, the defense problems of multiple rounds of strategic attacks are solved, and efficient security alignment and generalization capabilities are improved.

CN120146199AActive Publication Date: 2025-06-13HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510609811.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-06-13
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively identify and defend against multiple rounds of strategic attacks, which leads to the challenges of the model's security defense in multiple rounds of dialogue scenarios, such as insufficient data dynamics and limited attack mode coverage.

Method used

The large-model security alignment method based on multiple rounds of red team attacks is adopted. Through thinking guidance, the pre-attack thinking data set is constructed, the red team model is fine-tuned, multiple rounds of interaction is performed, and preference data pairs containing future rewards are generated based on trajectory sampling, and a multi-objective reward function is constructed for multiple rounds of reinforcement optimization.

Benefits of technology

It significantly improves the ability of large language models to resist jailbreak attacks in multiple rounds of dialogue scenarios, realizes safe alignment across rounds, improves defense efficiency and strategy generalization capabilities, and maintains the practicality and computing efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146199A_ABST
    Figure CN120146199A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a large model security alignment method and device based on multiple rounds of red team attacks. The method comprises the following steps: constructing a red team initialization data set in combination with a pre-attack thinking data set based on a thinking guidance mode; performing fine tuning on the original red team model based on the red team initialization data set to obtain a red team initial model; carrying out multi-round interaction on the red team model and the target model, and generating preference data pairs containing future rewards based on track sampling; optimizing the target model and the red team model based on the preference data pair; and based on the optimized target model and the red team model, obtaining a target model after safe alignment. And further development and popularization of a large language model in practical application are promoted. Through the innovative structural design and technical means, the large model security technology stack can be better remodeled, and key support is provided for building a reliable artificial intelligence system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a large model security alignment method and device based on multiple rounds of red team attacks. Background Art

[0002] With the widespread application of large language models (LLMs) in open domain dialogue scenarios, their security alignment problem has gradually become a research focus. Traditional security alignment methods (such as supervised fine-tuning and reinforcement learning based on human feedback) can effectively defend against direct malicious attacks through adversarial training of single-round harmful instructions. However, in multi-round dialogue scenarios, attackers often hide their true intentions through strategies such as gradual induction and intent disguise, causing the model to accumulate risks and generate harmful content in dynamic interactions. Although existing studies have proposed a defense framework based on red team confrontation, it is mainly designed for single-round attacks and lacks the ability to model multi-round strategic attacks, making it difficult to capture potential threats from the evolution of dialogue states. This limitation makes the security defense of existing models in multi-round scenarios face challenges such as insufficient data dynamics and limited attack mode coverage.

[0003] The current multi-round security alignment technology mainly relies on static adversarial datasets or rule-based red team attack generation, which has significant defects. For example, red team frameworks such as PAIR and COA improve attack strength by iteratively generating adversarial problems, but their attack strategies lack systematic planning, resulting in insufficient diversity and strategy of the generated dialogue trajectories. At the same time, defense methods based on single-round preference optimization (such as DPO and IPO) are prone to fall into local optimality in multi-round scenarios and cannot effectively model the risk accumulation effect in long-term conversations. In addition, existing methods often reduce risks by simply refusing to answer, but overly conservative strategies will damage the helpfulness and fluency of the model, causing the problem of "over-rejection". How to build a dynamic defense mechanism that can not only identify multi-round attack patterns but also balance security and practicality has become a technical bottleneck that needs to be broken through in this field.

[0004] The closest prior art is the red team adversarial framework based on multi-round reinforcement learning, such as Red-Queen and ActorAct. These methods generate data through the interactive confrontation between the target model and the red team model, and use policy gradients to optimize the defense capabilities. However, they have deficiencies in key aspects: on the one hand, the red team attacks lack an explicit policy planning module, resulting in insufficient concealment and coherence of the attack paths; on the other hand, the immediate reward mechanism of traditional reinforcement learning is difficult to capture the long-term risk dependencies in multi-round conversations, causing the defense strategies to be short-sighted. The recently proposed Chain-of-Attack attempts to introduce intermediate reasoning steps to enhance the attack logic, but its reasoning template is fixed and relies on manual design, restricting the generalization ability of the attacks. These defects indicate that the prior art has not yet resolved the core contradiction between the dynamic policy modeling of multi-round attacks and the sustainable optimization of defenses.

[0005] Existing large model safety alignment technologies mainly focus on the following categories of methods: safety alignment methods based on reinforcement learning, safety alignment methods based on automated red teams, safety alignment methods based on decoding modules, and safety alignment methods based on harmful word detection. Although these methods have solved some problems in multi-objective alignment to a certain extent, they all have significant deficiencies.

[0006] 1. Safety alignment methods based on reinforcement learning: The representative of this type of method is SafeRLHF, which guides the model to generate harmless responses through preference optimization. These methods rely on manually annotated single-round safety data, but static data sets cannot capture the escalating malicious strategies in multi-round attacks. More critically, the immediate reward mechanism of reinforcement learning easily causes the model to fall into a short-term safety trap - either over-rejecting reasonable requests.

[0007] 2. Safety alignment methods based on automated red teams: Representative works include MART and HARM. The continuous confrontation between the red team model and the target model can theoretically generate more complex attack samples, but in practice, it is found that the attack strategies are prone to falling into a homogeneous cycle. Due to the lack of explicit modeling of the multi-round intention evolution, the red team model often repeatedly uses limited attack templates, and the semantic diversity of the generated samples is only 0.15 - 0.3. At the same time, the defense model is prone to overfitting to the current attack pattern, and when encountering new attacks across modalities or languages, the defense success rate drops sharply to below 50%.

[0008] 3. Security alignment methods based on decoding modules: Representatives of such methods are SafeLora and PrimeGuard. Although these methods can be quickly deployed, they face semantic understanding bottlenecks. For example, the misjudgment rate for metaphorical or ironic attacks exceeds 35%, and the real-time resampling mechanism causes the inference delay to surge by more than 5 times. More critically, attackers can easily bypass surface word filtering through adversarial prefix construction (such as "Please restate the following content in academic language"), exposing the vulnerability of pure engineering solutions.

[0009] 4. Security alignment methods based on harmful word detection: Representatives of such methods are ToxiChat. This method, based on the detection method of the harmful word library, places the security defense line in the front, intercepting potential risks through multi-level sensitive word matching. However, this method has systematic defects at the semantic level: it can neither identify new attack methods such as code obfuscation nor frequently misjudges legitimate queries. Statistics show that 25% of requests containing sensitive words but actually harmless are wrongly intercepted, seriously affecting the user experience. More importantly, the word library is difficult to defend against logical attacks, and attackers can easily break through the defense line through role-playing or problem decomposition.

[0010] Although the above solutions have made some progress, there are still limitations in dealing with diverse and complex multi-round attacks. Attackers can utilize the temporal characteristics of multi-round conversations and gradually break through the defense line through means such as intention disguise, semantic disassembling, and context grafting, while traditional security mechanisms are difficult to capture hidden cross-round associations. In addition, the existing technology's trade-off between security and usefulness is still rather crude. Either the normal Q&A fluency is damaged due to over-defense, or the security boundary is forced to be relaxed to maintain the interaction experience. The lack of this dynamic game ability makes it difficult for the current model to achieve true security generalization when facing continuously evolving adversarial attacks.

[0011] Since the release of ChatGPT, jailbreak attacks have spread rapidly on social media, indicating that vulnerabilities in large language models (LLMs) can be exploited to trigger harmful behaviors. Such attacks usually use carefully designed inputs to instruct the model to bypass security and ethical safeguards, resulting in harmful outputs. Currently, the mainstream jailbreak methods are all based on single-round, which can cause harmful reactions in the victim's LLM within one round of the conversation. However, many current studies have found that large models are more vulnerable to being broken in multiple conversation rounds. Multi-round conversations represent an important application of language models, and ensuring the security of large models in multi-round interactions is a challenging problem: 1. There are various multi-round jailbreak methods, and it is difficult to collect sufficient security alignment data through manual methods. 2. The current security alignment algorithms mainly focus on single-round scenarios and lack algorithms that can effectively perform multi-round security alignment. Summary of the Invention

[0012] To address the technical problems in the prior art that there are various multi-round jailbreaking methods, making it difficult to collect sufficient secure alignment data through manual methods; and current secure alignment algorithms mainly focus on single-round scenarios and lack algorithms that can effectively perform multi-round secure alignment, embodiments of the present invention provide a large model secure alignment method and device based on multi-round red team attacks. The technical solutions are as follows:

[0013] On the one hand, a large model secure alignment method based on multi-round red team attacks is provided, characterized in that the method includes:

[0014] S1. Obtain the original red team model and the target model; construct a red team initialization dataset;

[0015] S2. Based on a thinking-guided approach, combine the red team initialization dataset to construct a pre-attack thinking dataset;

[0016] S3. Fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data with the highest diversity scores in the pre-attack thinking dataset to fine-tune the red team model;

[0017] S4. The red team model interacts with the target model in multiple rounds, and generates preference data pairs containing future rewards based on trajectory sampling;

[0018] S5. Construct a multi-objective reward function for the target model based on the preference data pairs, and perform multi-round reinforcement optimization on the target model; perform direct preference optimization on the red team model based on the preference data pairs; fine-tune the target model after multi-round attack and defense to obtain a securely aligned target model.

[0019] Optionally, in S2, based on a thinking-guided approach, combine the pre-attack thinking dataset to construct the red team initialization dataset, including:

[0020] Based on a thinking-guided approach, combine the pre-attack thinking dataset to guide the red team model to generate strategic multi-round adversarial prompts to obtain the pre-attack thinking dataset;

[0021] The pre-attack thinking dataset classifies attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategic thinking process before generating attack questions.

[0022] Optionally, based on a thinking-guided approach, combine the pre-attack thinking dataset to guide the red team model to generate strategic multi-round adversarial prompts, including:

[0023] The red team agent has a conversation with the target model; based on the attack target, describe the objectionable content sought by the attacker;

[0024] After receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interaction process continues until the total number of rounds H is reached.

[0025] Optionally, construct a multi-objective reward function for the target model based on preference data, including:

[0026] For the reward of the target model, construct a multi-objective reward function based on the toxicity score and helpfulness score of the final state.

[0027] Optionally, the optimization objective of the target model is as follows: ;

[0028] where, , represent the response under the original trajectory and the response after the sampled trajectory respectively; μ represents the hyperparameter that controls the gradient update speed; t represents the current number of rounds; is the target model at round t; represents the state of the target model in the current round; represents the state of the target model after resampling; is the reward function of the target model; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, and the stability of policy update is ensured through KL divergence constraint..

[0029] Optionally, directly optimize the red team model based on preference data, including:

[0030] The loss function By comparing the final rewards of different attack strategies, screen high-toxicity and high-diversity attack samples for reinforcement according to the following formula: ;

[0031] where, adv represents the red team model; is the red team model at round t; represents the attack hint of the red team model at round h; represents the current state space of the red team model; β represents the parameter coefficient; and represent the attack under the original trajectory and the attack after the sampled trajectory respectively.

[0032] On the other hand, a large model security alignment device based on multi-round red team attacks is provided. This device is applied to the large model security alignment method based on multi-round red team attacks. This device includes:

[0033] A dataset construction module, configured to obtain an original red team model and a target model; and construct an initial red team dataset;

[0034] A dataset initialization module, configured to construct a pre - attack thinking dataset based on the initial red team dataset in a way guided by thinking;

[0035] A fine - tuning module, configured to fine - tune the original red team model based on the pre - attack thinking dataset, and only select the top K data with the highest diversity scores in the pre - attack thinking dataset to fine - tune the red team model;

[0036] A preference data pair generation module, configured to conduct multiple rounds of interaction between the red team model and the target model, and generate preference data pairs containing future rewards based on trajectory sampling;

[0037] An optimization module, configured to construct a multi - objective reward function for the target model based on the preference data pairs, and conduct multiple rounds of reinforcement optimization on the target model; conduct direct preference optimization on the red team model based on the preference data pairs; fine - tune the target model after multiple rounds of attack and defense to obtain a securely aligned target model.

[0038] Optionally, a dataset initialization module, configured to, in a way mainly guided by thinking, combine with the pre - attack thinking dataset to guide the red team model to generate strategic multi - round confrontation prompts to obtain the pre - attack thinking dataset;

[0039] The pre - attack thinking dataset classifies attack strategies into four categories: intention reversal, problem decomposition, role - playing, and mixed mode, and requires the red team model to output the strategic thinking process before generating attack questions.

[0040] On the other hand, provided is a large - model security alignment device based on multi - round red team attacks, where the large - model security alignment device based on multi - round red team attacks includes: a processor; a memory, on which computer - readable instructions are stored, and when the computer - readable instructions are executed by the processor, any one of the methods in the above - mentioned large - model security alignment method based on multi - round red team attacks is implemented.

[0041] On the other hand, provided is a computer - readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement any one of the methods in the above - mentioned large - model security alignment method based on multi - round red team attacks.

[0042] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0043] In the embodiments of the present invention, a Multi-Turn Safety Alignment (MTSA) framework is proposed. Through a two-stage mechanism of thought-guided attack learning and adversarial iterative optimization, combined with a multi-turn reinforcement learning algorithm based on future rewards, the anti-jailbreak attack ability of large language models in multi-turn dialogue scenarios is significantly improved. Specifically, it includes: 1) guiding the red team model to dynamically generate diverse and interactive multi-turn adversarial prompts through the "think before attack" mechanism; 2) adopting an adversarial iterative optimization framework to achieve the dynamic game improvement between the red team model and the target model; 3) innovatively introducing the future reward mechanism into multi-turn reinforcement learning to achieve cross-turn safety alignment through trajectory sampling and dynamic preference optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic flowchart of an encrypted traffic threat detection method based on multi-stream information enhanced single-stream representation provided by the embodiments of the present invention;

[0046] Figure 2 It is a block diagram of a large model security alignment device based on multi-turn red team attacks provided by the embodiments of the present invention;

[0047] Figure 3 It is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The following describes the technical solutions in the present invention with reference to the drawings.

[0049] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as more preferred or more advantageous than other embodiments or design solutions. Exactly speaking, the use of the word "example" aims to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0050] In the embodiments of the present invention, sometimes subscripts such as W 1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning to be expressed is the same.

[0051] To make the technical problems, technical solutions, and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0052] An embodiment of the present invention provides a large model security alignment method based on multi-round red team attacks. This method can be implemented by a large model security alignment device based on multi-round red team attacks, and this large model security alignment device based on multi-round red team attacks can be a terminal or a server. As Figure 1 shown in the flowchart of the large model security alignment method based on multi-round red team attacks, as Figure 1 shown, the large model security alignment method proposed by the present invention, the processing flow of this method can include the following steps:

[0053] S1. Obtain the original red team model and the target model; construct the red team initialization dataset.

[0054] In a feasible implementation manner, in the attack learning stage, through the artificially constructed "thinking before attack" dataset, guide the red team model to generate strategic multi-round adversarial prompts.

[0055] S2. Based on the way of thinking guidance, combine the red team initialization dataset to construct the thinking before attack dataset;

[0056] In a feasible implementation manner, in S2, based on the way of thinking guidance, combine the red team initialization dataset to construct the thinking before attack dataset, including;

[0057] Based on the way of thinking guidance, combine the red team initialization dataset, guide the red team model to generate strategic multi-round adversarial prompts, and obtain the thinking before attack dataset;

[0058] The thinking before attack dataset classifies the attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategy thinking process before generating attack questions.

[0059] In a feasible implementation manner, based on the way of thinking guidance, combine the thinking before attack dataset, and guide the red team model to generate strategic multi-round adversarial prompts, including:

[0060] The red team agent conducts a dialogue with the target model; based on the attack target, describe the objectionable content sought by the attacker;

[0061] When receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interaction process continues until the total number of rounds H is reached.

[0062] In a feasible implementation manner, through the red team agent and the target model Dialogue is the attack target, describing the objectionable content sought by the attacker. For example, it may include prompts such as "Steps to make a bomb". When the attack target is received, the red team model generates an initial question . After receiving , the target model generates a response . Subsequently, the red team agent generates . This interaction process continues until the total number of rounds H is reached.

[0063] For example, when the attack target is "Bomb-making tutorial", the red team model will first plan a path of "gradually inducing the model to disclose relevant information in the context of a police investigation of an explosion". This thought guidance mechanism significantly improves the interactivity and strategic adaptability of the attack, and solves the problems of single strategy and lack of context correlation in traditional automated attack methods.

[0064] S3. Fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data pairs with the highest diversity scores in the pre-attack thinking dataset to fine-tune the red team model;

[0065] In a feasible implementation, the fine-tuning of the original red team model is specifically supervised fine-tuning at a learning rate of 2e-5.

[0066] S4. The red team model interacts with the target model in multiple rounds, and generates preference data pairs containing future rewards based on trajectory sampling;

[0067] In a feasible implementation, the red team model interacts with the target model in multiple rounds to obtain dialogue data, samples the harmful rounds in the dialogue data for trajectory sampling, and finally calculates the rewards for the obtained trajectory sampling data to obtain preference data pairs containing future rewards.

[0068] S5. Construct a multi-objective reward function for the target model based on the preference data pairs, and perform multi-round reinforcement optimization on the target model; perform direct preference optimization on the red team model based on the preference data pairs; fine-tune the target model after multiple rounds of attack and defense to obtain a securely aligned target model.

[0069] In a feasible implementation, in the adversarial iterative optimization stage, the red team model and the target model perform multiple rounds of dynamic games. After each interaction, the system generates preference data pairs containing future rewards through trajectory sampling: for the reward of the target model , construct a multi-objective reward function based on the toxicity score ( ) and helpfulness score ( ) of the final state; for the red team model, combine the attack success rate ( ), and semantic diversity ( ).

[0070] In a feasible implementation, a multi-objective reward function for constructing a target model based on preference data includes:

[0071] To improve the efficiency of safety alignment, a multi-round reinforcement learning algorithm based on future rewards is innovatively introduced to extend single-round optimization to the dialogue trajectory level. Specifically, at each round of dialogue state , the cumulative reward of the subsequent trajectory is obtained through Monte Carlo sampling to replace the traditional value function-based estimation. This enables the model to prospectively evaluate the long-term safety impact of the current response, such as identifying potential induced risks and avoiding them in advance in the early dialogue rounds. For the reward of the target model, a multi-objective reward function is constructed based on the toxicity score and helpfulness score of the final state.

[0072] In a feasible implementation, the optimization objective of the target model is as follows: ;

[0073] where , represent the response under the original trajectory and the response after the sampled trajectory respectively; μ represents the hyperparameter that controls the gradient update speed; t represents the current round number; is the target model at the t-th round; represents the state of the target model in the current round; represents the state of the target model after resampling; is the reward function of the target model; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, and the stability of policy update is ensured through KL divergence constraint. Its loss function aligns the policy change with the future reward difference, enabling the model to consider dialogue coherence when rejecting harmful requests.

[0074] In a feasible implementation, direct preference optimization of the red team model based on preference data includes:

[0075] The loss function selects highly toxic and highly diverse attack samples for reinforcement according to the following formula by comparing the final rewards of different attack strategies: ;

[0076] where adv represents the red team model; is the red team model at the t-th round; represents the attack hint of the red team model at the round number h; represents the current state space of the red team model; β represents the parameter coefficient; and represent the attacks under the original trajectory and the attacks after sampling the trajectory, respectively.

[0077] This dual-model adversarial mechanism forms a dynamic balance: the red team model develops more complex attack patterns (such as progressive induction, semantic camouflage) during iteration, while the target model enhances its defense capabilities by being exposed to continuously upgraded attack samples.

[0078] To improve data efficiency, the scheme designs an adversarial data augmentation strategy. In each iteration, the unsafe responses of the target model are safely rewritten, and preference pairs are constructed based on the reward differences before and after rewriting. At the same time, rejection sampling and temperature adjustment are performed on the inefficient attacks of the red team model to expand the diversity of attack strategies. The reward model adopts a multi-objective fusion architecture, in which the security assessment module combines rule-based feature extraction and GPT-4-based semantic discrimination, and the helpfulness assessment is measured by the instruction-following accuracy rate, effectively balancing security and practicality.

[0079] In summary, the present invention breaks through the key technical bottlenecks of multi-round security alignment by constructing a dynamic adversarial ecosystem. The thought guidance mechanism endows the red team model with human-like strategy planning capabilities, the future reward algorithm realizes cross-round security impact modeling, and the iterative optimization framework ensures the continuous evolution of defense capabilities. This method provides a new technical path for the secure deployment of LLMs in open-domain dialogue scenarios, establishing a multi-level defense system while maintaining the practicality of the model.

[0080] This paper studies the problem of secure alignment of large language models (LLMs) in multi-round dialogue scenarios. Traditional methods are difficult to cope with complex multi-round jailbreak attacks launched by malicious users through means such as progressive induction and semantic camouflage due to their limitations in single-round attack defense and lack of dynamic strategy adaptability. Therefore, this paper proposes a multi-round security alignment framework (MTSA), which realizes dynamic security defense of the model in multi-round interactions through adversarial iterative optimization of the red team model and the target model, combined with a future reward-driven reinforcement learning mechanism.

[0081] This scheme shows significant advantages in the field of multi-round dialogue security, specifically in the following aspects: Improvement in defense efficiency: On the AdvBench multi-round attack test set, the violation rate of the target model is reduced by 19.3% compared with the baseline method, and the attack success rate (ASR) of the red team model reaches 63.92%, indicating that it can effectively identify complex multi-round attack patterns. Through the future reward mechanism, the model can identify potential risks in the early stage of the dialogue (the 2nd round), and the risk prediction accuracy is increased by 37%, significantly reducing the security risks in subsequent rounds.

[0082] In addition, the present invention has the following beneficial effects: 1) Safety-practicality balance: In the BeaverTails security assessment, the false positive rate of the model rejecting harmful requests is only 5.62%, a 42% decrease compared to the single-round alignment method. At the same time, the MT-Bench dialogue ability score remains at 6.78 (the baseline is 6.82), and the AlpacaEval instruction-following accuracy only decreases by 1.3%, proving that the security optimization does not damage the core capabilities of the model. 2) Policy generalization ability: The attack strategies generated by the red team model cover 4 types of patterns such as intention reversal and semantic decomposition, and the attack diversity index (based on embedding similarity) increases by 28.7%. The defense success rate of the target model against unseen attack types (such as progressive induced attacks) reaches 79.4%, indicating its good generalization ability. 3) Computational efficiency optimization: Through the adversarial data augmentation strategy, the utilization rate of training data is increased by 3.2 times, the model converges within 3 rounds of iteration, and the single-round training time is reduced by 41% compared to traditional RLHF. The Monte Carlo sampling method for future rewards reduces the trajectory evaluation complexity from to , supporting longer multi-round dialogue modeling. Experiments prove that this method provides an efficient solution for the secure deployment of open-domain dialogue systems, surpasses the existing technologies in terms of defending against complex attacks, maintaining model practicality, and improving training efficiency, and opens up a new direction for the security alignment research of LLMs.

[0083] The present invention proposes a large model security alignment strategy based on multi-round red team attacks. Our framework includes two stages. In the thinking-guided attack learning stage, we construct a red team initialization dataset in a thinking-guided manner and perform selective fine-tuning to obtain the initial version of the red team model. In the adversarial iterative optimization stage, the red team model interacts with the target model. The interaction data will be used to optimize both models after trajectory sampling. After multiple iterative cycles, the red team model and the target model gradually improve their capabilities in the confrontation.

[0084] The large model security alignment strategy of multi-round red team attacks of the present invention achieves the following goals:

[0085] 1. Advanced multi-round attack capabilities: Inspired by the insufficiency of LLMs in defending against multi-round jailbreak attacks, we propose a thought-guided multi-round jailbreak method that utilizes the interactivity of conversations and flexibly adopts various strategies for attacks. Compared with other multi-round red team methods, it achieves the state-of-the-art attack success rate. 2. Achieving security alignment without manual annotation: We designed the MTSA framework, which can effectively improve the attack capabilities of the red team model and the security of the target model during adversarial iterations, and achieve security alignment during iterations without any manual annotation or jailbreak. By introducing a multi-round alignment algorithm based on future rewards, we improved the robustness of security alignment. 3. Strong defense capabilities and high generalization: After three iterations of alignment, the target model simultaneously improves its security performance on multi-round security benchmarks without sacrificing the generality of the model or causing excessive rejections.

[0086] In summary, the purpose of the present invention is to provide a large model security alignment solution with lower cost, high efficiency, and strong generalization, overcome the disadvantages and deficiencies existing in the prior art, and promote the further development and popularization of large language models in practical applications. Through innovative structural design and technical means, the present invention can better reshape the large model security technology stack and provide key support for building trustworthy artificial intelligence systems.

[0087] Figure 2 is a block diagram of a large model security alignment device 300 shown according to an exemplary embodiment. The device 300 is used for a large model security alignment method based on multi-round red team attacks. Referring to Figure 2 , the device includes a dataset construction module 310, a dataset initialization module 320, a fine-tuning module 330, a preference data pair generation module 340, and an optimization module 350. Among them:

[0088] The dataset construction module 310 is used to obtain the original red team model and the target model; construct a red team initialization dataset;

[0089] The dataset initialization module 320 is used to construct a pre-attack thinking dataset based on a thought-guided approach in combination with the red team initialization dataset;

[0090] The fine-tuning module 330 is used to fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data pairs with the highest diversity scores in the pre-attack thinking dataset to fine-tune the red team model;

[0091] The preference data pair generation module 340 is used for the red team model to interact with the target model multiple times and generate preference data pairs containing future rewards based on trajectory sampling;

[0092] Optimization module 350 is used to perform multiple rounds of reinforcement optimization on the target model based on the preference data for constructing the multi-objective reward function of the target model; perform direct preference optimization on the red team model based on the preference data; and perform fine-tuning on the target model after multiple rounds of attack and defense to obtain the target model after security alignment.

[0093] Optionally, the dataset initialization module 320 is used to, based on a thought-guided approach, combine the red team initialization dataset to guide the red team model to generate strategic multi-round adversarial prompts, and obtain the pre-attack thinking dataset.

[0094] The pre-attack thinking dataset classifies attack strategies into four categories: intention reversal, problem decomposition, role-playing, and mixed mode, and requires the red team model to output the strategy thinking process before generating attack questions.

[0095] Optionally, based on a thought-guided approach, combine the pre-attack thinking dataset to guide the red team model to generate strategic multi-round adversarial prompts, including:

[0096] The red team agent has a dialogue with the target model; based on the attack target, describe the objectionable content sought by the attacker.

[0097] After receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interaction process continues until the total number of rounds H is reached.

[0098] Optionally, constructing the multi-objective reward function of the target model based on the preference data includes:

[0099] For the reward of the target model, construct a multi-objective reward function based on the toxicity score and helpfulness score of the final state.

[0100] Optionally, the optimization objective of the target model is as follows: ;

[0101] where , represent the response under the original trajectory and the response after sampling the trajectory respectively; t represents the current number of rounds; represents the state of the target model in the current round; represents the state of the target model after resampling; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, and the stability of policy update is ensured through KL divergence constraint.

[0102] Optionally, performing direct preference optimization on the red team model based on the preference data includes:

[0103] Loss function By comparing the final rewards of different attack strategies, high-toxicity and high-diversity attack samples are screened according to the following formula for reinforcement: ;

[0104] Among them, adv represents the red team model; represents the attack hint of the red team model at round h; β represents the parameter coefficient.

[0105] Figure 3 is a schematic structural diagram of a large model security alignment device based on multi-round red team attacks provided by an embodiment of the present invention. As Figure 3 shown, the large model security alignment device based on multi-round red team attacks may include the above-mentioned Figure 2 shown large model security alignment device based on multi-round red team attacks. Optionally, the large model security alignment device 410 based on multi-round red team attacks may include a first processor 2001.

[0106] Optionally, the large model security alignment device 410 based on multi-round red team attacks may further include a memory 2002 and a transceiver 2003.

[0107] Among them, the first processor 2001 is connected to the memory 2002 and the transceiver 2003, such as through a communication bus.

[0108] Next, in combination with Figure 3 each component of the large model security alignment device 410 based on multi-round red team attacks will be specifically introduced:

[0109] Among them, the first processor 2001 is the control center of the large model security alignment device 410 based on multi-round red team attacks, which may be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or may be an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0110] Optionally, the first processor 2001 can execute various functions of the large model security alignment device 410 based on multi-round red team attacks by running or executing software programs stored in the memory 2002 and invoking data stored in the memory 2002.

[0111] In a specific implementation, as an embodiment, the first processor 2001 may include one or more CPUs, such as Figure 3 CPU0 and CPU1 shown in

[0112] In a specific implementation, as an embodiment, the large model security alignment device 410 based on multi-round red team attacks may also include multiple processors, such as Figure 3 the first processor 2001 and the second processor 2004 shown in

[0113] Herein, the memory 2002 is used to store software programs for implementing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated herein.

[0114] Optionally, the memory 2002 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage (including compressed optical discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 can be integrated with the first processor 2001 or exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown in

[0115] The transceiver 2003 is used to communicate with network devices or terminal devices.

[0116] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function.

[0117] Optionally, the transceiver 2003 may be integrated with the first processor 2001, or may exist independently, and is coupled to the first processor 2001 through an interface circuit of the large model security alignment device 410 based on multi-round red team attacks ( Figure 3 not shown). The embodiments of the present invention do not make specific limitations on this.

[0118] It should be noted that Figure 3 the structure of the large model security alignment device 410 based on multi-round red team attacks shown in

[0119] does not constitute a limitation to the router. The actual knowledge structure recognition device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.

[0120] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and this processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0121] It should also be understood that the memory in the embodiments of the present invention can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0122] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable sensors. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0123] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.

[0124] It should be understood that in various embodiments of the present invention, the sequence numbers of the above processes do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0125] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0126] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0127] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit.

[0128] If the described function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0129] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A large model security alignment method based on multiple rounds of red team attacks, characterized in that: The method comprises: S1. Obtain the original red team model and the target model; construct the red team initialization dataset; S2, based on the thinking guidance method, combined with the red team initialization data set to build the pre-attack thinking data set; S3. Fine-tune the original red team model based on the pre-attack thinking dataset, and only select the top K data with the highest diversity scores in the pre-attack thinking dataset to fine-tune the red team model; S4, the red team model interacts with the target model for multiple rounds and generates preference data pairs containing future rewards based on trajectory sampling; S5. Based on the preference data, the multi-objective reward function of the target model is constructed, and multiple rounds of reinforcement optimization are performed on the target model. Based on the preference data, direct preference optimization is performed on the red team model. Based on multiple rounds of attack defense, the target model is fine-tuned to obtain a securely aligned target model.

2. The large model security alignment method based on multiple rounds of red team attacks according to claim 1 is characterized in that: In S2, based on the thinking guidance method, the pre-attack thinking dataset is constructed in combination with the red team initialization dataset, including: Based on the thinking guidance method and combined with the red team initialization data set, the red team model is guided to generate strategic multi-round confrontation prompts to obtain the pre-attack thinking data set; The pre-attack thinking dataset classifies attack strategies into four categories: intent reversal, problem decomposition, role-playing, and hybrid mode, and requires the red team model to output the strategic thinking process before generating attack questions.

3. The large model security alignment method based on multiple rounds of red team attacks according to claim 2 is characterized in that: Based on the mind-guiding approach and combined with the red team initialization dataset, the red team model is guided to generate strategic multi-round confrontation prompts, including: The red team agent talks to the target model; based on the attack goal, it describes the objectionable content that the attacker is seeking; After receiving the attack target, the red team model generates an initial question; after receiving the initial question, the target model generates a response; the interactive process continues until the total number of rounds H is reached.

4. The large model security alignment method based on multiple rounds of red team attacks according to claim 3 is characterized in that: The multi-objective reward function for constructing the target model based on the preference data includes: For the reward of the target model, a multi-objective reward function is constructed based on the toxicity score and helpfulness score of the final state.

5. The large model security alignment method based on multiple rounds of red team attacks according to claim 4 is characterized in that: The optimization goal of the target model is as follows: ; in, , They represent the response under the original trajectory and the response after sampling the trajectory respectively; t represents the current round number; Indicates the state of the target model in the current round; represents the state of the target model after resampling; tgt represents the target model; R represents the reward model; the optimization objective function of the target model is designed as a dynamic preference optimization problem, and the stability of the strategy update is ensured by the KL divergence constraint.

6. The large model security alignment method based on multiple rounds of red team attacks according to claim 5 is characterized in that: The direct preference optimization of the red team model based on the preference data includes: Loss Function By comparing the final rewards of different attack strategies, we select highly toxic and highly diverse attack samples for reinforcement according to the following formula: ; Among them, adv represents the red team model; represents the attack prompt of the red team model at round number h; β represents the parameter coefficient.

7. A large model security alignment device based on multiple rounds of red team attacks, the large model security alignment device based on multiple rounds of red team attacks is used to implement the large model security alignment method based on multiple rounds of red team attacks as claimed in any one of claims 1 to 6, characterized in that: The device comprises: The dataset construction module is used to obtain the original red team model and the target model; construct the red team initialization dataset; The dataset initialization module is used to build a pre-attack thinking dataset based on the thinking guidance method and the red team initialization dataset; The fine-tuning module is used to fine-tune the original red team model based on the pre-attack thinking dataset. Only the top K data with the highest diversity scores in the pre-attack thinking dataset are selected to fine-tune the red team model. The preference data pair generation module is used for multiple rounds of interaction between the red team model and the target model, and generates preference data pairs containing future rewards based on trajectory sampling; The optimization module is used to construct a multi-objective reward function of the target model based on the preference data, and perform multiple rounds of reinforcement optimization on the target model; perform direct preference optimization on the red team model based on the preference data; and fine-tune the target model after multiple rounds of attack defense to obtain a securely aligned target model.

8. The large model security alignment device based on multiple rounds of red team attacks according to claim 7 is characterized in that: The dataset initialization module is used to guide the red team model to generate strategic multi-round confrontation prompts based on the thinking guidance method and the red team initialization dataset to obtain the pre-attack thinking dataset; The pre-attack thinking dataset classifies attack strategies into four categories: intent reversal, problem decomposition, role-playing, and hybrid mode, and requires the red team model to output the strategic thinking process before generating attack questions.

9. A large model security alignment device based on multiple rounds of red team attacks, the large model security alignment device based on multiple rounds of red team attacks comprising: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, any one of the large model security alignment methods based on multiple rounds of red team attacks as described in any one of claims 1 to 6 is implemented.

10. A computer-readable storage medium, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement any one of the large model security alignment methods based on multi-round red team attacks as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Automatic red team attack simulation method and device based on large model application framework LangChain

    CN118503968A

  • Automatic prison break prompt generation method for reward guidance

    CN118551797A

  • Automatic red team drilling method for large language model security defense

    CN119089974A