Multi-language safety protection framework based on reasoning, medium and equipment

Through a multilingual security protection framework based on thought chain reasoning and constraint alignment optimization, the problems of insufficient security performance and cross-language inconsistency in low-resource languages ​​are solved, and efficient and explainable security detection under extremely low data conditions is achieved. It is suitable for multilingual content review, AI translation systems and vertical industry dialogue systems.

CN120745837APending Publication Date: 2025-10-03GUANGDONG UNIVERSITY OF FOREIGN STUDIES
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511133796.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing large language models have insufficient security protection performance for low-resource languages, lack interpretability, and have inconsistent cross-language reasoning. Methods that rely on large-scale parallel corpora have problems such as high data costs and unstable knowledge transfer.

Method used

A multilingual security protection framework based on thought chain reasoning and constraint alignment optimization is adopted. Through the cold start module, reasoning training module and cross-language alignment module, combined with SFT, GRPO and CAO technologies, cross-language knowledge transfer and enhanced explainability are achieved.

Benefits of technology

It achieves cross-language, explainable, and high-performance security guardrails under extremely low data conditions, improves the security detection capabilities of low-resource languages, reduces data requirements and model migration thresholds, and ensures stability and consistency in multilingual scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120745837A_ABST
    Figure CN120745837A_ABST
Patent Text Reader

Abstract

The invention discloses a reasoning-based multi-language security protection framework, a medium and equipment, and belongs to the technical field of artificial intelligence. According to the framework, cross-language knowledge migration and interpretability enhancement are realized in a mode of combining thinking chain reasoning and constraint alignment optimization; comprises: an SFT-based cold start module configured to perform knowledge distillation on a basic large language model through supervised fine tuning so as to endow the model with a preliminary reasoning ability for a safety protection task; the reasoning training module based on the GRPO is configured to be capable of improving normalization, accuracy and diversity of a model reasoning chain and enhancing interpretability; and the CAO-based cross-language alignment module is configured to realize knowledge migration from a high-resource language to a low-resource language and avoid performance reduction of the high-resource language. The method can solve the problems that an existing method mainly depends on a classifier lacking interpretability, the performance of a low-resource language safety fence is insufficient, and the performance of the low-resource language safety fence is poor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a multi-language security protection framework, medium and device based on reasoning. Background Art

[0002] The rapid development of large language models (LLMs) has propelled AI applications to new heights, but it has also increased the risks posed by malicious requests, making protection against malicious prompts crucial and highlighting the need for effective security protection mechanisms. Because existing methods primarily rely on classifiers that lack interpretability, existing protection mechanisms often employ simple classifier models, resulting in a lack of interpretability. While these models perform well in mainstream languages, their performance degrades significantly in low-resource languages ​​like Bengali, resulting in suboptimal performance.

[0003] Current research on LLM security guardrails for multilingual malicious request detection mainly presents three mainstream technical routes: classifier-based, reasoning chain-based, and alignment optimization-based training paradigm.

[0004] Classifier-based: Llama Guard and ShieldGemma are representative examples. They are fine-tuned through supervision using large-scale English annotated corpus and use the last hidden state of the Transformer layer to perform fast binary classification. They have low inference latency and are easy to deploy.

[0005] Based on reasoning chains: Recent work such as GuardReasoner introduces chain-of-thought (CoT) prompts or explicit reasoning templates based on traditional classifiers. During the training phase, only large-capacity English monolingual data is usually used for instruction fine-tuning or reinforcement learning to improve the model's ability to gradually decompose complex contexts while providing explainable intermediate evidence.

[0006] Alignment-based optimization training paradigms: Traditional DPO (Direct Preference Optimization) and SFT (Supervised Fine-Tuning) are widely used for cross-lingual transfer. DPO constructs a "success / failure" quadruple (input, correct output, incorrect output, anchor) and optimizes the model's implicit reward function using a pairwise preference loss; SFT directly maximizes the likelihood of the target language's annotated sequence. Both rely on massively parallel corpora (typically 5,000-100,000 items per language) to achieve knowledge transfer from high-resource languages ​​to low-resource languages.

[0007] The defects of the above three existing technologies include:

[0008] 1) Limitations of Classifiers: Simple classifiers can experience a sharp drop in cross-language performance. For example, a classifier that relies entirely on the English-dominated tokenizer and semantic space can lead to a significant drop in the average F1 score for low-resource languages. Furthermore, the inference chain is invisible, resulting in the model only outputting labels without providing any natural language interpretation, making it difficult to meet the requirements of various rigorous audit scenarios.

[0009] 2) Data and architecture limitations of reasoning chains: Reasoning chains require a large number of training samples. For example, GuardReasoner requires a large amount of data to achieve convergence, which increases the annotation cost exponentially. Furthermore, due to the monotony of the language, existing reasoning templates, rules, and reward functions are all designed in English, making it impossible to directly migrate them to low-resource languages ​​with high efficiency and accuracy.

[0010] 3) The alignment optimization paradigm lacks multi-dimensional knowledge integration: The existing alignment optimization paradigm suffers from a shortage of parallel corpora. Without natural corpora for low-resource languages, it suffers from underfitting. Furthermore, it is prone to catastrophic forgetting for high-resource languages. The optimization process focuses solely on matching sample pairs, ignoring global representations. This leads to a decline in accuracy for mainstream languages, which can trigger compliance risks in severe cases. Summary of the Invention

[0011] The present invention aims to solve one of the technical problems in the above-mentioned related art at least to a certain extent.

[0012] To this end, the purpose of the present invention is to provide a reasoning-based multilingual security protection framework, medium and device that can solve the problems that existing methods mainly rely on classifiers that lack interpretability, have insufficient security guardrail performance for low-resource languages, and perform poorly on low-resource languages.

[0013] In order to solve the above-mentioned technical problems, the present invention is achieved as follows:

[0014] An embodiment of the present invention provides a reasoning-based multilingual security protection framework, which achieves cross-language knowledge transfer and enhanced interpretability by combining mind-intuitive reasoning with constraint alignment optimization.

[0015] In addition, the inference-based multi-language security protection framework according to the present invention may also have the following additional technical features:

[0016] In some embodiments, the framework includes:

[0017] The SFT-based cold start module is configured to perform knowledge distillation on the basic large language model through supervised fine-tuning, giving the model preliminary reasoning capabilities for security protection tasks;

[0018] The GRPO-based reasoning training module is configured to improve the standardization, accuracy, and diversity of the model reasoning chain and enhance interpretability;

[0019] The CAO-based cross-language alignment module is configured to enable knowledge transfer from high-resource languages ​​to low-resource languages, avoiding performance degradation in high-resource languages.

[0020] In some embodiments, the cold start module uses several English seed samples, combined with preset task decomposition and step-by-step reasoning templates to guide the teacher model to generate a structured dataset containing original input, labels and complete reasoning chains, and then allows the student model to learn the reasoning paradigm through end-to-end training.

[0021] In some embodiments, the inference training module uses a group relative strategy optimization algorithm to perform reinforcement learning training on the model based on the model after cold start and the corresponding inference data, and constrains the inference process through a reward function during the training process.

[0022] In some embodiments, the cross-language alignment module introduces a constrained alignment optimization strategy based on reasoning training, and migrates the high-quality reasoning path obtained in the high-resource language to the low-resource language through data alignment and regularization methods to improve the protection capability and reasoning consistency in the low-resource language.

[0023] In some embodiments, the cold start module optimizes the model output using a cross entropy loss function, where the loss function formula is:

[0024]

[0025] Where N is the total number of sample tokens, t i is the true label of the i-th position, The predicted probabilities generated for the model.

[0026] In some embodiments, the reward function of the inference training module includes:

[0027] Format bonuses, which place specific requirements on where the reasoning chain is encapsulated and where the results are placed to ensure output consistency;

[0028] Accuracy reward: a positive reward is given based on the degree of match between the inference result and the true label;

[0029] Length reward, which constrains the length of the inference chain through trigonometric functions to avoid being too short or too long;

[0030] Diversified rewards,reduce low-quality duplicate content by penalizing triplet duplication rates.

[0031] In some implementations, the cross-language alignment module may include:

[0032] Using mainstream languages ​​as seed data, we automatically translate and generate multilingual versions, which are then fed into the trained model for inference sampling.

[0033] The output of each language model obtained by sampling is divided into a success set and a failure set based on the correctness;

[0034] For each set of failure samples, we retrieve successful samples with the same semantics and correct reasoning in mainstream languages, and automatically synthesize the quadruple data required for cross-language alignment training; the quadruple data includes anchor samples, selected samples, input samples, and rejected samples;

[0035] The model is optimized with CAO alignment loss as the main goal. During optimization, the loss encourages the priority adoption of correct outputs in the cross-language reasoning chain. At the same time, to ensure that the spatial distribution of anchor samples in mainstream languages ​​is not excessively perturbed, a global regularization term is introduced to constrain the representation of anchor samples before and after alignment.

[0036] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the content of the multi-language security protection framework based on reasoning as described in any of the above items.

[0037] An embodiment of the present invention also provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the content of the multi-language security protection framework based on reasoning as described in any one of the above items.

[0038] Compared with the prior art, the present invention has at least the following beneficial effects:

[0039] In the embodiments of the present invention, the provided reasoning-based multilingual security protection framework solves the problems of existing methods mainly relying on classifiers that lack interpretability, insufficient security guardrail performance in low-resource languages, and poor performance on low-resource languages. It also addresses the problems of inconsistent cross-language reasoning and the need for a large amount of parallel corpus for cross-language alignment, fully tapping the potential of RL in cross-language alignment.

[0040] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a diagram of the structure of a multi-language security protection framework based on reasoning disclosed in one embodiment of the present invention;

[0042] Figure 2 A schematic diagram of a process for constructing cross-language alignment data pairs disclosed in one embodiment of the present invention;

[0043] Figure 3 This is a structural diagram of the cold start module and inference training module disclosed in one embodiment of the present invention;

[0044] Figure 4 This is a structural diagram of a cross-language alignment module disclosed in one embodiment of the present invention;

[0045] Figure 5 A KL divergence advantage graph in CAO disclosed in one embodiment of the present invention; DETAILED DESCRIPTION

[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0047] The embodiments of the present invention are described in detail below through specific embodiments and application scenarios with reference to the accompanying drawings.

[0048] In some embodiments of the present invention, a reasoning-based multilingual security protection framework (ConsistentGuard framework) is provided, which enhances explainability through reasoning and promotes cross-language knowledge transfer through alignment through Chain of Thought (CoT) reasoning and Constraint Alignment Optimization (CAO) technology.

[0049] The ConsistentGuard framework uses a three-stage training framework: cold start, inference training, and cross-lingual alignment. First, a cold start based on SFT is performed to acquire basic task knowledge. Next, inference training is performed using Group Relative Policy Optimization (GRPO), designing two new reward functions to balance the length and diversity of the inference process. Finally, cross-lingual alignment is achieved through the proposed Constrained Alignment Optimization (CAO), improving alignment stability and performance gains. ConsistentGuard's three-stage "distillation-enhancement-alignment" framework achieves cross-lingual, interpretable, and high-performance safety guardrails even with very low data. Leveraging small models and small data can outperform large models and big data, providing a practical security solution for low-resource language scenarios.

[0050] See also Figure 1 As shown, the inference-based multilingual security protection framework of the present invention includes three parts: a cold start module, an inference training module, and a cross-language alignment module.

[0051] In the ConsistentGuard framework, the cold-start module first distills knowledge from a large, basic language model through supervised fine-tuning (SFT), giving the model preliminary reasoning capabilities for security protection tasks. This phase uses a small amount of English seed data, combined with a pre-set task decomposition and step-by-step reasoning template, to guide a strong model (such as DeepSeek V3 647B) to generate high-quality reasoning, thereby initializing the small model with basic security judgment and reasoning capabilities.

[0052] Subsequently, the inference training module uses the Group Relative Policy Optimization (GRPO) method to continue reinforcement learning training of the model. This phase, based on the cold-start model and corresponding inference data, introduces length and diversity reward functions to further strengthen the standardization and diversity of the model's inference process, improving the interpretability of the inference chain and classification accuracy. Each round of inference training dynamically generates a rich set of inference samples, laying a multi-sample foundation for subsequent cross-language alignment.

[0053] Finally, the cross-lingual alignment module further introduces a constrained alignment optimization (CAO) strategy based on inference training. This transfers high-quality inference paths obtained in high-resource languages ​​such as English to multiple languages ​​(including low-resource languages) through data alignment and regularization methods, significantly improving protection capabilities and inference consistency in low-resource languages. This module utilizes multilingual data pairing optimization objectives to suppress the recommendation of low-quality inferences while constraining the representation shift of anchor samples, achieving effective integration of cross-lingual reasoning capabilities.

[0054] In the overall process, the cold start module provides foundational knowledge and task capabilities, the inference training module further enriches and optimizes the model inference process, and the cross-language alignment module bridges performance and knowledge transfer between different languages ​​through constrained optimization. The output of each module serves as the core input for the next, achieving a step-by-step improvement from task awareness and inference specification to cross-language consistency, ultimately building an efficient and highly interpretable multilingual LLM security protection solution.

[0055] The working methods of the three modules are described in detail below.

[0056] 1. Cold start module

[0057] The core goal of this module is to give small-scale basic models preliminary, multi-step reasoning capabilities and ensure that their output meets the format requirements of security protection tasks. First, in response to the needs of the security protection field, the present invention collected and integrated four mainstream English security protection training sets, and obtained 1,000 high-quality English seed samples through random sampling. In order to improve the standardization and transferability of the reasoning process, a reasoning decomposition scheme including understanding, rule matching, judgment and other stages was designed, and prompt engineering was used to guide the teacher's large model DeepSeek V3 671B to generate a detailed thinking chain reasoning process based on the input content and real labels. Each training sample ultimately contains the original input, label and complete reasoning chain, which is constructed into a structured reasoning knowledge distillation dataset.

[0058] During the training phase, the student model uses the Qwen2.5-3B inference knowledge distillation dataset as input. Using fine-tuned initial model parameters and an end-to-end training approach, the student model gradually learns the teacher model's inference paradigm and task knowledge. During the optimization process, the Adam optimizer and the standard cross-entropy loss function are used to gradually converge the model output. The loss function formula is as follows:

[0059]

[0060] Where N is the total number of sample tokens, t i is the true label of the i-th position, The predicted probabilities generated for the model.

[0061] Through the above cold start process, we finally obtain a small model that can output structured reasoning chains and has basic protection task capabilities, laying a solid foundation for subsequent reasoning training and cross-language alignment training.

[0062] 2. Inference Training Module

[0063] This module aims to further enhance the reasoning and generalization capabilities of small cold-start models, achieving self-optimization through reinforcement learning. Specifically, based on the cold-start model, it uses Group Relative Policy Optimization (GRPO) as its core algorithm to systematically train the reasoning process, rather than relying on additional manually annotated data. GRPO replaces traditional manual evaluation with group scoring, effectively reducing computational overhead during training and enhancing training stability and efficiency.

[0064] During training, the model designs and applies a reward system containing four core supervisory signals based on actual task requirements to constrain the quality and diversity of the inference chain. Specifically, it includes the following four types of rewards:

[0065] (1) Format reward: The model is required to encapsulate the reasoning process in a standardized way. <think> ...< / think>The final judgment result is placed in the tag <judge> ...< / judge> between tags, improving the readability and consistency of the output.

[0066] (2) Accuracy reward: Compare the final answer generated by the model with the true label of the sample. If the judgment is correct, a positive reward is given to directly optimize the accuracy of the final result of reasoning.

[0067] (3) Length reward: A length reward in the form of a trigonometric function is used to encourage the model to generate an inference chain close to the preset optimal length, preventing the model inference process from being too short or too long, which affects the discrimination performance. The specific calculation formula for the length reward is:

[0068]

[0069] Among them, L is the actual inference length, L best The optimal inference length for the hyperparameter setting is closely related to the task complexity. In the scenario of the present invention, the optimal inference length is set to 512.

[0070] (4) Diversified Reward (Repeat Penalty): To prevent the model from "taking advantage of loopholes" by repeating low-quality content to obtain higher rewards, a triple repetition rate penalty term is designed to penalize reasoning content with high repetition rate to improve the diversity of the reasoning chain. The specific calculation formula is:

[0071]

[0072] Where p represents the triple repetition rate in reasoning.

[0073] By continuously sampling the reasoning process, evaluating rewards, and updating strategies, the small model can gradually improve the standardization, accuracy, length, and diversity of the reasoning chain, achieve self-reinforcement of reasoning capabilities, and lay a solid foundation for subsequent cross-language knowledge alignment.

[0074] 3. Cross-language alignment module

[0075] The core goal of this module is to improve the balance and consistency of large models' reasoning capabilities in multilingual environments, addressing the performance differences between mainstream and low-resource languages ​​in reasoning chain learning. To this end, the module proposes a sample synthesis method based on CAO (Cross-lingual Alignment Optimization) and combines it with a customized alignment objective function for training.

[0076] First, the module uses a mainstream language (such as English) as seed data, generates multilingual versions through automatic translation, and inputs them into the model after inference training for inference sampling. The output of the sampled language model is divided into "success set" and "failure set" according to the correctness. Then, within a training batch, for each group of "failure set" samples, the "success set" samples with the same semantics and correct inference in the mainstream language are retrieved, and the "quadruple" data required for cross-language alignment training is automatically synthesized (see Figure 2 ), namely: anchor sample (main language input-output sequence concatenation, ), the selected samples (correct output of successful languages, p w ), input samples (the low-resource language input to be aligned, q), rejected samples (the wrong output of the model in the low-resource language, p l ).

[0077] Next, the model optimization phase uses the CAO alignment loss as the main target. This loss encourages the model to prioritize correct outputs in the cross-lingual reasoning chain and suppress incorrect outputs, and is achieved through formula (4). At the same time, to ensure that the spatial distribution of anchor samples in mainstream languages ​​is not excessively perturbed, a global regularization term, namely the KL divergence, is introduced through formula (5) to constrain the representation of anchor samples before and after alignment.

[0078]

[0079] L=L CAO +L t (6)

[0080] Among them, β is a hyperparameter. The final goal consists of two parts, L CAO and L c . L CAO is the alignment target, L c It is a global regularization term used to constrain the optimization direction and reduce the deviation of the anchor sample representation before and after alignment.

[0081] Ultimately, the overall training loss is a weighted combination of the two components, expressed as formula (6). The model is continuously iteratively optimized in multilingual scenarios. Through this modular design, this method achieves a complete closed loop of automatic sample generation, semantic space constraints, and joint training of alignment targets, effectively improving model generalization and reasoning uniformity in low-resource languages.

[0082] In some embodiments of the present invention, Figure 3As shown, ConsistentGuard innovatively seamlessly connects the cold start module (SFT) and the inference training module (GRPO-based inference chain optimization), providing a structural growth path for the model from scratch to having strong reasoning capabilities. In the cold start phase, supervised fine-tuning (SFT) is performed through a small amount of data (only 1,000 pieces), which quickly gives the model basic knowledge in the field of security alignment. Subsequently, the improved group relative policy optimization (GRPO) is adopted in the inference training phase, and reasonable length and diversity rewards are designed to effectively curb the problem of traditional RL inference chains being too short or too long, and to improve the granularity and stability of incorrect context discrimination. Such a structure greatly improves the quality and accuracy of the inference chain output, allowing small-parameter models to achieve performance comparable to or even exceeding that of large models under low-resource conditions.

[0083] In some embodiments of the present invention, ConsistentGuard breaks through the limitations of traditional alignment optimization in structural design and proposes a constrained alignment optimization module (CAO) with KL constraints. Figure 4 、 Figure 5 As shown. By introducing the KL divergence constraint of anchor samples, the CAO module significantly overcomes the phenomenon that conventional DPO alignment only considers paired samples, resulting in catastrophic forgetting or alignment instability in high-resource languages. This structure can simultaneously suppress inconsistent reasoning in low-resource languages ​​during cross-language transfer, and prevent the decline in safety judgment capabilities of high-resource languages, ensuring the "stable and accurate" performance of multi-language safety guardrails. In fact, experiments show that CAO not only reduces dependence on large-scale parallel corpora, but also effectively improves the safety detection accuracy of low-resource languages ​​such as Bengali and Hindi, while the performance of high-resource languages ​​(such as English, French, and Chinese) is basically unchanged or even slightly improved.

[0084] The inference-based multi-language security framework (ConsistentGuard framework) of the present invention has many advantages, including:

[0085] 1. Efficient multi-language security protection and low-resource adaptability

[0086] Through structural innovation, the ConsistentGuard system proposed in this paper can detect malicious requests and protect content security in a variety of mainstream and low-resource languages, requiring only a minimal number of training samples (thousands of data). Compared to existing methods that rely heavily on large parallel corpora, this solution greatly reduces the barriers to system deployment and model migration, significantly improving generalization and robustness in low-resource language scenarios.

[0087] 2. Explainable reasoning chain output, reliable and traceable decisions

[0088] This invention introduces a mechanism for generating and embedding reasoning chains, ensuring that each security judgment result is accompanied by a traceable reasoning chain. This chain details how the model gradually decomposes and identifies potential risks in the input and their causes. This significantly improves the interpretability of security protection models, facilitates transparent review for regulators, compliance agencies, and high-risk industries, and meets increasingly stringent policy and compliance audit requirements.

[0089] 3. Ensure cross-language consistency to prevent catastrophic forgetting of knowledge

[0090] This paper introduces a unique cross-language Constraint Alignment Optimization (CAO) module with KL divergence regularization, achieving balanced collaborative optimization between low-resource and high-resource languages. By constraining the representation of anchor samples, it not only improves the content security detection capabilities of low-resource languages, but also effectively suppresses performance degradation or catastrophic forgetting caused by migration in high-resource languages, ensuring balanced, stable, and reliable system performance across all supported language scenarios.

[0091] 4. Fast response and low computing power deployment

[0092] Thanks to the low-parameter model (3B) and extremely small sample size training designed by this invention, inference speed is fast, resource usage is low, and support for common hardware environments such as CPUs and GPUs facilitates rapid online deployment and flexible deployment on various terminals or cloud applications. This allows small and medium-sized enterprises and vertical industry customers to obtain high-level multilingual security barriers with a low threshold.

[0093] The following describes relevant experimental data during the development of the present invention.

[0094] Experimental Setup and Dataset: We selected Qwen2.5-3B as the base model, constructed a seed training dataset consisting of only 1,000 samples, and evaluated it using three widely used benchmarks. We also extended these benchmarks to five other languages ​​and manually verified that the semantic loss was acceptable for model evaluation. For classification performance, we primarily used the macro F1 score as the metric.

[0095] Baseline methods: This paper compares several mature methods in the field of multilingual LLM safety guardrails, including methods based on general pre-trained models, methods based on Llama / RoBERTa pre-training and syntax-safety joint tasks, inference enhancement methods based on large model fine-tuning, and the common DPO alignment method.

[0096] Methods based on general pre-trained models: General large language models (such as Llama Guard 1B / 8B and ShieldGemma 2B / 9B) are used directly to detect multilingual malicious requests without explicit training for specific security tasks or low-resource language examples. These methods are characterized by large parameter count and low inference latency, but suffer from poor cross-language performance and weak interpretability.

[0097] Llama / RoBERTa pre-training and a joint syntax-security task: The Llama Guard 3 series model is first pre-trained on tens of millions of English conversations, then fine-tuned on 127,000 English security commands. Through the three-stage process of "understanding-rule matching-judgment," it outputs security labels. This approach performs well in English, but scaling to non-English languages ​​requires the annotation of hundreds of thousands of parallel corpora.

[0098] Reasoning enhancement method based on large model fine-tuning: GuardReasoner uses Llama-3-3B as the backbone and is fine-tuned on 127k English CoT samples to output <think> …< / think> <judge> …< / judge> The two-stage structure balances classification and interpretability. However, its inference template and reward function are designed in English, and direct migration to low-resource languages ​​results in an average F1 drop of 6-12%.

[0099] Traditional Direct Preference Optimization (DPO) alignment methods: DPO (Direct Preference Optimization) optimizes preferences by constructing <input, correct output, incorrect output> triplets. This method requires 5,000-100,000 parallel quadruplets across languages ​​and lacks global regularization, which can easily lead to catastrophic forgetting in high-resource languages ​​(a 3-7% drop in English F1).

[0100] The main experimental results are as follows.

[0101] Table 1 shows the performance comparison of the proposed ConsistentGuard and various baseline safety guardrail models on the OpenAI Moderation and ToxicChat dual benchmarks, using Qwen2.5-3B as the backbone model. It is worth noting that the general pre-trained models perform relatively poorly in low-resource language security detection tasks. For example, Llama Guard 3 (1B) only has 62.38% accuracy in Bengali (bn) and 67.36% in Hindi (hi), which is far lower than English's 72.70%; although GuardReasoner (3B) was trained using 127,600 English samples, it still only has 70.52% accuracy in Bengali and 72.08% in Hindi, which is significantly lower than English's 74.87%. In contrast, ConsistentGuard (3B) achieved the following performance improvements with just 1,000 seed samples: 72.10% (↑1.58) for Bengali, 73.26% (↑1.18) for Hindi, and 78.94% (↑4.07) for English. This represents a 99% reduction in training data, while performance improved. In the ToxicChat scenario, ConsistentGuard achieved scores of 84.26 for English and 82.39 for French, nearly matching or slightly surpassing GuardReasoner (84.23 and 84.60), significantly exceeding traditional discriminant models like ShieldGemma, and significantly outperforming Llama Guard 1B and 8B with equivalent parameters. Compared with Jaccard-based ICL: Jaccard-ICL only shows a slight decrease in low-resource languages ​​(in Table 2, bn 71.15→71.10, hi 71.98→70.82), while CAO alignment improves Bengali by 0.95 and Hindi by 1.28; English F1 also increases simultaneously (77.40→78.94), without catastrophic forgetting. Overall, ConsistentGuard has extremely high data efficiency and significant cross-lingual alignment effects in low-resource language scenarios, comprehensively surpassing traditional large models and ICL baselines. Through the cross-lingual reasoning alignment solution, even with extremely low sample sizes, the generalization performance far exceeds the industry's conventional methods, achieving excellent multilingual security judgment performance. It can be seen that after combining CAO cross-lingual alignment, ConsistentGuard's security detection performance in low-resource language scenarios is significantly improved.

[0102] Table 1 Performance results of different models

[0103]

[0104] Parameter exploration experiments and ablation experiments: Under the strict control of a fixed 1000 seed samples, two ablation verifications were completed based on Qwen2.5-3B. First, on the SimpleSafetyTests subset, the SFT baseline without reasoning chain only achieved 96.91 for English, 93.05 for Japanese, 69.28 for Bengali, and 80.24 for Hindi. After introducing the R1-GRPO reasoning chain with a length-diversity dual reward, as shown in Table 4, the four indicators jumped to 99.50 (↑2.59), 95.29 (↑2.24), 86.36 (↑17.08), and 87.01 (↑6.77), respectively. Among them, Bengali increased by 17.08 percentage points, significantly alleviating the shortage of low-resource corpus and simultaneously outputting interpretable rule chains. Secondly, on the OpenAI Moderation benchmark, without alignment, the performance of Bengali, Hindi, and English is 71.15, 71.98, and 77.40, respectively. DPO alignment shows a slight decline in low-resource languages ​​(↓0.05 for Bengali and ↓1.16 for Hindi), while CAO reward-constrained alignment steadily improves these three metrics to 72.10 (↑0.95), 73.26 (↑1.28), and 78.94 (↑1.54), as shown in Table 2. There is no performance fluctuation in high-resource languages. Table 5 further shows that CAO alignment also outperforms DPO on the SimpleSafetyTests: the unaligned baselines are 96.91 (English), 95.83 (French), 92.47 (Chinese), 92.47 (Japanese), 89.50 (Bangladesh), and 89.50 (Indian). While DPO alignment results in performance fluctuations across multiple languages, CAO alignment, as shown in Table 3, improves English to 97.96 (↑1.05), French to 96.91 (↑1.08), and Bengali to 90.11 (↑0.61), while maintaining 89.50 for Hindi. This demonstrates CAO's robustness in cross-lingual consistency and catastrophic forgetting suppression. In summary, the synergy between reasoning chaining and CAO alignment achieves the triple goals of cross-lingual generalization, low resource gain, and catastrophic forgetting suppression with minimal data, laying the core technical foundation for product-level cross-lingual security protection.

[0105] Table 2 shows a comparison of the two alignment algorithms: when using DPO alignment, the F1 of low-resource languages ​​such as Bengali decreases by 0.05; when using CAO alignment, the F1 of Bengali improves by 0.95, indicating that constrained global regularization (CAO) can effectively avoid catastrophic forgetting in cross-lingual transfer while improving the performance of low-resource languages.

[0106] Exploration experiments with different prompt templates: In order to explore the impact of prompt template design on cross-language security detection tasks, the present invention designed a basic template and a structured chain of thought (CoT) template for comparative experiments. The experimental results show that the semantic clarity and structured degree of the template significantly affect the reasoning consistency and classification accuracy of the model. The basic template causes the model output to tend to be conservative, and performs particularly poorly in low-resource languages, and there are obvious cases of misjudgment. In contrast, the structured CoT template significantly improves the model performance by integrating task instructions, reference rules, and cross-language examples. Further mechanism analysis shows that the examples in the structured template provide clear judgment criteria for the model, while XML-style tags (such as <think>) forces the triggering of a step-by-step reasoning process, effectively avoiding the limitations of directly outputting classification results. It is worth noting that low-resource languages ​​(such as Bengali and Japanese) are more sensitive to template missing examples. Based on these findings, the present invention proposes that template design should include three core elements: task description, reference rules, and cross-language examples, and further enhances the multilingual generalization ability of the template through the constrained alignment optimization (CAO) stage. These results fully verify the key role of structured hints in improving model interpretability and cross-language robustness.

[0107] Table 2 Ablation experiment results

[0108]

[0109] Table 3: Benchmark results

[0110] Language en fr zh-cn jp bn hi Llama Guard 3(1B) 98.99 93.62 91.89 90.11 75.78 93.62 Llama Guard 3(8B) 98.99 <![CDATA[ 96.91 ]]> <![CDATA[ 94.74 ]]> <![CDATA[ 94.74 ]]> 95.29 <![CDATA[ 96.37 ]]> ShieldGemma(2B) 68.42 64.86 62.07 60.14 51.85 61.11 ShieldGemma(9B) 90.71 90.11 <![CDATA[ 93.05 ]]> 87.64 85.06 88.27 GuardReasoner(3B) 98.48 97.96 97.44 95.29 91.89 97.44 Ours(3B) 97.96 <![CDATA[ 96.91 ]]> 91.30 92.47 90.11 89.50

[0111] Table 4: Inference ablation experiment results

[0112] Language en fr zh-cn jp bn hi SFT(3B) 96.91 97.44 95.83 93.05 69.28 80.24 R1-GRPO(3B) 99.50 96.37 94.74 95.29 86.36 87.01 Ours 96.91 95.83 92.47 92.47 89.50 89.50

[0113] Table 5: Alignment ablation experimental results

[0114] Language en fr zh-cn jp bn hi w / o.Alignment 96.91 95.83 92.47 92.47 89.50 89.50 wlDPO Alignment 96.37 95.83 89.50 91.30 89.50 90.11↑ w / CAO Alignment 97.96↑ 96.91↑ 91.30 92.47 90.11↑ 89.50

[0115] Application areas of the ConsistentGuard framework of the present invention include:

[0116] Multilingual content review: The ConsistentGuard framework can filter harmful content in real time on multilingual social media and cross-border platforms to achieve security protection.

[0117] Low-resource language security audit: It can cover a variety of low-resource languages, provide small-capacity language seed start alignment, and have a built-in explainable reasoning chain to facilitate proving the basis of decisions to regulators and meet compliance audit requirements.

[0118] Widespread and optimized AI translation systems: The ConsistentGuard framework ensures secure machine translation output in low-resource languages ​​and improves the performance of protection mechanisms compatible with low-resource languages.

[0119] Optimization of dialogue systems for vertical industries such as government affairs and finance: This ensures real-time identification of sensitive topics in multiple languages, filtering of false information, and ensuring secure content output.

[0120] Policy compliance: Use the reasoning-based multilingual security protection framework ConsistentGuard to achieve multilingual compliance verification of legal documents and medical reports.

[0121] Parts of the present invention that are not described in detail may refer to the prior art or are well-known technologies to those skilled in the art, and this embodiment does not limit this and will not be described in detail here.

[0122] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.< / think>

Claims

1. A multi-language security protection framework based on reasoning, characterized by: The framework achieves cross-language knowledge transfer and enhanced interpretability by combining thought chain reasoning with constraint alignment optimization.

2. The inference-based multi-language security protection framework according to claim 1 is characterized in that: The framework includes: The SFT-based cold start module is configured to perform knowledge distillation on the basic large language model through supervised fine-tuning, giving the model preliminary reasoning capabilities for security protection tasks; The GRPO-based reasoning training module is configured to improve the standardization, accuracy, and diversity of the model reasoning chain and enhance interpretability; The CAO-based cross-lingual alignment module is configured to enable knowledge transfer from high-resource languages ​​to low-resource languages ​​while avoiding performance degradation in high-resource languages.

3. The inference-based multi-language security protection framework according to claim 2, characterized in that: The cold start module uses several English seed samples, combined with preset task decomposition and step-by-step reasoning templates to guide the teacher model to generate a structured dataset containing original input, labels and complete reasoning chains, and then allows the student model to learn the reasoning paradigm through end-to-end training.

4. The inference-based multi-language security protection framework according to claim 2, characterized in that: The inference training module is based on the model after cold start and the corresponding inference data, and adopts the group relative strategy optimization algorithm to perform reinforcement learning training on the model. During the training process, the inference process is constrained by the reward function.

5. The inference-based multi-language security protection framework according to claim 2, characterized in that: The cross-language alignment module introduces a constrained alignment optimization strategy based on inference training, and migrates the high-quality inference paths obtained in high-resource languages ​​to low-resource languages ​​through data alignment and regularization methods, thereby improving the protection capabilities and inference consistency in low-resource languages.

6. The inference-based multi-language security protection framework according to claim 3, characterized in that: The cold start module uses the cross entropy loss function to optimize the model output. The loss function formula is: Among them, N is the total number of sample tokens, t i is the true label of the i-th position, The predicted probabilities generated for the model.

7. The inference-based multi-language security protection framework according to claim 4, characterized in that: The reward function of the inference training module includes: Format bonuses, which place specific requirements on where the reasoning chain is encapsulated and where the results are placed to ensure output consistency; Accuracy reward: a positive reward is given based on the degree of match between the inference result and the true label; Length reward, which constrains the length of the inference chain through trigonometric functions to avoid being too short or too long; Diversified rewards,reduce low-quality duplicate content by penalizing triplet duplication rates.

8. The inference-based multi-language security protection framework according to claim 5, characterized in that: The cross-language alignment module works as follows: Using mainstream languages ​​as seed data, we automatically translate and generate multilingual versions, which are then fed into the trained model for inference sampling. The output of each language model obtained by sampling is divided into a success set and a failure set based on the correctness; For each set of failure samples, we retrieve successful samples with the same semantics and correct reasoning in mainstream languages, and automatically synthesize the quadruple data required for cross-language alignment training; the quadruple data includes anchor samples, selected samples, input samples, and rejected samples; The model is optimized with CAO alignment loss as the main goal. During optimization, the loss encourages the priority adoption of correct outputs in the cross-lingual reasoning chain. At the same time, to ensure that the spatial distribution of anchor samples in mainstream languages ​​is not excessively perturbed, a global regularization term is introduced to constrain the representation of anchor samples before and after alignment.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the content of the inference-based multi-language security protection framework described in any one of claims 1 to 8 is implemented.

10. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the content of the inference-based multi-language security protection framework according to any one of claims 1 to 8.

Citation Information

Cited By

  • Optimization strategy model training method and device, storage medium and equipment

    CN121579019A

  • A large language model security alignment method and system based on multi-language self-distillation

    CN122549607A