Large model content security multi-level defense method
By constructing a multi-level collaborative defense framework across the entire chain, the problems of dynamic attacks and missed detections in AI-generated content have been solved, achieving efficient security interception and compliance, and meeting the standard of "Basic Requirements for Security of Generative Artificial Intelligence Services".
Patent Information
- Application Number
- CN202510732752.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies struggle to effectively handle contextual manipulation and dynamic attacks in multi-turn dialogues, as well as output-end detection issues, in the security protection of AI-generated content. Furthermore, traditional solutions cannot meet the mandatory requirement of dual input and output verification in the "Basic Requirements for Security of Generative Artificial Intelligence Services".
A multi-level collaborative defense framework is constructed across the entire input-processing-output chain, including a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. Through mixed dataset training and supervised fine-tuning plus human feedback reinforcement learning, the model's ability to identify illegal content is optimized, enabling dynamic routing and security interception.
It achieves full-process security control over the content of large models, reduces the rate of missed detections and false interceptions, improves the ability to identify adversarial attacks, and meets the compliance and accuracy requirements of the "Basic Requirements for Security of Generative Artificial Intelligence Services".
Smart Images

Figure CN120880683A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence security technology, and in particular to a multi-level defense method for large model content security. Background Technology
[0002] In real-time risk interception and security enhancement scenarios for generative AI services such as AI (Artificial Intelligence) assistants, education, and social media, user input may contain explicit violations (such as political or pornographic keywords) and implicit inducements (such as jailbreak attacks asking "how to bypass the review"). Traditional rule engines and single classifiers are difficult to cover dynamically changing attack methods.
[0003] The generated content of large models has high uncertainty and contains logical traps (such as discriminatory reasoning) and hidden value biases (such as biased descriptions of historical events), so secondary interception at the output end is required.
[0004] China's "Basic Requirements for Security of Generative Artificial Intelligence Services" (TC260-003) explicitly requires the establishment of a dual-review mechanism for inputs and outputs, and full-scale testing of 31 types of risks, including illegality, ethics, and bias. It must meet stringent requirements such as "comprehensive coverage of the rejection test question bank" (e.g., 5000+ rejection samples) and "dynamic update frequency" (at least once a month).
[0005] Currently, the shortcomings of existing multi-level protection methods for large model content security include:
[0006] Single-point defense failure: Existing protection solutions rely on front-end classifiers to intercept high-risk inputs, but cannot handle contextual guidance in multi-turn dialogues (such as users guiding the model to generate illegal content step by step).
[0007] Output-end omissions: Existing protection schemes rely solely on the model's own secure alignment (such as RLHF), which still results in "illusory outputs" (such as fictitious legal clauses), requiring a separate post-filtering module to fill the gap.
[0008] Dynamic countermeasure requirements: Malicious users constantly iterate their attack methods (such as jailbreak attacks and pre-filled answers to induce attacks), making a single static defense model prone to failure. Summary of the Invention
[0009] Embodiments of the present invention provide a multi-level security defense method for large model content, so as to effectively perform multi-level security defense for large model content.
[0010] To achieve the above objectives, the present invention adopts the following technical solution.
[0011] A multi-level defense method for content security of large models includes:
[0012] A multi-level collaborative defense framework is constructed across the entire input-processing-output chain. The multi-level collaborative defense framework includes a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. The multi-level collaborative defense framework is then deployed in a large model.
[0013] The full-link multi-level collaborative defense framework receives user input data to be detected. The user input data is first input to the sensitive word detection module, and the content that passes the sensitive word detection is transmitted to the intent recognition module.
[0014] The intent recognition module determines the risk level of the user input data transmitted by the sensitive word detection module. The risk level includes high risk, medium risk and low risk. The user input data identified as medium or high risk is routed to the risk label classification module.
[0015] The risk label classification module classifies the user input data transmitted from the intent recognition module into risk categories, and then transmits the risk-classified user input data to the security enhancement module.
[0016] The security enhancement module generates compliant system response content through domain fine-tuning and reinforcement learning strategies. After performing security interception processing on the system response content through interception strategies, it returns it to the user.
[0017] Preferably, the multi-level collaborative defense framework further includes a security filtering module deployed at the output port of the multi-level collaborative defense framework. The security filtering module filters the system response content generated by the security enhancement module, intercepts any missed violations, and returns the filtered system response content to the user.
[0018] Preferably, the multi-level collaborative defense framework for constructing the entire input-processing-output chain includes a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. Deploying this multi-level collaborative defense framework within a large model includes:
[0019] A multi-level collaborative defense framework with a complete input-processing-output chain is constructed. This framework includes a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. A hybrid dataset is built, containing an adversarial sample library, red team test data, and regulatory compliance question-and-answer pairs. The adversarial sample library collects jailbreak attack cases and phishing question variations. The regulatory compliance question-and-answer pairs include a fusion of regulatory compliance question-and-answer pairs, security adversarial samples, and value alignment datasets. The multi-level collaborative defense framework is trained using this hybrid dataset, employing a two-stage training process of supervised fine-tuning SFT + human feedback reinforcement learning RLHF. SFT optimizes the model's ability to identify illegal content based on compliance question-and-answer pairs and adversarial samples. RLHF optimizes the reward function using the PPO algorithm. An adversarial prompt is generated by deploying the attack model through red team adversarial training. The reward function of the training process of the multi-level collaborative defense framework includes comprehensive security, value consistency, and logical coherence. The trained multi-level collaborative defense framework is obtained.
[0020] The trained multi-level collaborative defense framework is deployed in the large model.
[0021] Preferably, the end-to-end multi-level collaborative defense framework receives user input data to be detected. The user input data is first input to the sensitive word detection module, and the content that passes the sensitive word detection is transmitted to the intent recognition module, including:
[0022] The end-to-end multi-level collaborative defense framework receives user input data to be detected. The user input data is first fed into the sensitive word detection module, which pre-defines various sensitive words. For text-image multimodal data, the sensitive word detection module needs to extract joint feature vectors through CLIP-style cross-modal alignment.
[0023] Embed(I) is the embedding vector of the image, and Embed(T) is the embedding vector of the text;
[0024] The sensitive word detection module uses a hybrid matching method based on keywords and semantic similarity to match the joint feature vector with pre-set keywords, identify security risks in multimodal data, directly block high-risk content detected by the sensitive word detection module, and transmit the content that passes the sensitive word detection to the intent recognition module.
[0025] Preferably, the intent recognition module determines the risk level of the user input data transmitted by the sensitive word detection module. This risk level includes high risk, medium risk, and low risk. User input data identified as medium or high risk is then routed to the risk label classification module, including:
[0026] The intent recognition module classifies user input data transmitted from the sensitive word detection module into risk levels: high, medium, and low. For medium and low-risk user input data, the intent recognition module uses a high-risk corpus and a semantic vector matching algorithm to identify the user's intent and dynamically routes the user input data to a general model or security enhancement module based on the user's intent. User input data identified as medium or high-risk is routed to the risk label classification module.
[0027] Preferably, the risk label classification module classifies the user input data transmitted from the intent recognition module into risk categories, and transmits the risk-classified user input data to the security enhancement module, including:
[0028] The risk label classification module categorizes user input data transmitted from the intent recognition module into risk categories and provides corresponding risk levels. For text-image multimodal content, the risk label classification module performs multimodal risk aggregation and conducts a joint risk assessment of text and image.
[0029] RiskScore = β × Risk text +(1-β)×Risk image
[0030] RiskScore represents the risk level. text Risk is the probability of text risk. image β represents the probability of risk for the image, and β represents the confidence threshold for both text and image.
[0031] The risk label classification module takes different actions based on the policy configured by the decision engine for different risk categories. These actions include blocking, answering on behalf of others, or allowing. The risk label classification module sets different rejection thresholds according to the sensitivity of risk categories, and directly rejects answers to high-risk data. The risk label classification module transmits the user input data that has been classified by risk category to the security enhancement module.
[0032] Preferably, the security enhancement module generates compliant system response content through domain fine-tuning and reinforcement learning strategies, performs security interception processing on the system response content through an interception strategy, and then returns it to the user, including:
[0033] The risk label classification module transmits the user input data, which has been classified into risk categories, to the security enhancement module. The security enhancement module generates compliant system response content through domain fine-tuning and reinforcement learning strategies. Based on the classification model and intent recognition model, the security enhancement module obtains the optimal threshold and strategy combination through benchmark evaluation and formulates content blocking strategies.
[0034] The security enhancement module's output content interception decision is as follows:
[0035]
[0036] s i For the score of the i-th filtering rule, w i Here, τ represents the rule weight, and τ represents the threshold.
[0037] The security enhancement module will return the processed system response content to the user through content interception strategies. As can be seen from the technical solution provided by this invention, this invention constructs a large-scale model security enhancement solution covering the entire process of input, inference, and output through a combination of hierarchical defense, dynamic routing, and security enhancement module training, possessing both compliance, efficiency, and scalability.
[0038] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of the invention. Attached Figure Description
[0039] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 A schematic diagram illustrating the implementation principle of a multi-level defense method for large-scale content security based on dynamic routing and security enhancement modules, provided in an embodiment of the present invention.
[0041] Figure 2 This is a flowchart illustrating a multi-level defense method for large-scale content security based on dynamic routing and security enhancement modules, provided as an embodiment of the present invention. Detailed Implementation
[0042] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0043] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or couplings. The term “and / or” as used herein includes any and all combinations of one or more of the associated listed items.
[0044] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as herein.
[0045] To facilitate understanding of the embodiments of the present invention, the following will provide further explanation and description with reference to the accompanying drawings and several specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.
[0046] The implementation principle of a multi-level defense method for large-scale content security based on dynamic routing and security enhancement modules provided in this invention is as follows: Figure 1 As shown, the specific processing flowchart is as follows: Figure 2 As shown, the processing steps include the following:
[0047] Step S10: Construct a multi-level collaborative defense framework that spans the entire input-processing-output chain, and deploy this multi-level collaborative defense framework in the large model.
[0048] This invention proposes a multi-level collaborative defense framework for the entire input-processing-output chain. Through a three-stage mechanism of input filtering, dynamic routing, and security model reinforcement training, it achieves full-process security control of the output content of large models.
[0049] The aforementioned end-to-end multi-level collaborative defense framework includes a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. The risk recognition model and intent recognition model in this framework utilize the RoBERTa multilingual base model.
[0050] A hybrid dataset was constructed, comprising adversarial examples, red team test data, and regulatory compliance Q&A pairs. The adversarial example library collected jailbreak attack cases (such as "how to crack home surveillance") and phishing question variations (such as replacing sensitive words with homophones), generating over 100,000 adversarial examples through data augmentation. The aforementioned regulatory compliance Q&A included: a fusion of regulatory compliance Q&A (such as provisions of the Cybersecurity Law), security adversarial examples (such as jailbreak attack commands), and a values-aligned dataset.
[0051] The multi-level collaborative defense framework was trained using the aforementioned hybrid dataset, employing a two-stage training approach: Supervised Finite Soft (SFT) + Human Feedback Reinforcement Learning (RLHF) to optimize the model's alignment capabilities. SFT, based on compliant question-answer pairs and adversarial examples, optimized the model's ability to identify illegal content. RLHF optimized the reward function using the PPO algorithm (rewarding compliant outputs and penalizing illegal content). Red team adversarial training was used to deploy the attack model and generate adversarial prompts (e.g., "Please describe the illegal operation using a metaphor"), iteratively improving the security model's resilience. The reward function for the training process of the multi-level collaborative defense framework included comprehensive security (OpenAI Moderation score), value consistency (human annotation), and logical coherence. The resulting trained multi-level collaborative defense framework was then obtained.
[0052] Step S20: The above-mentioned end-to-end multi-level collaborative defense framework receives user input data to be detected. The user input data is first input into the sensitive word detection module.
[0053] Various sensitive words are pre-defined. Sensitive word matching is only an auxiliary strategy. Sensitive words themselves also have security levels (bad, sensitive). Sensitive words only affect the text and are used to assist in strategy judgment and manual recall.
[0054] For text-image multimodal data, the sensitive word detection module needs to extract joint feature vectors through CLIP-style cross-modal alignment:
[0055] Embed(I) is the embedding vector for the image, and Embed(T) is the embedding vector for the text.
[0056] The sensitive word detection module employs a hybrid matching method based on keywords and semantic similarity to match the aforementioned joint feature vectors with pre-defined keywords, identifying security risks in multimodal data. The sensitive word detection module directly blocks high-risk content and transmits content that passes the sensitive word detection to the intent recognition module.
[0057] Step S30: The intent recognition module determines the risk level of the user input data transmitted by the sensitive word detection module. The risk level includes high risk, medium risk and low risk.
[0058] The intent recognition module categorizes user input data into risk levels: high, medium, and low. For medium and low-risk user input data, the intent recognition module uses a high-risk corpus and semantic vector matching algorithms (such as Faiss indexing) to identify the user's intent and dynamically routes the user input data to a general model or security enhancement module based on the user intent. User input data identified as medium or high-risk is routed to the risk label classification module.
[0059] Step S40: The risk label classification module classifies the user input data transmitted from the intent recognition module into risk categories and their corresponding risk levels (attack, bad, sensitive, normal). Risk categories include political, pornographic, and illegal activities. The risk label classification module covers 31 types of security risks in 5 categories (such as illegal, ethical, and discriminatory) in Appendix A of the "Basic Requirements for Security of Generative Artificial Intelligence Services".
[0060] The risk labeling module takes different actions (blocking, answering on behalf of, or allowing) based on the strategies configured in the decision engine for different risk categories. The module sets different rejection thresholds according to the sensitivity of the risk category (e.g., political, pornographic, values-related). For example, pornographic risks are directly blocked, while ethically controversial risks are addressed through a security enhancement module that may answer on behalf of the user or require human review. High-risk data (such as illegal or politically sensitive information) is directly rejected to prevent malicious manipulation of data.
[0061] For multimodal text-image content, the risk label classification module performs multimodal risk aggregation and joint risk assessment of text and images:
[0062] RiskScore = β × Risk text +(1-β)×Risk image
[0063] RiskScore represents the risk level. text Risk is the probability of text risk. image β represents the probability of risk for the image, and β represents the confidence threshold for text and image. Generally, text has higher precision and thus higher confidence, with β greater than 0.5.
[0064] Step S50: The risk label classification module transmits the user input data, after risk category classification, to the security enhancement module. The security enhancement module generates compliant system response content through domain fine-tuning and reinforcement learning strategies, isolating the risks of the base model. Based on the classification model and intent recognition model, the security enhancement module obtains the optimal threshold and strategy combination through benchmark evaluation, and formulates a long text (model output scenarios typically involve relatively long texts) interception strategy.
[0065] The security enhancement module replaces high-risk content (such as "inciting ethnic conflict") with a security warning (such as "We regret that the relevant content you mentioned is currently unavailable. Let's continue our discussion on a topic of mutual interest.") and triggers a manual review process.
[0066] The security enhancement module's output content interception decision is as follows:
[0067]
[0068] s i For the score of the i-th filtering rule, w i τ represents the rule weight, and τ represents the threshold.
[0069] The security enhancement module automatically incorporates user reports and interception logs into a labeled data pool. After manual review, samples are added to the training data pool, and the model is incrementally fine-tuned weekly. Based on the interception rate, the module traces and identifies high-risk users and optimizes classifier thresholds (e.g., adjusting the sensitivity of matching politically sensitive terms).
[0070] Step S60: At the output port of the multi-level collaborative defense framework, the system response content is filtered by the security filtering module to intercept any missed violations. The filtered system response content is then returned to the user.
[0071] The compliance evaluation of the aforementioned multi-level collaborative defense framework was conducted using the TC260-003 evaluation set and benchmark datasets (such as Chinese SafetyQA). The framework's interception metrics across dimensions such as "political," "pornographic," "illegal," "ethical," and "biased" were periodically tested (target miss rate <1%, accuracy >98%). Simulated red team attacks (such as "pre-filled answer inducement") were performed to calculate the attack resistance success rate of the security enhancement module (target >95%).
[0072] The application scenarios of the method in the embodiments of the present invention include:
[0073] AI-generated content scenarios enhance content security in fields such as finance (anti-fraud consulting), education (distorting historical facts), and government affairs (policy interpretation).
[0074] The educational Q&A section will provide answers to questions involving sensitive historical events, ensuring that the responses conform to the "Implementation Plan for the Reform and Innovation of Ideological and Political Theory Courses in Schools in the New Era," and will recall any content that distorts historical facts.
[0075] Commercial potential and standardized output: It can be packaged into API services or SaaS platforms to provide functions such as content security identification, tag classification, security scoring, and risk tracing reports.
[0076] Compliance toolchain: Includes red team testing tools (such as automated jailbreak attack generators) and security datasets to form a complete solution.
[0077] In summary, the embodiment of the present invention, which combines input hierarchical interception, security model-assisted response, and independent output filtering, reduces the false negative rate by 70% compared to traditional solutions.
[0078] Compared to traditional keyword filtering, the introduction of multi-level judgment and routing mechanisms can identify variant attacks (such as homophones and metaphors), greatly reducing the false blocking rate. Coupled with a security enhancement module to answer on behalf of the user, the missed blocking rate is further reduced, while the rejection rate is also significantly lowered, improving the user experience.
[0079] Anti-attack generalization
[0080] Through adversarial training and policy injection, the model's success rate in intercepting attacks that "have no illegal words in the question stem but contain implicit bias" has increased to 92% (compared to 65% for the baseline model).
[0081] Compliance assurance
[0082] Strictly aligned with the "Basic Requirements for Security of Generative Artificial Intelligence Services", it supports automated evaluation of security test question banks (such as 5000+ rejected answer samples) to meet the requirements of filing and review.
[0083] Model performance balance
[0084] By using a layered approach (the base model handles general problems, while the security enhancement module handles sensitive tasks), performance degradation caused by fine-tuning the entire model is avoided.
[0085] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.
[0086] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0087] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0088] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A multi-level defense method for content security of large models, characterized in that, include: A multi-level collaborative defense framework is constructed across the entire input-processing-output chain. The multi-level collaborative defense framework includes a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. The multi-level collaborative defense framework is then deployed in a large model. The full-link multi-level collaborative defense framework receives user input data to be detected. The user input data is first input to the sensitive word detection module, and the content that passes the sensitive word detection is transmitted to the intent recognition module. The intent recognition module determines the risk level of the user input data transmitted by the sensitive word detection module. The risk level includes high risk, medium risk and low risk. The user input data identified as medium or high risk is routed to the risk label classification module. The risk label classification module classifies the user input data transmitted from the intent recognition module into risk categories, and then transmits the risk-classified user input data to the security enhancement module. The security enhancement module generates compliant system response content through domain fine-tuning and reinforcement learning strategies. After performing security interception processing on the system response content through interception strategies, it returns it to the user.
2. The method according to claim 1, characterized in that, The multi-level collaborative defense framework also includes a security filtering module deployed at the output port of the multi-level collaborative defense framework. This security filtering module filters the system response content generated by the security enhancement module, intercepts any missed violations, and returns the filtered system response content to the user.
3. The method according to claim 1 or 2, characterized in that, The aforementioned multi-level collaborative defense framework, encompassing the entire input-processing-output chain, includes a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. Deploying this multi-level collaborative defense framework within a large model includes: A multi-level collaborative defense framework with a complete input-processing-output chain is constructed. This framework includes a sensitive word detection module, an intent recognition module, a risk label classification module, and a security filtering module. A hybrid dataset is built, containing an adversarial sample library, red team test data, and regulatory compliance question-and-answer pairs. The adversarial sample library collects jailbreak attack cases and phishing question variations. The regulatory compliance question-and-answer pairs include a fusion of regulatory compliance question-and-answer pairs, security adversarial samples, and value alignment datasets. The multi-level collaborative defense framework is trained using this hybrid dataset, employing a two-stage training process of supervised fine-tuning SFT + human feedback reinforcement learning RLHF. SFT optimizes the model's ability to identify illegal content based on compliance question-and-answer pairs and adversarial samples. RLHF optimizes the reward function using the PPO algorithm. An adversarial prompt is generated by deploying the attack model through red team adversarial training. The reward function of the training process of the multi-level collaborative defense framework includes comprehensive security, value consistency, and logical coherence. The trained multi-level collaborative defense framework is obtained. The trained multi-level collaborative defense framework is deployed in the large model.
4. The method according to claim 3, characterized in that, The described end-to-end multi-level collaborative defense framework receives user input data to be detected. The user input data is first fed into the sensitive word detection module, and the content that passes the sensitive word detection is transmitted to the intent recognition module, including: The end-to-end multi-level collaborative defense framework receives user input data to be detected. The user input data is first fed into the sensitive word detection module, which pre-defines various sensitive words. For text-image multimodal data, the sensitive word detection module needs to extract joint feature vectors through CLIP-style cross-modal alignment. Embed(I) is the embedding vector of the image, and Embed(T) is the embedding vector of the text; The sensitive word detection module uses a hybrid matching method based on keywords and semantic similarity to match the joint feature vector with pre-set keywords, identify security risks in multimodal data, directly block high-risk content detected by the sensitive word detection module, and transmit the content that passes the sensitive word detection to the intent recognition module.
5. The method according to claim 4, characterized in that, The intent recognition module determines the risk level of the user input data transmitted from the sensitive word detection module. This risk level includes high risk, medium risk, and low risk. User input data identified as medium or high risk is routed to the risk label classification module, which includes: The intent recognition module classifies user input data transmitted from the sensitive word detection module into risk levels: high, medium, and low. For medium and low-risk user input data, the intent recognition module uses a high-risk corpus and a semantic vector matching algorithm to identify the user's intent and dynamically routes the user input data to a general model or security enhancement module based on the user's intent. User input data identified as medium or high-risk is routed to the risk label classification module.
6. The method according to claim 5, characterized in that, The risk label classification module classifies the user input data transmitted from the intent recognition module into risk categories, and transmits the risk-classified user input data to the security enhancement module, including: The risk label classification module categorizes user input data transmitted from the intent recognition module into risk categories and provides corresponding risk levels. For text-image multimodal content, the risk label classification module performs multimodal risk aggregation and conducts a joint risk assessment of text and image. RiskScore=β×Risk text +(1-β)×Risk image RiskScore represents the risk level. text Risk is the probability of text risk. image β represents the probability of risk for the image, and β represents the confidence threshold for both text and image. The risk label classification module takes different actions based on the policy configured by the decision engine for different risk categories. These actions include blocking, answering on behalf of others, or allowing. The risk label classification module sets different rejection thresholds according to the sensitivity of risk categories, and directly rejects answers to high-risk data. The risk label classification module transmits the user input data that has been classified by risk category to the security enhancement module.
7. The method according to claim 6, characterized in that, The security enhancement module generates compliant system response content through domain fine-tuning and reinforcement learning strategies. After performing security interception processing on the system response content using an interception strategy, it returns it to the user, including: The risk label classification module transmits the user input data, which has been classified into risk categories, to the security enhancement module. The security enhancement module generates compliant system response content through domain fine-tuning and reinforcement learning strategies. Based on the classification model and intent recognition model, the security enhancement module obtains the optimal threshold and strategy combination through benchmark evaluation and formulates content blocking strategies. The security enhancement module's output content interception decision is as follows: s i For the score of the i-th filtering rule, w i Here, τ represents the rule weight, and τ represents the threshold. The security enhancement module will intercept and process the system response content and return it to the user using content interception strategies.
Citation Information
Cited By
Safety control method, system and equipment for interaction data of large language model and medium
CN121051738A
Content risk detection filtering method and device, electronic equipment and storage medium
CN121808765A