Large model content security two-link triple dynamic defense method and device, equipment, storage medium
Patent Information
- Application Number
- CN202610891738.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-08-18
AI Technical Summary
然而,现有安全技术在实践中存在不足:一方面,大模型因其生成能力,在处理复杂、模糊或隐含恶意意图的用户输入时,可能出现对潜在安全风险的审核拦截遗漏;另一方面,模型生成的内容也可能在无意中输出包含偏见、虚假信息、不符合伦理或特定地区监管政策的内容
[0020] The two-stage, three-tiered dynamic defense method, apparatus, device, and storage medium for large model content security provided in this application, in response to receiving model input content, performs a security audit on the model input content to obtain an input audit result. If the input audit result is risky, the model input content is intercepted and risk response information is returned. If the input audit result is safe, the model output content corresponding to the model input content is determined through the large model. The model output content is then subjected to a security audit to obtain an output audit result. If the output audit result is risky, the model output content is intercepted and risk response information is returned. This application comprehensively achieves model content security defense by performing security audits at both the model input and output stages, ensuring that the final output content is safe.
Smart Images

Figure CN122601334A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of large model technology, and to, but is not limited to, a two-stage, three-fold dynamic defense method, apparatus, device, and storage medium for the security of large model content. Background Technology
[0002] With the development of large-scale model capabilities and their widespread application in key areas such as search, dialogue, content generation, and decision support, their security has become a core issue concerning social stability, ethics, and legal compliance. However, existing security technologies have shortcomings in practice: on the one hand, due to their generative capabilities, large-scale models may overlook potential security risks when processing complex, ambiguous, or malicious user input; on the other hand, the content generated by the models may unintentionally output biased, false, unethical, or regionally specific regulatory policies. This lack of security assurance may not only lead to serious real-world harm but also makes it difficult for the application of large-scale models to meet increasingly stringent global regulatory frameworks and the fundamental societal requirement that technological development align with shared human values. Summary of the Invention
[0003] In view of this, the two-stage triple dynamic defense method, apparatus, device, and storage medium for large model content security provided in the embodiments of this application can achieve comprehensive and accurate two-stage triple dynamic defense for large model content security.
[0004] The large model content security two-stage triple dynamic defense method, device, equipment, and storage medium provided in this application embodiment are implemented as follows: One aspect of this application provides a two-stage, three-layer dynamic defense method for large model content security, the method comprising: In response to receiving model input content, the model input content is subjected to security audit, and the input audit result is obtained; If the input review result is a risk result, the model input content is intercepted and risk response information is returned; If the input review result is a safe result, the model output content corresponding to the model input content is determined through the large model; The output content of the model is subjected to security audit to obtain the output audit result; If the output review result is a risk result, the model output content is intercepted and risk response information is returned.
[0005] In one possible implementation, the step of performing security audits on the model input content to obtain the input audit result includes: The model input content is input into the three-layer risk review module for risk review, and the input review result is obtained; The three-tiered risk review module includes a large model review layer, a small model review layer, and a rule review layer.
[0006] In one possible implementation, the step of inputting the model input content into a three-layer risk review module for risk review and obtaining the input review result includes: The input content of the model is reviewed by the large model review layer, the small model review layer, and the rule review layer, respectively. If at least one of the audit results in the large model audit layer, small model audit layer, and rule audit layer is considered risky, the input audit result is determined to be a risky result. If all the review results of the large model review layer, the small model review layer, and the rule review layer are deemed safe, then the input review result is determined to be a safe result.
[0007] In one possible implementation, the step of reviewing the model input content through the large model review layer, the small model review layer, and the rule review layer respectively includes: The input content of the model is reviewed sequentially through the large model review layer, the small model review layer, and the rule review layer. If the review result of the large model review layer or the small model review layer is deemed risky, the review process shall be stopped.
[0008] In one possible implementation, the large model review layer is used to perform risk review on the model input content through a large model, the small model review layer is used to perform risk review on the model input content based on rules through a lightweight neural network model, and the rule review layer is used to perform risk review on the model input content through a real-time rule base.
[0009] In one possible implementation, the method further includes: The training set is periodically determined based on historical input review results and / or output review results; The three-layer risk review module is adjusted based on the training set.
[0010] In one possible implementation, adjusting the three-layer risk review module based on the training set includes: Positive and negative samples are obtained from the training set; The large model audit layer is supervised and fine-tuned based on the positive and negative samples. The small model review layer is updated by extracting risk features from the negative samples; Temporary or emergency risk rules are determined by extracting risk characteristics from the negative samples and then updated to the implementation rule base of the rule review layer.
[0011] Another aspect of this application embodiment provides a two-stage, three-layer dynamic defense device for large model content security, the device comprising: The first review module is used to perform a security review on the received model input content in response to receiving the model input content, and obtain the input review result; The first response module is used to intercept the model input content and return risk response information when the input review result is a risk result; The output content matching module is used to determine the model output content corresponding to the model input content by using a large model, provided that the input review result is a safe result. The second review module is used to perform security review on the output content of the model and obtain the output review result. The second response module is used to intercept the model's output content and return risk response information when the output review result is a risk result.
[0012] In one possible implementation, the first audit module is further configured to: The model input content is input into the three-layer risk review module for risk review, and the input review result is obtained; The three-tiered risk review module includes a large model review layer, a small model review layer, and a rule review layer.
[0013] In one possible implementation, the first audit module is further configured to: The input content of the model is reviewed by the large model review layer, the small model review layer, and the rule review layer, respectively. If at least one of the audit results in the large model audit layer, small model audit layer, and rule audit layer is considered risky, the input audit result is determined to be a risky result. If all the review results of the large model review layer, the small model review layer, and the rule review layer are deemed safe, then the input review result is determined to be a safe result.
[0014] In one possible implementation, the first audit module is further configured to: The input content of the model is reviewed sequentially through the large model review layer, the small model review layer, and the rule review layer. If the review result of the large model review layer or the small model review layer is deemed risky, the review process shall be stopped.
[0015] In one possible implementation, the large model review layer is used to perform risk review on the model input content through a large model, the small model review layer is used to perform risk review on the model input content based on rules through a lightweight neural network model, and the rule review layer is used to perform risk review on the model input content through a real-time rule base.
[0016] In one possible implementation, the device further includes: The training set acquisition module is used to periodically determine the training set based on historical input review results and / or output review results; The model adjustment module is used to adjust the three-layer risk review module based on the training set.
[0017] In one possible implementation, the model adjustment module is further configured to: Positive and negative samples are obtained from the training set; The large model audit layer is supervised and fine-tuned based on the positive and negative samples. The small model review layer is updated by extracting risk features from the negative samples; Temporary or emergency risk rules are determined by extracting risk characteristics from the negative samples and then updated to the implementation rule base of the rule review layer.
[0018] The electronic device provided in this application includes a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements the method described in this application.
[0019] The computer-readable storage medium provided in this application embodiment stores a computer program thereon, which, when executed by a processor, implements the method provided in this application embodiment.
[0020] The two-stage, three-tiered dynamic defense method, apparatus, device, and storage medium for large model content security provided in this application, in response to receiving model input content, performs a security audit on the model input content to obtain an input audit result. If the input audit result is risky, the model input content is intercepted and risk response information is returned. If the input audit result is safe, the model output content corresponding to the model input content is determined through the large model. The model output content is then subjected to a security audit to obtain an output audit result. If the output audit result is risky, the model output content is intercepted and risk response information is returned. This application comprehensively achieves model content security defense by performing security audits at both the model input and output stages, ensuring that the final output content is safe. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a two-stage, three-fold dynamic defense method for large model content security according to an embodiment of this application is shown. Figure 2 This diagram illustrates the structure of a three-tier risk assessment module according to an embodiment of this application. Figure 3 This diagram illustrates a two-stage, three-layer dynamic defense process for large-scale content security according to an embodiment of this application. Figure 4 A schematic diagram of a two-stage, triple-dynamic defense device for large model content security according to an embodiment of this application is shown. Figure 5 A schematic diagram of an electronic device according to an embodiment of this application is shown. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the specific technical solutions of this application will be further described in detail below with reference to the accompanying drawings of the embodiments of this application. The following embodiments are used to illustrate this application, but are not intended to limit the scope of this application.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0025] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0026] It should be noted that the terms "first, second, third" used in the embodiments of this application are used to distinguish similar or different objects and do not represent a specific order of objects. It can be understood that "first, second, third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0027] The two-stage, three-layer dynamic defense method for large model content security in this application embodiment can be executed by any electronic device, including but not limited to mobile phones, wearable devices (such as smartwatches, smart bracelets, smart glasses, etc.), tablets, laptops, in-vehicle terminals, PCs (Personal Computers), etc. The functionality implemented by this method can be achieved by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. Therefore, the electronic device includes at least a processor and a storage medium.
[0028] The two-stage, three-fold dynamic defense method for large model content security in this application can be used in any large model content security defense scenario, such as application scenarios for security defense of policy-sensitive information input or output of large models, or application scenarios for security defense of confidentiality-sensitive data input or output of large models.
[0029] In existing two-stage, three-tiered dynamic defense methods for large-scale model content security, current mainstream security protection measures typically suffer from the following limitations: First, relying on a single static rule base or keyword filtering makes it difficult to cope with the complexity and variability of semantics within the context of large models. They are also weak in identifying risky content that is disguised, metaphorical, or relies on contextual understanding, making them easily bypassed. Second, deployment at only a single input or output stage lacks dynamic monitoring and defense throughout the entire process. If any stage fails, risky content can easily penetrate the system. Third, the model's own security policies are updated lagging behind, making it difficult to respond quickly to emerging attack patterns and sudden security vulnerabilities, and failing to meet emergency response needs requiring minute-level or even real-time updates. Fourth, the lack of effective closed-loop feedback and continuous optimization mechanisms makes it impossible to systematically evaluate the effectiveness of security policies (whether there are false positives or false negatives), and it is difficult to use actual business data to drive iterative upgrades of models and rules, leading to a gradual rigidity and ineffectiveness of the security protection system. Therefore, there is an urgent need for a multi-layered, intelligent security protection system that can deeply understand semantics, dynamically adapt to risk changes, cover the entire input-output chain, continuously evolve, and ensure high consistency with human values and regulatory requirements.
[0030] To address the shortcomings of the existing technologies, this application's embodiments construct a multi-layered, in-depth defense and closed-loop feedback mechanism, significantly reducing false positives and false negatives, and ensuring that the output content of the large model meets the requirements of security supervision and alignment with human values.
[0031] The following describes in detail the two-stage, three-fold dynamic defense scheme for large model content security in this application embodiment, with reference to the accompanying drawings.
[0032] Figure 1 A flowchart illustrating a two-stage, three-tiered dynamic defense method for large model content security according to an embodiment of this application is shown. Figure 1As shown, the large model content security two-stage triple dynamic defense method of this application embodiment may include the following steps S10-S50.
[0033] For ease of description, the two-stage, three-layer dynamic defense method for large model content security in this application embodiment is described using an electronic device as the execution subject. It should be understood that the execution subject in this application embodiment can also be a processor or chip in an electronic device, and this application embodiment does not impose any limitations.
[0034] Step S10: In response to receiving the model input content, perform a security audit on the model input content to obtain the input audit result.
[0035] In one possible implementation, the electronic device can perform a security audit on the model input content received from the user, thus obtaining an input audit result. The model input content serves as instructions or contextual information to guide the large model in generating corresponding answers. Examples include questions like "What is quantum mechanics?", "Summarize this passage into three key points", and "You are now a patient psychologist; please listen to my troubles and give advice." Any type of model input content may contain illegal information, discriminatory remarks, privacy breaches, extremist content, or other content that does not conform to human societal values, or violate data privacy protection regulations and content security laws. Therefore, upon receiving the model input content, the electronic device can determine whether the model input content complies with preset normative documents or public order and good morals of human society by performing a security audit, thereby obtaining the corresponding input audit result.
[0036] Optionally, in this embodiment, the electronic device is pre-deployed with a three-layer risk review module for content risk review. This module includes three layers: a large model review layer, a small model review layer, and a rule review layer. The large model review layer performs risk review on the model input content using a large model; the small model review layer performs risk review on the model input content using a lightweight neural network model based on rules; and the rule review layer performs risk review on the model input content using a real-time rule base. In some embodiments, the large model review layer is configured to leverage the deep semantic understanding capabilities of the large model to perform value-alignment review of content, identifying potential security risks in complex contexts, such as illegal information, discriminatory speech, privacy breach inducements, and extremist content. This layer serves as the core review layer, undertaking the primary deep review tasks. The small model review layer is configured to utilize lightweight small models based on real-time updated rules to agilely intercept risky content that the large model's precise review layer may have missed. This layer, based on a streamlined and efficient small model, supports minute-level policy updates, enabling rapid response to sudden security vulnerabilities and new attack patterns. The rule review layer is configured to rapidly respond to and intercept sudden security vulnerabilities and new attack patterns based on a real-time rule base that can be updated minute-level. This layer serves as the outermost rapid defense line, providing immediate interception of urgent security events.
[0037] Based on the pre-deployed three-layer risk review module, electronic devices can input model input content into the three-layer risk review module for risk review when they receive model input content from users through human-computer interaction or through other devices, and obtain the input review result.
[0038] In some embodiments, after the electronic device inputs the model input content into the three-layer risk review module, the input content can be reviewed by the large model review layer, the small model review layer, and the rule review layer respectively. If at least one of the review results from the large model review layer, the small model review layer, and the rule review layer is considered risky, the input review result is determined to be a risky result. If all the review results from the large model review layer, the small model review layer, and the rule review layer are considered safe, the input review result is determined to be a safe result.
[0039] In other embodiments, the three different review layers in the three-layer risk review module also have corresponding review sequences, where the review sequence can be to review the model input content sequentially through the large model review layer, the small model review layer, and the rule review layer. If the review result of the large model review layer or the small model review layer is risky, the review process stops. That is, if the large model review layer determines that the model input content is risky, the review process stops and the input review result is directly determined to be risky; if the large model review layer determines that the model input content is safe, the small model review layer performs a second review of the model input content. If the small model review layer determines that the model input content is risky, the review process stops and the input review result is directly determined to be risky; if the small model review layer determines that the model input content is safe, the rule review layer performs a third review of the model input content, and the third review result is directly determined as the input review result.
[0040] Step S20: If the input review result is a risk result, intercept the model input content and return risk response information.
[0041] In one possible implementation, after the electronic device performs risk review on the model input content through a three-layer risk review module, the input review result can include a risk result and a safety result. The safety result indicates that the model input content does not violate preset normative documents or public order and good morals of human society, while the risk result indicates that the model input content violates preset normative documents or public order and good morals of human society. If the input review result is a risk result, the electronic device can intercept the model input content and return risk response information.
[0042] Optionally, the risk response information in this application embodiment can be a preset catch-all information. The content of the preset catch-all information can be a safe, neutral, and compliant general response text, such as "I'm sorry, I cannot answer this question or obtain relevant information about this question. If you have any other questions, I'd be happy to help." This is used to replace the output when intercepting risky content, preventing the large model from generating risky and harmful information based on risky model input content.
[0043] Step S30: If the input review result is a safe result, determine the model output content corresponding to the model input content through the large model.
[0044] In one possible implementation, in this embodiment of the application, after the electronic device performs risk review on the model input content through a three-layer risk review module, and the input review result is a safe result, the large model determines the model output content corresponding to the model input content. For different model input content, the large model generates corresponding different types of model output content, which may include direct answers that answer facts or explain concepts, task-oriented answers that complete specific tasks, and analytical answers obtained through comparison, induction, and deduction.
[0045] Step S40: Perform a security audit on the output content of the model to obtain the output audit result.
[0046] In one possible implementation, after the electronic device receives the model output content, it also needs to conduct a security audit on the model output content to avoid risks such as non-compliance or violation of public order and good morals in the model output content generated by the large model, and obtain the output audit result. The risk audit of the model output content can also be performed through a three-layer risk audit module. The audit process for the model output content is the same as that for the model input content, and will not be elaborated further here.
[0047] Step S50: If the output audit result is a risk result, intercept the model output content and return risk response information.
[0048] In one possible implementation, after the electronic device performs risk review on the model output content through a three-layer risk review module, the output review result may include both a risk result and a safety result. The safety result indicates that the model output content does not violate preset normative documents or public order and good morals of human society, while the risk result indicates that the model output content violates preset normative documents or public order and good morals of human society. If the output review result is a risk result, the electronic device can intercept the model output content and return risk response information. If the output review result is a safety result, the electronic device can directly return and output the model output content.
[0049] Figure 2 This diagram illustrates the structure of a three-tier risk assessment module according to an embodiment of this application. Figure 2 As shown, this embodiment of the application establishes a three-layer risk review module to obtain a triple dynamic protection system. The large model review layer serves as the first core defense line for precise review; the small model review layer serves as the second specialized enhancement defense line for rapid risk interception; and the rule review layer serves as the third agile response defense line for emergency risk mitigation. If any layer of the three defense lines determines that the input content of the input model (user input content) or the model output content determined based on the model input content has a violation risk, risk interception is performed, and risk response information (fallback content) is output. If neither the model input content nor the model output content determined based on the model input content has a violation risk, the model output content is output.
[0050] In some embodiments, to ensure the real-time accuracy of the three-layer risk review module, the module can be periodically adjusted. This process can periodically determine a training set based on historical input and / or output review results, and adjust the three-layer risk review module based on the training set. Optionally, the electronic device acquires positive and negative samples from the training set, performs supervised fine-tuning of the large model review layer based on the positive and negative samples, updates the small model review layer by extracting risk features from the negative samples, determines temporary or emergency risk rules by extracting risk features from the negative samples, and updates the implementation rule base of the rule review layer.
[0051] Figure 3 This diagram illustrates a two-stage, three-tiered dynamic defense process for large-scale content security according to an embodiment of this application. Figure 3As shown, after each review of model input or output content, the electronic device can use a combination of automated evaluation tools and manual review to assess the correctness of the interception or release operations for the three-layer risk review model input or output content. If the evaluation indicates correct processing or no false positives or false negatives, the electronic device records it as a positive example in the security evaluation data, resulting in a positive sample. If the evaluation indicates incorrect processing or false positives or false negatives, it records it as a negative example in the security evaluation data, resulting in a negative sample. For example, when the electronic device correctly intercepts a model input containing illegal information, the processing result is marked as a positive sample; when the electronic device incorrectly intercepts a legitimate inquiry, the processing result is marked as a negative sample; when the electronic device fails to intercept a model output containing biased content, the processing result is also marked as a negative sample.
[0052] Furthermore, electronic devices can periodically acquire a training set for the current period based on determined positive and negative samples, and adjust the three-layer risk assessment model based on the training set. Specifically, the electronic devices can use positive and negative samples as training samples to perform supervised fine-tuning of the large model's assessment layer, improving its assessment accuracy and enabling it to learn from erroneous cases to avoid similar missed or false positives. New risk features are extracted from negative samples to update the rule base of the small model's assessment layer, enabling it to identify newly emerging attack patterns. Temporary or emergency risk rules are extracted from negative samples to update the real-time rule base of the rule assessment layer, achieving minute-level emergency response. Through these methods, the entire protection system forms a virtuous cycle of "error detection - strategy correction - avoiding repeated errors," continuously evolving.
[0053] Based on the aforementioned technical features, this application's embodiments construct a multi-layered, in-depth defense system comprising a large model review layer, a small model review layer, and a rule review layer. This achieves full coverage from deep semantic understanding to agile rule response, significantly improving the ability to identify complex risky content. Simultaneously, deploying this triple defense mechanism at both the input and output ends forms a dynamic monitoring system across the entire chain, effectively preventing the leakage of risky content due to the failure of a single link. Furthermore, by constructing a closed-loop security evaluation feedback system combining automatic and manual methods, positive and negative sample data are collected to drive model fine-tuning and rule updates, enabling continuous dynamic evolution of the protection system and agile response to new attack patterns and sudden security vulnerabilities. Moreover, with risk review conducted through a three-layered sequential review mechanism, any level detecting a risk terminates subsequent reviews and intercepts the vulnerability, ensuring both comprehensiveness of the review and efficiency of response. Negative sample data is used to reverse-optimize the review strategies at each layer, forming a virtuous cycle of "error detection - strategy correction - avoiding repeated errors," significantly reducing the false positive and false negative rates.
[0054] It should be understood that although the steps in the above flowcharts are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the above flowcharts may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0055] Based on the foregoing embodiments, this application provides a two-stage, three-layer dynamic defense device for large model content security. The device includes various modules and units included in each module, which can be implemented by a processor; of course, it can also be implemented by specific logic circuits. In the implementation process, the processor can be a central processing unit (CPU), microprocessor (MPU), digital signal processor (DSP), or field programmable gate array (FPGA), etc.
[0056] Figure 4 This diagram illustrates a two-stage, triple-dynamic defense device for large-scale content security according to an embodiment of this application. Figure 4 As shown, the two-stage, three-layer dynamic defense device for large model content security in this application embodiment includes: The first review module 40 is used to perform security review on the model input content in response to receiving the model input content, and obtain the input review result; The first response module 41 is used to intercept the model input content and return risk response information when the input review result is a risk result; The output content matching module 42 is used to determine the model output content corresponding to the model input content by using a large model when the input review result is a safe result. The second review module 43 is used to perform security review on the output content of the model and obtain the output review result; The second response module 44 is used to intercept the model output content and return risk response information when the output review result is a risk result.
[0057] In one possible implementation, the first audit module 40 is further configured to: The model input content is input into the three-layer risk review module for risk review, and the input review result is obtained; The three-tiered risk review module includes a large model review layer, a small model review layer, and a rule review layer.
[0058] In one possible implementation, the first audit module 40 is further configured to: The input content of the model is reviewed by the large model review layer, the small model review layer, and the rule review layer, respectively. If at least one of the audit results in the large model audit layer, small model audit layer, and rule audit layer is considered risky, the input audit result is determined to be a risky result. If all the review results of the large model review layer, the small model review layer, and the rule review layer are deemed safe, then the input review result is determined to be a safe result.
[0059] In one possible implementation, the first audit module 40 is further configured to: The input content of the model is reviewed sequentially through the large model review layer, the small model review layer, and the rule review layer. If the review result of the large model review layer or the small model review layer is deemed risky, the review process shall be stopped.
[0060] In one possible implementation, the large model review layer is used to perform risk review on the model input content through a large model, the small model review layer is used to perform risk review on the model input content based on rules through a lightweight neural network model, and the rule review layer is used to perform risk review on the model input content through a real-time rule base.
[0061] In one possible implementation, the device further includes: The training set acquisition module is used to periodically determine the training set based on historical input review results and / or output review results; The model adjustment module is used to adjust the three-layer risk review module based on the training set.
[0062] In one possible implementation, the model adjustment module is further configured to: Positive and negative samples are obtained from the training set; The large model audit layer is supervised and fine-tuned based on the positive and negative samples. The small model review layer is updated by extracting risk features from the negative samples; Temporary or emergency risk rules are determined by extracting risk characteristics from the negative samples and then updated to the implementation rule base of the rule review layer.
[0063] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0064] It should be noted that, in the embodiments of this application... Figure 4 The module division of the large-scale model's two-stage, three-layer dynamic defense device for content security is illustrative and represents only one logical functional division; in actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of this application can be integrated into a single processing unit, exist as separate physical entities, or be integrated into a single unit. The integrated units can be implemented in hardware, as software functional units, or a combination of both.
[0065] It should be noted that, in the embodiments of this application, if the above-described methods are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0066] Figure 5 A schematic diagram of an electronic device according to an embodiment of this application is shown. For example... Figure 5 As shown in the figure, this application provides an electronic device, which can be a server, and its internal structure diagram can be as follows. Figure 5 As shown, the electronic device includes a processor 520, a memory, and a transceiver 540 connected via a system bus 510. The processor 520 provides computing and control capabilities. The memory includes a non-volatile storage medium 531 and internal memory 532. The non-volatile storage medium 531 stores an operating system, computer programs, and a database. The internal memory 532 provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium 531. The database stores data. The transceiver 540 communicates with external terminals via a network connection. When the computer program is executed by the processor 520, it implements the methods described above.
[0067] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor 520, implements the steps of the method provided in the above embodiments.
[0068] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the steps in the method provided in the above-described method embodiments.
[0069] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0070] In one possible implementation, the apparatus provided in this application can be implemented as a computer program, which can be configured as follows: Figure 5 The device operates on the electronic device shown. The memory of the electronic device can store various program modules that make up the above-described apparatus. The computer program composed of the various program modules causes the processor 520 to execute the steps of the methods in the various embodiments of this application described in this specification.
[0071] It should be noted that the descriptions of the storage medium and device embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0072] It should be understood that the phrases "one embodiment," "an embodiment," or "some embodiments" mentioned throughout the specification mean that a specific feature, structure, or characteristic related to an embodiment is included in at least one embodiment of this application. Therefore, phrases such as "in one possible implementation," "in one embodiment," or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. The descriptions of the various embodiments above tend to emphasize the differences between the various embodiments; their similarities or commonalities can be referred to mutually, and for the sake of brevity, they will not be repeated here.
[0073] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three kinds of relationships. For example, object A and / or object B can represent three situations: object A exists alone, object A and object B exist simultaneously, and object B exists alone.
[0074] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, apparatus, article, or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, apparatus, article, or device. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, apparatus, article, or device that includes that element.
[0075] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple modules or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or modules can be electrical, mechanical, or other forms.
[0076] The modules described above as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules. They may be located in one place or distributed across multiple network units. Some or all of the modules may be selected to achieve the purpose of this embodiment according to actual needs.
[0077] In addition, each functional module in the various embodiments of this application can be integrated into one processing unit, or each module can be a separate unit, or two or more modules can be integrated into one unit; the integrated modules can be implemented in hardware or in the form of hardware plus software functional units.
[0078] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0079] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause an electronic device to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0080] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0081] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0082] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0083] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A two-stage, three-tiered dynamic defense method for large-scale model content security, characterized in that: The method includes: In response to receiving model input content, the model input content is subjected to security audit, and the input audit result is obtained; If the input review result is a risk result, the model input content is intercepted and risk response information is returned; If the input review result is a safe result, the model output content corresponding to the model input content is determined through the large model; The output content of the model is subjected to security audit to obtain the output audit result; If the output review result is a risk result, the model output content is intercepted and risk response information is returned.
2. The method according to claim 1, characterized in that, The process of performing security audits on the model input content to obtain the input audit results includes: The model input content is input into the three-layer risk review module for risk review, and the input review result is obtained; The three-tiered risk review module includes a large model review layer, a small model review layer, and a rule review layer.
3. The method according to claim 2, characterized in that, The step of inputting the model input content into the three-layer risk review module for risk review and obtaining the input review result includes: The input content of the model is reviewed by the large model review layer, the small model review layer, and the rule review layer, respectively. If at least one of the audit results in the large model audit layer, small model audit layer, and rule audit layer is considered risky, the input audit result is determined to be a risky result. If all the review results of the large model review layer, the small model review layer, and the rule review layer are deemed safe, then the input review result is determined to be a safe result.
4. The method according to claim 3, characterized in that, The process of reviewing the model input content through the large model review layer, the small model review layer, and the rule review layer includes: The input content of the model is reviewed sequentially through the large model review layer, the small model review layer, and the rule review layer. If the review result of the large model review layer or the small model review layer is deemed risky, the review process shall be stopped.
5. The method according to claim 2, characterized in that, The large model review layer is used to perform risk review on the model input content through a large model, the small model review layer is used to perform risk review on the model input content based on rules through a lightweight neural network model, and the rule review layer is used to perform risk review on the model input content through a real-time rule base.
6. The method according to claim 2, characterized in that, The method further includes: The training set is periodically determined based on historical input review results and / or output review results; The three-layer risk review module is adjusted based on the training set.
7. The method according to claim 6, characterized in that, The adjustment of the three-layer risk review module based on the training set includes: Positive and negative samples are obtained from the training set; The large model audit layer is supervised and fine-tuned based on the positive and negative samples. The small model review layer is updated by extracting risk features from the negative samples; Temporary or emergency risk rules are determined by extracting risk characteristics from the negative samples and then updated to the implementation rule base of the rule review layer.
8. A two-stage, three-layer dynamic defense device for the content security of large models, characterized in that, The device includes: The first review module is used to perform a security review on the received model input content in response to receiving the model input content, and obtain the input review result; The first response module is used to intercept the model input content and return risk response information when the input review result is a risk result; The output content matching module is used to determine the model output content corresponding to the model input content by using a large model, provided that the input review result is a safe result. The second review module is used to perform security review on the output content of the model and obtain the output review result. The second response module is used to intercept the model's output content and return risk response information when the output review result is a risk result.
9. An electronic device comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.