Content risk detection filtering method and device, electronic equipment and storage medium
By combining multiple models for content risk detection and filtering, the high cost and insufficient robustness of AI-based risk interception technology have been addressed. This approach enables efficient and accurate risk assessment and filtering, improves system security and adaptability, and optimizes user experience and compliance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-07
AI Technical Summary
Existing AI-based risk interception technologies suffer from problems such as high alignment costs, strong data dependence, insufficient generalization and robustness, lack of objectivity in the evaluation system, insufficient interpretability, and trade-offs between cost and performance, making it difficult to effectively deal with attacks in cross-domain and cross-language scenarios.
A multi-model content risk detection and filtering method is adopted. By combining attack filtering detection and risk label assessment with lenient and strict interception modes, different model strategies are used for differentiated processing, including the invocation of standard language, security big model and business big model, to achieve efficient and accurate risk assessment and filtering of input information.
It improves the system's security and flexibility, ensures that the most appropriate response measures are taken in different risk scenarios, enhances the accuracy of risk assessment and the system's adaptability, reduces operating costs, optimizes user experience and compliance, and enhances the overall performance of the system.
Smart Images

Figure CN121808765A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information security and content filtering technology, in particular, the present application relates to a content risk detection filtering method and device, electronic equipment and storage medium. BACKGROUND
[0002] Current artificial intelligence technology is developing at an unprecedented speed, and its ability in programming, reasoning and content creation has surpassed human experts in some tests. However, while this technology brings great convenience, it also poses serious security challenges.
[0003] Artificial intelligence faces the risk of generating content security and privacy leakage. The security risk of large models not only affects the development of artificial intelligence business, but also brings great compliance risk and social risk to the artificial intelligence application side.
[0004] In recent years, artificial intelligence technology represented by large language models (LLMs) has made breakthrough progress and shown great application potential in various industries. However, while LLM brings convenience, its inherent security risks are increasingly prominent.
[0005] Based on human scoring of model output, a reward model (RM) is trained, and reinforcement learning (commonly used PPO) is used to optimize the original model strategy, so that the generated results are more in line with human preferences and safety specifications.
[0006] Deep Reinforcement Learning from Human Preferences (Christiano P F, Leike J, Brown T, et al. Deep reinforcement learning from human preferences [J]. Advances in neural information processing systems, 2017, 30.) establishes the basic idea of reinforcement learning through human preferences, providing strong academic support for subsequent work.
[0007] Training language models to follow instructions (Ouyang L, Wu J, Jiang X, et al. Training language models to follow instructions with human feedback[J]. Advances in neural information processing systems, 2022, 35: 27730-27744.) uses RLHF to improve the ability of language to follow instructions, which is the core method of InstructGPT and other series.
[0008] This approach can significantly improve the compliance and safety of the model to instructions, but the cost is high, the quality of data is uneven, and the diversity of the model's results is large, and it is still bypassed by adversarial prompts.
[0009] The dialogue governance mode of constitutional artificial intelligence can make the model value judge, correct or reject the content in the dialogue through the set rules and specifications, so as to realize the clear goal-oriented self-restraint and safety improvement.
[0010] Constitutional ai: Harmlessness from ai feedback (Bai Y, Kadavath S, Kundu S, et al. Constitutional ai: Harmlessness from ai feedback[J]. arXiv preprint arXiv:2212.08073, 2022.) realizes the governance output of "constitutional artificial intelligence corresponding rules" in the dialogue framework, reduces the generation probability of risky content, and improves the self-restraint ability of sensitive information.
[0011] This approach has higher explainability and controllability through constitutional artificial intelligence, and is more robust in cross-domain and cross-language applicability, but the design, update and coverage of constitutional artificial intelligence corresponding clauses are still challenges, and need to work with human preference data to avoid excessive conservatism.
[0012] In summary, the existing artificial intelligence risk interception has the following problems: High cost and strong data dependence: RLHF / DRLHF requires high-quality artificial annotation data and continuous human review, which is costly and difficult to scale to cover all scenarios.
[0013] Insufficient generalization and robustness: The robustness of adversarial hints, hint injection, and cross-domain / cross-language scenarios is still insufficient, making it easy to be breached by "bypass" or "jailbreak" attacks.
[0014] Challenges of the evaluation system: There is a lack of comprehensive, objective and repeatable evaluation benchmarks, making it difficult to quantify the effectiveness and side effects of fences in different domains and user groups (such as excessive conservatism leading to a decrease in output usability).
[0015] Insufficient explainability and transparency: The multi-layered governance chain is complex to implement, and it is difficult for outsiders to reliably explain why outputs are rejected or corrected, affecting user trust and the traceability of compliance audits.
[0016] Cost and performance trade-offs: Introducing multi-layered screening, RM, PPO updates, etc., will bring system latency, computing costs, and deployment complexity, requiring a trade-off between security and performance.
[0017] Construction and maintenance challenges: Continuously updating the constitutional provisions on artificial intelligence, risk classifications, and adaptations to emerging fields (such as privacy protection, copyright compliance, and data source transparency) requires a high level of engineering and governance capabilities. Summary of the Invention
[0018] The technical problem to be solved by the present invention is to provide a content risk detection and filtering method, apparatus, electronic device and storage medium, which aims to solve at least one of the above-mentioned technical problems.
[0019] Firstly, the technical solution of the present invention to solve the above-mentioned technical problems is as follows: a content risk detection and filtering method, the method comprising: Obtain the input information to be processed; The input information to be processed is subjected to attack filtering detection. The information that passes the attack filtering detection is subjected to input content risk detection to determine the risk label of the input information to be processed. The risk label is used to characterize the risk level of the input information to be processed. The information that has passed the attack filtering detection is subjected to content risk detection using a processing strategy corresponding to the risk label, and the output information is obtained.
[0020] The beneficial effects of this invention are as follows: By constructing a content risk detection and filtering method combining multiple models, this invention achieves efficient and accurate risk assessment and filtering of input information. First, after acquiring the input information to be processed, attack filtering detection is performed, which can effectively identify and block potential malicious content, thereby ensuring the security and stability of the system. Subsequently, the information that passes the attack filtering detection undergoes input content risk detection, and risk labels are determined. Finally, a processing strategy corresponding to the risk labels is used for content risk detection to obtain output information. This differentiated processing strategy based on risk levels not only improves the system's flexibility and adaptability but also ensures that the most appropriate countermeasures can be taken in different risk scenarios.
[0021] Based on the above technical solution, the present invention can be further improved as follows.
[0022] Furthermore, the aforementioned processing strategy corresponding to the risk label is used to perform content risk detection on the information detected by the attack filter, resulting in output information including: If the risk label indicates a high risk level, the output information is generated using standard language. If the risk label represents a medium risk level, a large security model is used to replace the information detected by the attack filter to obtain the output information. If the risk label indicates a safe risk level, the business big data model is used to respond to the information detected by the attack filter and obtain the output information.
[0023] Furthermore, if the risk label indicates a safe risk level, the business big data model is used to respond to the information detected by the attack filter, resulting in output information including: If the risk label represents a safe risk level, the business big model is used to respond to the information detected by the attack filter to obtain the response information; The system performs output content filtering and detection on the response information. If the response information contains sensitive information, it is de-identified and the output information is obtained. If the response information contains risky content, the response is rejected. Output information is generated based on standard scripts. If the response information passes the output content filtering and detection, it is used as the output information.
[0024] Furthermore, the input information to be processed undergoes attack filtering and detection, including: By using a pre-trained attack detection model, attack risks in the input information to be processed can be identified.
[0025] Furthermore, the method also includes: Configure different interception modes, including lenient interception mode, moderately lenient interception mode, strict interception mode, and moderately strict interception mode. Each interception mode corresponds to a different model invocation strategy. Based on the risk label, determine the interception mode corresponding to the risk label; Based on the interception mode corresponding to the risk label, determine the calling strategy of the lenient classification model and the strict classification model to perform input content risk detection.
[0026] Furthermore, based on the interception mode corresponding to the risk label, the above-mentioned invocation strategies for the lenient classification model and the strict classification model are determined, including: In the lenient interception mode, only the lenient classification model is used for detection; In the more lenient interception mode, the lenient classification model is called first, and the strict classification model is called for secondary detection of content whose detection results are within the preset edge confidence interval. In strict interception mode, only the strict classification model is used for detection; In a stricter interception mode, the strict classification model is used first for detection, and the lenient classification model is used for risk downgrading and review of content that is determined to be high-risk.
[0027] Secondly, to solve the above-mentioned technical problems, the present invention also provides a content risk detection and filtering device, the device comprising: The acquisition module is used to acquire the input information to be processed. The first detection module is used to perform attack filtering detection on the input information to be processed, perform input content risk detection on the information that passes the attack filtering detection, and determine the risk label of the input information to be processed. The risk label is used to characterize the risk level of the input information to be processed. The second detection module is used to perform content risk detection on information that has passed the attack filtering detection using a processing strategy corresponding to the risk label, and to obtain output information.
[0028] Thirdly, in order to solve the above-mentioned technical problems, the present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the content risk detection and filtering method of the present application.
[0029] Fourthly, in order to solve the above-mentioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the content risk detection and filtering method of the present application.
[0030] Additional aspects and advantages of this application will be set forth in part in the description which follows, and will become apparent from the description or may be learned by practice of this application. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below.
[0032] Figure 1 This is a flowchart illustrating a content risk detection and filtering method according to an embodiment of the present invention. Figure 2 This is a flowchart illustrating another content risk detection and filtering method provided in an embodiment of the present invention; Figure 3 A schematic diagram of a system architecture provided for one embodiment of the present invention; Figure 4 This is a schematic diagram of the structure of a content risk detection and filtering device according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of an electronic device provided in one embodiment of the present invention. Detailed Implementation
[0033] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0034] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.
[0035] The solution provided in this invention can be applied to any application scenario requiring risky content detection and filtering. The solution provided in this invention can be executed by any electronic device, such as a user's terminal device, including at least one of the following: smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, smart TV, or smart in-vehicle device.
[0036] This invention provides a possible implementation, such as... Figure 1 The diagram shows a flowchart of a content risk detection and filtering method. This method can be executed by any electronic device, such as a terminal device, or by both a terminal device and a server. For ease of description, the method provided in this embodiment will be described below using a terminal device as the execution subject. Figure 1The flowchart shown indicates that the method may include the following steps: S10, Obtain the input information to be processed; S20, perform attack filtering detection on the input information to be processed, perform input content risk detection on the information that passes the attack filtering detection, and determine the risk label of the input information to be processed. The risk label is used to characterize the risk level of the input information to be processed. S30 employs a processing strategy corresponding to the risk label to perform content risk detection on the information detected by the attack filter and obtains output information.
[0037] This invention constructs a content risk detection and filtering method combining multiple models, achieving efficient and accurate risk assessment and filtering of input information. First, after acquiring the input information to be processed, attack filtering detection is performed, effectively identifying and blocking potential malicious content, thereby ensuring the security and stability of the system. Subsequently, the information that passes the attack filtering detection undergoes input content risk detection, and risk labels are determined. Finally, a processing strategy corresponding to the risk labels is applied to perform content risk detection, yielding output information. This risk-level-based differentiated processing strategy not only improves the system's flexibility and adaptability but also ensures that the most appropriate countermeasures can be taken in different risk scenarios.
[0038] The following specific embodiments further illustrate the solution of the present invention. The present invention proposes a risk content detection and filtering system through multi-model risk identification. It mainly utilizes the different risk sensitivities of different models and keyword databases to classify and filter risk content. This solution enables diversified risk judgment of risk content and ensures the security and compliance of artificial intelligence input and output content.
[0039] On the other hand, it enables the identification of multiple risk types for the input and output content of the model (the risk types cover the risk classification of generative content and the identification of prompt word attack content in the "Basic Requirements for Security of Generative Artificial Intelligence Services"), thereby providing security protection for the input and output content of the model and ensuring the security of the content generated by the large model.
[0040] Based on this, in this embodiment, combined with Figure 2 The provided content risk detection and filtering method may include the following steps: S10, Obtain the input information to be processed. After the input information to be processed is processed by the solution of this application, it is necessary to give the answer information (output information). The input information to be processed refers to text content that needs to undergo risk detection and filtering. This content is usually input by the user and may include natural language text, instructions, questions, or any other form of text data. Specifically, it can include the following common scenarios: User input in the intelligent customer service system: The text content entered by the user when asking questions or sending instructions to the intelligent customer service system. For example, "I want to check my account balance" or "Please book a flight to Shanghai for tomorrow".
[0041] User messages in a chatbot: Text messages sent by users when interacting with the chatbot. For example, "How's the weather today?" or "Tell me a joke."
[0042] Content to be published in a content management system: In a content management system, this refers to content that users or editors are preparing to publish, such as news articles, blog posts, and comments.
[0043] User instructions in a financial trading system: Trading instructions or query requests entered by users in a financial trading system, such as "buy 100 shares of a certain company's stock" or "query my trading records".
[0044] Patient questions in a medical consultation system: Questions that patients ask doctors or smart assistants in a medical consultation system, such as "I've been feeling unwell lately, what could be the reason?".
[0045] Any text input requiring security assessment: In various application scenarios, any text content entered by the user, as long as it requires security and compliance assessment, can be considered "input information to be processed".
[0046] S20, perform attack filtering detection on the input information to be processed, perform input content risk detection on the information that passes the attack filtering detection, and determine the risk label of the input information to be processed. The risk label is used to characterize the risk level of the input information to be processed. Attack filtering detection refers to a preliminary security check performed on the input information to be processed, with the aim of identifying and blocking potentially offensive or malicious input content. This detection typically targets attempts to attack the system through input information, such as injection attacks, malicious code, SQL injection, and cross-site scripting (XSS) attacks. Through attack filtering detection, the system can proactively prevent these potential attacks, thereby protecting the system's security and stability.
[0047] Optionally, one implementation method for attack filtering and detection of the input information to be processed is as follows: By using pre-trained attack detection models (such as cue word attack detection models), attack risks in the input information to be processed can be identified.
[0048] Alternatively, another implementation of the above-mentioned attack filtering and detection of the input information to be processed is as follows: The input information to be processed is filtered and detected for attacks using a pre-built risk lexicon.
[0049] Alternatively, another implementation of the above-mentioned attack filtering and detection of the input information to be processed is as follows: First, the input information to be processed is subjected to attack filtering and detection using a pre-built risk lexicon. Then, a pre-trained attack detection model is used to identify the attack risks in the information after the attack filtering and detection by the risk lexicon.
[0050] The risk lexicon is a predefined set of keywords containing words, phrases, or patterns that may indicate risk or sensitive content. It is primarily used to quickly identify whether there are obvious risk signals in the input content. By constructing the risk lexicon and risk warning lexicon, an outer interception system is formed to ensure initial risk assessment of the content. The attack detection model is a machine learning or deep learning-based model used to identify aggressive patterns or potential malicious behavior in the input content. It can detect more complex and covert attack methods, such as injection attacks and cross-site scripting attacks.
[0051] S30 employs a processing strategy corresponding to the risk label to perform content risk detection on the information detected by the attack filter and obtains output information.
[0052] Different risk labels correspond to different processing strategies. Specifically, one implementation of S30 is as follows: If the risk label indicates a high risk level, the output information is generated using standard language. If the risk label represents a medium risk level, a large security model is used to replace the information detected by the attack filter to obtain the output information. If the risk label indicates a safe risk level, the business big data model is used to respond to the information detected by the attack filter and obtain the output information.
[0053] In this context, standard responses refer to pre-defined text with fixed semantics. For high-risk input information, to prevent it from harming the user or system, the system may refuse to respond and instead use standard responses.
[0054] The "security big data model" refers to a specially designed and trained artificial intelligence model used to process and respond to medium-risk input information, ensuring that the generated output meets both security standards and basic user needs. As an example, suppose a user enters the following medium-risk content into an intelligent customer service system: "I want to check my account balance. My account number is 123456789." The security big data model's processing: The model will recognize that the account information is sensitive content and will not directly respond with the specific account balance. Instead, it will generate a safer response, such as: "Hello, to protect your account security, please log in to your account to check your balance." In this way, the security big data model ensures the security of the response while also providing useful guidance to the user, balancing risk control and user experience.
[0055] The business big model is typically used to handle low-risk or safe input, with the primary goal of providing high-quality, efficient services that meet the user's business needs. It prioritizes the accuracy and relevance of the content over risk control. As an example, in an intelligent customer service system, assuming a user enters the following: "I want to check my order status, order number ABC123456," the business big model generates a detailed and business-relevant answer based on the query results, ensuring the answer is accurate, useful, and meets the user's expectations. The generated answer might look like this: "Hello, your order ABC123456 has been shipped and is expected to arrive within 3 days. Thank you for your patience!"
[0056] Optionally, if the risk level represented by the risk label is safe, the business big data model is used to respond to the information detected by the attack filter, and the output information includes: If the risk label represents a safe risk level, the business big model is used to respond to the information detected by the attack filter to obtain the response information; The system performs output content filtering and detection on the response information. If the response information contains sensitive information, it is de-identified and the output information is obtained. If the response information contains risky content, the response is rejected. Output information is generated based on standard scripts. If the response information passes the output content filtering and detection, it is used as the output information.
[0057] "Sensitive information" refers to information that requires special protection and may not be disclosed or revealed without authorization. This information typically involves personal privacy, trade secrets, content restricted by laws and regulations, or other data that may negatively impact individuals, organizations, or society. Protecting sensitive information is a crucial component of information security, especially in scenarios involving user data, financial services, and medical records. Anonymization refers to processing sensitive information so that it can still be used for analysis, presentation, or other legitimate purposes without disclosing the original information. The goal of anonymization is to ensure data usability while protecting privacy and security. Specifically, anonymization methods can include replacing sensitive information and encrypting sensitive information.
[0058] Response information filtering and detection refers to the process of performing a series of security and compliance checks on response information after it has been generated. This ensures that the response content does not disclose sensitive information, complies with laws, regulations, and industry standards, and does not pose any other potential risks. This process is a crucial component of a content risk detection and filtering system, ensuring that the information ultimately delivered to users is safe, reliable, and compliant.
[0059] Optionally, the method further includes: Configure different interception modes, including lenient interception mode, moderately lenient interception mode, strict interception mode, and moderately strict interception mode. Each interception mode corresponds to a different model invocation strategy. Based on the risk label, determine the interception mode corresponding to the risk label; Based on the interception mode corresponding to the risk label, determine the calling strategy of the lenient classification model and the strict classification model to perform input content risk detection.
[0060] By combining risk-sensitive and risk-lenient models, four different risk rules are formed, including: lenient interception mode, relatively lenient interception mode, strict interception mode, and relatively strict interception mode, to ensure that risky content can be effectively handled under different interception rules.
[0061] Optionally, the above-mentioned determination of the interception mode corresponding to the risk label based on the risk label includes: If the risk label is low risk, select the lenient blocking mode. If the risk label is medium risk, select the relatively lenient blocking mode. If the risk label is high risk, select the strict blocking mode. If the risk label is very high risk, select the relatively strict blocking mode.
[0062] In cases of extremely high risk, standard scripts can also be used to generate output information.
[0063] Optionally, the above-mentioned strategy for determining the invocation of the lenient classification model and the strict classification model based on the interception mode corresponding to the risk label includes: In the lenient interception mode, only the lenient classification model is used for detection; In the more lenient interception mode, the lenient classification model is called first, and the strict classification model is called for secondary detection of content whose detection results are within the preset edge confidence interval. In strict interception mode, only the strict classification model is used for detection; In a stricter interception mode, the strict classification model is used first for detection, and the lenient classification model is used for risk downgrading and review of content that is determined to be high-risk.
[0064] See Figure 3 The system architecture diagram shown in this solution illustrates that, in order to achieve risk detection of the input information to be processed, pre-built risk detection strategies can also be adopted. For example, attack detection, sensitive issue detection, general violation detection, and custom violation detection can be used for attack filtering detection.
[0065] Optionally, in this application, gateway configuration can be performed in advance. By configuring the access method and the format of the request and response content of the protection model, content security filtering can be seamlessly achieved.
[0066] Optionally, in this application, service monitoring configuration can also be performed in advance to identify and block the distribution of risky content in real time through monitoring the gateway, and to manage the risky content.
[0067] To better illustrate and understand the principle of the method provided by this invention, the following description uses an optional specific embodiment to illustrate the solution of this invention. It should be noted that the specific implementation of each step in this specific embodiment should not be construed as a limitation of the solution of this invention. Other implementations that can be conceived by those skilled in the art based on the principle of the solution provided by this invention should also be considered within the scope of protection of this invention.
[0068] The configured gateway receives the content to be filtered by artificial intelligence and processes the user input.
[0069] 1. User input is processed through an attack filtering model to detect any potential attack risks; 2. If it passes, the risk content will be checked and the corresponding risk label will be returned; if it fails, the business party will be refused a response.
[0070] 3. For medium-risk content in the risk tags, a security model will provide a proxy answer. For high-risk content, a rejection mechanism will be implemented. If the content is deemed safe, a business model will provide the answer.
[0071] 4. Implement risk filtering for the content answered by the business model. If sensitive information is found, it should be anonymized. If risky content is found, the answer should be rejected. If the content is safe, the information should be output.
[0072] The solution of the present invention has the following beneficial effects: 1. Enhanced content security: This invention uses a multi-model combination approach to perform multi-level risk detection and filtering on input information, which can effectively identify and block potential malicious content and sensitive information, thereby significantly improving system security and preventing malicious attacks and data leaks.
[0073] 2. Enhanced Risk Assessment Accuracy: By utilizing a risk terminology database for initial screening and combining it with risk-sensitive and risk-lenient models for in-depth detection, this invention can more comprehensively and accurately assess the risk level of input information. This multi-dimensional risk assessment mechanism effectively reduces the false positive rate and ensures the reliability of risk detection.
[0074] 3. Optimize User Experience: By employing differentiated processing strategies based on different risk levels, this invention ensures both security and user experience. For example, for low-risk content, the system can respond quickly and provide accurate answers; for medium- to high-risk content, it generates answers that meet security standards through a comprehensive security model, avoiding direct rejection of user requests, thus achieving a good balance between security and usability.
[0075] 4. Enhanced System Flexibility and Adaptability: By configuring different interception modes, this invention can flexibly adjust model invocation strategies based on different business scenarios and risk assessment results. This flexibility enables the system to better adapt to various complex application scenarios, effectively addressing different risk challenges in fields such as intelligent customer service, content management, and financial transactions.
[0076] 5. Reduced Operating Costs and Management Complexity: The multi-model content risk detection and filtering method of this invention achieves automated risk assessment and processing, reducing reliance on manual review and thus lowering operating costs. Simultaneously, the systematic risk management and interception mode configuration simplifies management processes and improves management efficiency.
[0077] 6. Promoting Compliance and Privacy Protection: When processing sensitive information, this invention ensures that the output content complies with relevant laws, regulations, and privacy protection requirements through anonymization and strict risk control. This not only helps enterprises avoid compliance risks but also enhances users' trust in the system.
[0078] 7. Improved overall system performance: Through reasonable model division of labor and interception mode selection, this invention optimizes system resource allocation and improves processing efficiency. When faced with a large number of input requests, the system can quickly and accurately complete risk detection and filtering, ensuring stable system operation.
[0079] In summary, the solution of this invention has significant beneficial effects in improving content security, enhancing the accuracy of risk assessment, optimizing user experience, improving system flexibility and adaptability, reducing operating costs and management difficulty, promoting compliance and privacy protection, and improving overall system performance, providing comprehensive and effective protection for the safe operation of artificial intelligence applications.
[0080] Based on and Figure 1 Based on the same principle as the method shown, this embodiment of the invention also provides a content risk detection and filtering device 20, such as... Figure 4 As shown, the content risk detection and filtering device 20 may include an acquisition module 210, a first detection module 220, and a second detection module 230, wherein: The acquisition module 210 is used to acquire the input information to be processed; The first detection module 220 is used to perform attack filtering detection on the input information to be processed, perform input content risk detection on the information that passes the attack filtering detection, and determine the risk label of the input information to be processed. The risk label is used to characterize the risk level of the input information to be processed. The second detection module 230 is used to perform content risk detection on the information detected by the attack filter using a processing strategy corresponding to the risk label, and to obtain output information.
[0081] Optionally, when the second detection module 230 performs content risk detection on the information detected by the attack filter using a processing strategy corresponding to the risk label, and obtains the output information, it is specifically used for: If the risk label indicates a high risk level, the output information is generated using standard language. If the risk label represents a medium risk level, a large security model is used to replace the information detected by the attack filter to obtain the output information. If the risk label indicates a safe risk level, the business big data model is used to respond to the information detected by the attack filter and obtain the output information.
[0082] Optionally, when the risk level represented by the risk label is "safe," the second detection module 230 uses a large business model to respond to the information detected by the attack filter and obtains the output information, specifically for: If the risk label represents a safe risk level, the business big model is used to respond to the information detected by the attack filter to obtain the response information; The system performs output content filtering and detection on the response information. If the response information contains sensitive information, it is de-identified and the output information is obtained. If the response information contains risky content, the response is rejected. Output information is generated based on standard scripts. If the response information passes the output content filtering and detection, it is used as the output information.
[0083] Optionally, when performing attack filtering detection on the input information to be processed, the first detection module 220 is specifically used for: A pre-trained prompt attack detection model is used to identify attack risks in the input information to be processed.
[0084] Optionally, the device further includes: The configuration module is used to configure different interception modes, including lenient interception mode, moderately lenient interception mode, strict interception mode, and moderately strict interception mode. Each interception mode corresponds to a different model invocation strategy. Based on the risk label, the module determines the interception mode corresponding to the risk label. Based on the interception mode corresponding to the risk label, the module determines the invocation strategy of the lenient classification model and the strict classification model to perform input content risk detection.
[0085] Optionally, when determining the invocation strategy for the lenient classification model and the strict classification model based on the interception mode corresponding to the risk label, the above configuration module is specifically used for: In the lenient interception mode, only the lenient classification model is used for detection; In the more lenient interception mode, the lenient classification model is called first, and the strict classification model is called for secondary detection of content whose detection results are within the preset edge confidence interval. In strict interception mode, only the strict classification model is used for detection; In a stricter interception mode, the strict classification model is used first for detection, and the lenient classification model is used for risk downgrading and review of content that is determined to be high-risk.
[0086] The content risk detection and filtering device of this invention can execute the content risk detection and filtering method provided in this invention. The implementation principle is similar. The actions performed by each module and unit in the content risk detection and filtering device in each embodiment of this invention correspond to the steps in the content risk detection and filtering method in each embodiment of this invention. For detailed functional descriptions of each module of the content risk detection and filtering device, please refer to the descriptions in the corresponding content risk detection and filtering methods shown above. They will not be repeated here.
[0087] The aforementioned content risk detection and filtering device can be a computer program (including program code) running on a computer device, such as an application software; the device can be used to execute the corresponding steps in the method provided in the embodiments of the present invention.
[0088] In some embodiments, the content risk detection and filtering device provided in this invention can be implemented using a combination of hardware and software. As an example, the content risk detection and filtering device provided in this invention can be a processor in the form of a hardware decoding processor, which is programmed to execute the content risk detection and filtering method provided in this invention. For example, the processor in the form of a hardware decoding processor can be one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.
[0089] In other embodiments, the content risk detection and filtering device provided in this invention can be implemented in software. Figure 4 A content risk detection and filtering device stored in a memory is shown. It can be software in the form of programs and plug-ins, and includes a series of modules, including an acquisition module 210, a first detection module 220 and a second detection module 230, for implementing the content risk detection and filtering method provided in the embodiments of the present invention.
[0090] The modules described in the embodiments of the present invention can be implemented in software or hardware. The names of the modules are not, in some cases, limiting the scope of the module itself.
[0091] Based on the same principles as the methods shown in the embodiments of the present invention, the embodiments of the present invention also provide an electronic device, which may include, but is not limited to: a processor and a memory; the memory for storing computer programs; and the processor for executing the methods shown in any embodiment of the present invention by invoking the computer programs.
[0092] In one alternative embodiment, an electronic device is provided, such as Figure 5 As shown, Figure 5The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of the present invention.
[0093] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this invention. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0094] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0095] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0096] The memory 4003 stores application code (computer program) for executing the present invention, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.
[0097] Among these, electronic devices can also be terminal devices. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of the present invention.
[0098] This invention provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.
[0099] According to another aspect of the present invention, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various embodiments described above.
[0100] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0101] It should be understood that the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of methods and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0102] The computer-readable storage medium provided in this invention can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0103] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the method shown in the above embodiments.
[0104] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.
Claims
1. A content risk detection and filtering method, characterized in that, include: Obtain the input information to be processed; The input information to be processed is subjected to attack filtering detection, and the information that passes the attack filtering detection is subjected to input content risk detection to determine the risk label of the input information to be processed. The risk label is used to characterize the risk level of the input information to be processed. The information that has passed the attack filter detection is subjected to content risk detection using a processing strategy corresponding to the risk label, and output information is obtained.
2. The method according to claim 1, characterized in that, The process employs a processing strategy corresponding to the risk label to perform content risk detection on the information detected by the attack filter, and obtains output information, including: If the risk label represents a high-risk level, output information is generated using standard scripts. If the risk level represented by the risk label is medium risk, the security big model is used to replace the information detected by the attack filter to obtain the output information. If the risk level represented by the risk label is safe, the business big model is used to respond to the information detected by the attack filter to obtain output information.
3. The method according to claim 2, characterized in that, If the risk label represents a safe risk level, the business big data model is used to respond to the information detected by the attack filter, and the output information includes: If the risk level represented by the risk label is safe, the business big model is used to respond to the information detected by the attack filter to obtain the response information; The response information is subjected to output content filtering detection. If the response information contains sensitive information, it is desensitized to obtain output information. If the response information contains risky content, the response is rejected. Output information is generated based on standard scripts. If the response information passes the output content filtering detection, it is used as output information.
4. The method according to any one of claims 1 to 3, characterized in that, The attack filtering and detection of the input information to be processed includes: The attack risks in the input information to be processed are identified by a pre-trained attack detection model.
5. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Configure different interception modes, including lenient interception mode, relatively lenient interception mode, strict interception mode, and relatively strict interception mode, each of which corresponds to a different model invocation strategy; Based on the risk label, determine the interception mode corresponding to the risk label; Based on the interception mode corresponding to the risk label, the calling strategy of the lenient classification model and the strict classification model is determined in order to perform risk detection on the input content.
6. The method according to claim 5, characterized in that, The step of determining the invocation strategy for the lenient classification model and the strict classification model based on the interception mode corresponding to the risk label includes: In the lenient interception mode, only the lenient classification model is used for detection; In the more lenient interception mode, the lenient classification model is called first, and the strict classification model is called for secondary detection of content whose detection results are within the preset edge confidence interval. In strict interception mode, only the strict classification model is invoked for detection; In a stricter interception mode, the strict classification model is used first for detection, and the lenient classification model is used for risk downgrading and review of content that is determined to be high-risk.
7. A content risk detection and filtering device, characterized in that, include: The acquisition module is used to acquire the input information to be processed. The first detection module is used to perform attack filtering detection on the input information to be processed, perform input content risk detection on the information that passes the attack filtering detection, and determine the risk label of the input information to be processed. The risk label is used to characterize the risk level of the input information to be processed. The second detection module is used to perform content risk detection on the information that has passed the attack filtering detection using a processing strategy corresponding to the risk label, and to obtain output information.
8. The apparatus according to claim 7, characterized in that, When the second detection module performs content risk detection on the information that has passed the attack filter detection using the processing strategy corresponding to the risk label and obtains output information, it is specifically used for: If the risk label represents a high-risk level, output information is generated using standard scripts. If the risk level represented by the risk label is medium risk, the security big model is used to replace the information detected by the attack filter to obtain the output information. If the risk level represented by the risk label is safe, the business big model is used to respond to the information detected by the attack filter to obtain output information.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the method of any one of claims 1-6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-6.
Citation Information
Patent Citations
Sensitive information detection method and device, storage medium and computer device
CN110598411A
Prompt word attack detection method and device for large language model
CN118445815A
Attack defense method and device for e-commerce intelligent customer service large model
CN120470582A
Large model content security multi-level defense method
CN120880683A