Content security reasoning auditing method based on knowledge enhancement
By constructing a standardized knowledge rule base and using supervised fine-tuning and reinforcement learning guided by teacher models, the problem of insufficient reasoning ability and high resource consumption in the detection of harmful Chinese content in existing technologies is solved. This achieves efficient and interpretable content security review, which is applicable to real-time content security governance on Internet platforms.
Patent Information
- Application Number
- CN202511731288.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-24
- Publication Date
- 2025-12-23
AI Technical Summary
Existing Chinese harmful content detection technologies lack explicit reasoning processes, knowledge support, and verifiable reasoning chains, making it difficult to explain and verify the judgment logic. Furthermore, large language models suffer from high inference latency and resource consumption, making it difficult to meet the needs of real-time deployment and large-scale content security governance.
By constructing a standardized knowledge rule base and generating implicit knowledge through teacher models, student models are guided to undergo supervised fine-tuning and reinforcement learning, thereby achieving knowledge injection and ability compensation, outputting explicit judgment reasons and basis, and improving the model's reasoning ability and interpretability.
While maintaining lightweight and high efficiency, it significantly improves detection accuracy and robustness, meets the needs of internet platforms for real-time performance, compliance and interpretability, and reduces computing and deployment costs.
Smart Images

Figure CN121189459A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of natural language processing and machine learning, and specifically to a content security reasoning and auditing method based on knowledge enhancement. Background Technology
[0002] In the early days of the internet, platforms primarily relied on human review teams to evaluate suspicious content posted by users one by one. While this method performed well in terms of accuracy, with the explosive growth of data, human review could no longer meet the requirements of "minute-level discovery and hour-level processing." Furthermore, human review suffers from high labor costs, inconsistent review standards, and the psychological damage to reviewers caused by long-term handling of inappropriate content, making it difficult to support the long-term operation of large-scale platforms.
[0003] With the development of deep learning technology, researchers have begun to explore the use of small-scale neural network models for text classification and detection. While these models possess some semantic understanding capabilities compared to keyword matching, they still suffer from significant drawbacks: their judgment process typically relies on direct output of results based on internal semantic representations, lacking explicit reasoning chains or judgment criteria, making them typical "black box decisions." In complex contexts or with evasive expressions, the models often cannot review their own judgment logic or provide clear reasons for violations, leading to difficulties in manual review and a lack of traceability and consistency in the review results.
[0004] In recent years, large language models have made significant breakthroughs in semantic understanding and cross-task generalization. Some studies have attempted to directly utilize large language models for text violation detection through prompting engineering. These methods possess zero-shot or small-shot capabilities and can achieve good detection results without additional training. However, these methods still have significant limitations: on the one hand, the judgments of large language models mainly rely on internal statistical patterns, lacking a systematic introduction of external knowledge, regulations, or review standards. The reasoning process is not based on a verifiable knowledge chain, making its judgment logic difficult to explain and verify. On the other hand, although some large language models can generate judgment reasons, these explanations are often lengthy, subjective, and lack structured expression, deviating from actual content review standards and making it difficult to directly support manual review and accountability. In addition, models with hundreds of billions of records in actual deployment have high inference costs and high response latency, making it difficult for small and medium-sized platforms to bear the computing power and compliance costs.
[0005] In summary, existing Chinese harmful content detection technologies suffer from the following shortcomings: most methods lack explicit reasoning processes, only directly outputting "violation / non-violation" conclusions, making it difficult to review the judgment logic and provide reliable evidence; existing methods fail to systematically incorporate external knowledge or domain rules, relying on internal semantic representations for model judgments, lacking knowledge support and verifiable reasoning chains; while large language models possess certain detection and interpretation capabilities, their high inference latency and resource consumption make them unsuitable for real-time deployment and large-scale content security governance. Therefore, existing technologies cannot simultaneously guarantee detection accuracy while ensuring inference transparency, knowledge traceability, and inference efficiency, failing to meet the "causally traceable and verifiable" intelligent review requirements of large-scale content security governance. Summary of the Invention
[0006] To address the shortcomings of existing technologies, particularly the general lack of reasoning ability and knowledge support in existing methods, this invention provides a content security reasoning and auditing method based on knowledge enhancement. By introducing expert knowledge bases and implicit knowledge generated by teacher models, it achieves knowledge injection and capability compensation for small-scale pre-trained models, enabling them to possess reasoning ability and interpretability while maintaining lightweight and high efficiency.
[0007] To achieve the above-mentioned objectives, an embodiment provides a content security reasoning auditing method based on knowledge enhancement, comprising: A standardized knowledge rule base is constructed by content review experts based on internet content security standards; An attribute system is established based on the violations in the standardized knowledge rule base, and then the attribute system is structured into an attribute set; A prompt template generation function is designed based on attribute sets to generate prompt text. The prompt text is used to call the teacher model to infer and generate candidate answers. A synthetic dataset is built using candidate answers and a standardized knowledge rule base. Synthetic datasets are used to guide student models in supervised fine-tuning to obtain content moderation category judgments and explanatory texts. Then, reinforcement learning is used to update the category judgments and explanatory texts generated by the student models to train and optimize the student models. The trained and optimized student model is deployed to the Chinese Internet content review scenario. Taking any content as input, it generates reasoning explanations for manual review and accountability, and finally outputs violation category tags to complete the content security reasoning review.
[0008] In one embodiment, the standardized knowledge rule base includes categories of content that violates and does not violate the rules. For each content category, content review experts, based on domain knowledge and in conjunction with internet content security standards, mark the corresponding rules and judgment points to determine the identification standards and triggering conditions for violating content.
[0009] In one embodiment, the attribute system includes user profile features, text features, avoidance strategy features, and knowledge rule features; The user profile features include gender, age, occupation, and education level, which are used to simulate different writing styles and contexts; The text features include text length, narrative perspective, and publishing platform, which are used to reflect differences in content format; The avoidance strategy features avoidance detection methods that include emoticons, homophones, and variant symbols, used to reproduce hidden expressions; The knowledge rule feature is a standardized knowledge rule base that is constructed to indicate the specific norms violated by the content text.
[0010] In one embodiment, the design process of the prompt template generation function includes: setting up a prompt for the prompt template generation function, and providing task options and generation requirements based on an attribute set to guide the generation of prompt text; The prompt is used to clearly indicate that the role of the prompt template generation function is a content review expert. The task options based on the attribute set include: user profile task options including gender, age, occupation, and education level; text feature task options including whether it violates regulations, violation category, rule violation, content text length, narrative perspective, and publishing platform; and avoidance strategy task options including avoidance methods and corresponding explanations. The generation requirements are specifically: to generate content text that conforms to the user profile task options and text feature task options, and to correctly apply the avoidance strategy task options when using them.
[0011] In one embodiment, the use of synthetic datasets to guide supervised fine-tuning of a student model to obtain content moderation category determinations and explanatory texts includes: Using violation text and a standardized knowledge rule base as input, and reasoning text and category labels generated by the teacher model as output, the student model is guided to learn the joint distribution of candidate answers and the standardized knowledge rule base. The candidate answers contain implicit knowledge, and the standardized knowledge rule base contains explicit knowledge. A text sequence loss is established to supervise and fine-tune the student model, enabling the student model to have preliminary semantic understanding and rule alignment capabilities, and output the category judgment and explanation text for content review.
[0012] In one example, the text sequence loss is represented by the cross-entropy loss function, calculated as follows: , , In the formula, For text sequence loss, To infer the length of the text sequence, For the length of the category text sequence, This indicates the reasoning and interpretation of the text in the first part. Each word element, The first in the category text Each word element, Indicates that the attribute category belongs to the synthetic dataset. The Input of each sample, This represents the operation of the parameter value to obtain the minimum value of the function. The optimal student model parameter values are those that minimize the loss function. The number of attribute categories, For attribute categories The number of samples in This represents the parameters of the student model.
[0013] In one example, group-relative policy optimization is used as a reinforcement learning optimization strategy to optimize student model parameters, specifically including: By setting a reward function Used to determine the student model's response to the input. The generated inference explanation text and category determination constitute a complete sequence. The accuracy and format of the judgment results are verified, and the final reward is output. ; For each input Sampling from the old strategy Candidate output sequences And calculate its corresponding reward group. If all rewards in this group If all are the same, discard the data in that group and repeat the sampling process until a reward group with a valid gradient signal is obtained. Based on the reward group with the effective gradient signal The objective function is defined as follows by optimizing based on the group relative strategy: , in, Indicates sample input, This indicates the corresponding answer. To follow the old strategy obtained from sampling One of the candidate output sequences; For the student model parameters of the old strategy; probability ratio For the current strategy Generate sequence The Middle The ratio of the probability of a word to the probability of that word being generated by the old strategy; advantage function estimation. For sequence Rewards Subtract groups The mean of the rewards for that group is then divided by the standard deviation of the rewards for that group. The lower bound clipping factor for decoupling. The upper limit pruning coefficient is used to decouple the policy and limit the step size of policy updates.
[0014] In one embodiment, after the trained and optimized student model is deployed to the Chinese Internet content review scenario, it takes any content as input, uses a greedy decoding strategy for inference, and generates the word corresponding to the current maximum probability value, and outputs the corresponding violation category label.
[0015] The present invention also provides a knowledge-enhanced content security reasoning and review device, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the knowledge-enhanced content security reasoning and review method.
[0016] The present invention also provides a computer-readable storage medium storing a computer program, which, when used with a computer, is used to implement the knowledge-enhanced content security reasoning and auditing method described above.
[0017] Compared with the prior art, the beneficial effects of the present invention include at least the following: (1) This invention introduces a standardized knowledge rule base based on domain knowledge construction and implicit knowledge generated by the teacher model. It guides the student model to perform supervised fine-tuning through knowledge distillation, thereby realizing knowledge injection and capability compensation for small-scale pre-trained models. Under the premise of ensuring detection accuracy and robustness, it explicitly outputs the judgment reasons and basis, and also enables it to have reasoning ability and interpretability while maintaining lightweight and high efficiency.
[0018] (2) Further optimization through reinforcement learning training can improve the accuracy of model output and the completeness and consistency of explanatory text output. It can significantly improve accuracy and robustness in the identification of multiple categories and forms of illegal information, enhance the transparency and consistency of review results, significantly reduce inference and deployment costs, and meet the comprehensive needs of Internet platforms in terms of real-time performance, compliance and interpretability. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.
[0020] Figure 1 This is a flowchart illustrating the knowledge-enhanced content security reasoning and review method provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the following description is provided in conjunction with the accompanying drawings and... The embodiments further illustrate the present invention in detail. It should be understood that the specific embodiments described herein are merely illustrative of the invention and do not limit the scope of protection of the invention.
[0022] like Figure 1 As shown, the present invention proposes a content security reasoning and auditing method based on knowledge enhancement, which includes the following steps: S1. A standardized knowledge rule base is constructed by content review experts based on Internet content security standards.
[0023] In this embodiment, content moderation experts first construct a standardized knowledge rule base based on internet content security standards. This standardized knowledge rule base covers common categories of both illegal and non-illegal content. For each category, content moderation experts summarize and annotate the corresponding rules and judgment points, clarifying the identification standards and triggering conditions for illegal content, providing standardized knowledge support for subsequent model training and inference. Specifically: S2. Establish an attribute system based on the violations in the standardized knowledge rule base, and then systematize the attribute system into an attribute set.
[0024] To address different types of infringing content, this embodiment designs and constructs a hierarchical, fine-grained attribute system to comprehensively characterize the multidimensional features of internet content text. This attribute system includes: user profile features, such as gender, age, occupation, and education level, used to simulate different writing styles and contexts; text features, including text length, narrative perspective, and publishing platform, used to reflect differences in content format; avoidance strategy features, including common avoidance detection methods such as emoticons, homophones, and variant symbols, used to reproduce hidden expressions; and knowledge rule features, which are standardized knowledge rule bases constructed by content review experts based on internet content security standards, used to indicate specific norms that internet content text may violate. These attributes are organized in a structured manner as an attribute set as input, used to prompt template construction and synthetic data generation, ensuring the model can learn multidimensional features and knowledge constraints.
[0025] S3. Design a prompt template generation function based on the attribute set to generate prompt text. Use the prompt text to call the teacher model to infer and generate candidate answers. Use the candidate answers and the standardized knowledge rule base to build a synthetic dataset.
[0026] In this embodiment, a prompt template generation function is designed based on a structured attribute set to integrate the sampled attributes with corresponding knowledge rules and generate prompt text. The specific steps are as follows: For each attribute category... The following samples Build attribute set ,in Several attributes are randomly selected from the aforementioned attribute system to comprehensively describe the multidimensional features of the sample. The attribute set... An input prompt template generation function is used to guide the teacher model to generate candidate outputs that conform to attribute and rule constraints. The design process of the prompt template generation function includes: establishing a personality prompt for the prompt template generation function, and providing task options and generation requirements based on an attribute set to guide the generation of prompt text; wherein, the personality prompt clarifies that the role of the prompt template generation function is a content review expert; the task options based on the attribute set include: user profile task options including gender, age, occupation, and education level; text feature task options including whether it violates regulations, violation category, rule violation, content text length, narrative perspective, and publishing platform; and avoidance strategy task options including avoidance methods and corresponding explanations; the generation requirements specifically require generating content text that conforms to the user profile task options and text feature task options, and correctly applying the avoidance strategy task options when used.
[0027] Next, the teacher model is invoked. Reasoning based on the prompt text to generate candidate answers . Indicates the teacher model in the sample Attribute Category The candidate outputs generated contain implicit knowledge and can be used to train student models, enabling them to learn the reasoning abilities of teacher models while following the rules.
[0028] Regarding the teacher model, this invention employs Deep Search, generating knowledge-enhanced synthetic training data through its API interface. During data generation, the teacher model's temperature parameter is set to 1.0, and the top-k value is 1 to ensure output diversity and coverage of potential variant expressions. 3000 synthetic data points are sampled for each violation category to ensure the dataset's class balance and diversity, thereby improving the training effect and generalization ability of the student model.
[0029] After data generation, this invention performs rigorous quality control on candidate samples, eliminating duplicates, invalid samples, and samples rejected by the teacher model. Subsequently, balanced sampling is performed according to categories to ensure a balance and diversity in sample quantity, style, and features across different categories. The high-quality samples selected constitute the final synthetic dataset, used in the subsequent student model training phase.
[0030] S4. Use synthetic datasets to guide the student model for supervised fine-tuning to obtain content moderation category judgments and explanation texts. Then, use reinforcement learning to update the category judgments and explanation texts generated by the student model to train and optimize the student model.
[0031] This invention uses the Qwen-2.5 series language models as student models, specifically including two versions with different parameter scales: Qwen-2.5-3B and Qwen-2.5-7B. Both can be rapidly deployed and efficiently inferred in resource-constrained hardware environments.
[0032] Next, after the synthetic dataset is built, it is divided, with one part used for supervised fine-tuning (SFT) and the other for reinforcement learning (RL). For the data used for SFT, it needs to be based on the teacher model output. Generate corresponding reasoning text This is used to guide student models in learning the joint distribution of explicit rules implied in a standardized knowledge rule base and implicit knowledge from the teacher model. For data used in RL, it is not necessary to generate inference text. This is used to improve the robustness and consistency of the model in determining categories and interpreting outputs.
[0033] Regarding the training framework, this invention employs different frameworks for different training stages. The SFT stage uses the LLaMA-Factory framework to support efficient fine-tuning and distributed training of small-scale language models; the RL stage uses the Verl framework to meet the reinforcement optimization training requirements based on Group Relative Policy Optimization (DAPO). The training environment is configured as a single-machine multi-GPU solution, specifically including 8 computing chips, each with 80 bytes of GPU memory, thereby ensuring the stability, efficiency, and reproducibility of the training process.
[0034] Furthermore, this invention employs a joint training paradigm of "supervised fine-tuning + reinforcement learning": The SFT stage uses distilled data to guide the student model in learning the joint distribution of explicit rules R and implicit knowledge A from the teacher model, outputting a category decision. and explanatory text The optimization objective is text sequence loss, represented by the cross-entropy loss function, to enable the model to possess preliminary semantic understanding and rule alignment capabilities. The loss function is calculated as follows: , , In the formula, For text sequence loss, To infer the length of the text sequence, For the length of the category text sequence, This indicates the reasoning and interpretation of the text in the first part. Each word element, The first in the category text Each word element, Indicates that the attribute category belongs to the synthetic dataset. The Input of each sample, This represents the operation of the parameter value to obtain the minimum value of the function. The optimal student model parameter values are those that minimize the loss function. The number of attribute categories, For attribute categories The number of samples in This represents the parameters of the student model.
[0035] Furthermore, hyperparameters were strictly configured to ensure training stability and reproducibility. Specifically, the batch size per card was set to 4, the gradient accumulation step count was 2, equivalent to a global batch size of 64; and the learning rate was set to 1.0 × 10⁻⁶. -5 A cosine annealing scheduling strategy with a warm-up ratio of 0.1 was employed to smoothly adjust the learning rate and avoid instability in the early stages of student model training. The AdamW optimizer was used with a weight decay coefficient of 0.01 to reduce the risk of overfitting. The number of training epochs was set to 3. During training, a bfloat16 mixed precision strategy was adopted to improve training speed and reduce memory usage while maintaining computational accuracy. Cross-entropy was used as the loss function, and dropout=0.1 was introduced into the network as a regularization measure to further prevent overfitting.
[0036] The RL phase, based on SFT, employs DAPO as a reinforcement learning optimization strategy to enhance the robustness of the student model in class determination and interpretation output. The reward function comprehensively considers: (i) class determination accuracy; (ii) format accuracy. Specifically, this includes: By setting a reward function Used to determine the student model's response to the input. The generated inference explanation text and category determination constitute a complete sequence. The accuracy and format of the judgment results are verified, and the final reward is output. ; For each input Sampling from the old strategy Candidate output sequences And calculate its corresponding reward group. If all rewards in this group If all are the same, discard the data in that group and repeat the sampling process until a reward group with a valid gradient signal is obtained. Based on the reward group with the effective gradient signal The objective function is defined as follows by optimizing based on the group relative strategy: , in, Indicates sample input, This indicates the corresponding answer. To follow the old strategy obtained from sampling One of the candidate output sequences; For the student model parameters of the old strategy; probability ratio For the current strategy Generate sequence The Middle The ratio of the probability of a word to the probability of that word being generated by the old strategy; advantage function estimation. For sequence Rewards Subtract groups The mean of the rewards for that group is then divided by the standard deviation of the rewards for that group. The lower bound clipping factor for decoupling. The upper limit pruning coefficient is used to decouple the policy and limit the step size of policy updates.
[0037] The student model updates its predictions and interpretations based on the reward signal, without using a KL penalty term, making the model training more exploratory. Training sampling and batch parameter settings are as follows: 16 responses are generated per cue, the training batch size is 512, the training microbatch size is 32, and the learning rate is set to 1×10⁻⁶. -6 A 10-step learning rate warm-up is used in the initial training phase. The training environment maintains the same single-machine multi-GPU solution as the SFT phase, supporting a maximum input / output length of up to 32k to ensure training stability and efficiency.
[0038] During the model inference phase, this invention uniformly employs a greedy decoding strategy and sets the temperature parameter to 0 to ensure the stability and consistency of the output results. Under this strategy, each prediction generated by the student corresponds to the word with the highest probability at the current time, thus ensuring that the output results are repeatable under the same input, facilitating reliable deployment in actual content review scenarios.
[0039] Through the above configuration, the present invention achieves a balance between training efficiency, inference speed and detection performance in small-scale models, providing an efficient, stable and reproducible implementation scheme for practical applications, thereby improving the model's inference ability and output robustness.
[0040] S5. Deploy the trained and optimized student model to the Chinese Internet content review scenario. With any content as input, generate reasoning explanations for manual review and accountability. Finally, output violation category tags to complete the content security reasoning review.
[0041] During the inference and deployment phases, the model output order is fixed as follows: first, an inference explanation is generated (e.g., "The text contains the violating word 'high-paying part-time job,' which conforms to rule R2"), followed by the final violation category label. This explanation result is returned in a structured form for easy manual review and accountability.
[0042] The embodiment also provides a knowledge-enhanced content security reasoning review device, including a memory and one or more processors. The memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the knowledge-enhanced content security reasoning review method.
[0043] The embodiment also provides a computer-readable storage medium storing a computer program, which, when used with a computer, is used to implement the knowledge-enhanced content security reasoning and auditing method.
[0044] This invention achieves significant improvements in technical effectiveness. By introducing a knowledge-enhanced harmful content detection method, it enables reasoning-based content security review even on medium-sized language models (such as Qwen-2.5-3B and 7B). This invention verifies that detection performance comparable to ultra-large models can be achieved with relatively small models, effectively resolving the contradiction between detection accuracy and computational cost in existing technologies. Experimental results show that the method of this invention can significantly improve accuracy and robustness in the identification of multi-category and multi-form illegal information, demonstrating engineering feasibility and broad application value.
[0045] This invention enhances the ability to identify and intercept various types of harmful information on the network, enabling efficient monitoring and control of illegal and other sensitive content. By applying knowledge augmentation strategies to a medium-sized language model, this invention significantly improves the accuracy and robustness of detecting concealed illegal content in complex contexts, assisting internet platforms in quickly and stably identifying and processing illegal content, shortening the time from information release to handling, and improving the cleanliness of cyberspace and the security of the information environment.
[0046] This invention significantly reduces computational and deployment costs while maintaining high detection accuracy. Compared to solutions relying on large language models with hundreds of billions of parameters, this invention achieves comparable performance using only a language model with a medium parameter scale, thereby greatly reducing computational power consumption, hardware investment, and maintenance costs. By optimizing model size and training strategies, the method of this invention can run efficiently in small to medium-sized servers or cloud environments, ensuring inference speed and stability. Furthermore, this invention provides a reproducible training and inference process, facilitating rapid replication and expansion to different platforms.
[0047] The specific embodiments described above illustrate the technical solution and beneficial effects of the present invention in detail. It should be understood that the above description is only the most preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, additions, and equivalent substitutions made within the scope of the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A content security reasoning and auditing method based on knowledge enhancement, characterized in that, include: A standardized knowledge rule base is constructed by content review experts based on internet content security standards; An attribute system is established based on the violations in the standardized knowledge rule base, and then the attribute system is structured into an attribute set; A prompt template generation function is designed based on attribute sets to generate prompt text. The prompt text is used to call the teacher model to infer and generate candidate answers. A synthetic dataset is built using candidate answers and a standardized knowledge rule base. Synthetic datasets are used to guide student models in supervised fine-tuning to obtain content moderation category judgments and explanatory texts. Then, reinforcement learning is used to update the category judgments and explanatory texts generated by the student models to train and optimize the student models. The trained and optimized student model is deployed to the Chinese Internet content review scenario. Taking any content as input, it generates reasoning explanations for manual review and accountability, and finally outputs violation category tags to complete the content security reasoning review.
2. The content security reasoning and auditing method based on knowledge enhancement according to claim 1, characterized in that, The standardized knowledge rule base includes content categories that are both illegal and non-illegal. For each content category, content review experts, based on domain knowledge and combined with internet content security standards, mark the corresponding rules and judgment points to determine the identification standards and triggering conditions for illegal content.
3. The content security reasoning and auditing method based on knowledge enhancement according to claim 2, characterized in that, The attribute system includes user profile features, text features, avoidance strategy features, and knowledge rule features; The user profile features include gender, age, occupation, and education level, which are used to simulate different writing styles and contexts; The text features include text length, narrative perspective, and publishing platform, which are used to reflect differences in content format; The avoidance strategy features avoidance detection methods that include emoticons, homophones, and variant symbols, used to reproduce hidden expressions; The knowledge rule feature is a standardized knowledge rule base that is constructed to indicate the specific norms violated by the content text.
4. The content security reasoning and auditing method based on knowledge enhancement according to claim 1, characterized in that, The design process of the prompt template generation function includes: setting the prompt of the prompt template generation function, and providing task options and generation requirements based on the attribute set to guide the generation of prompt text; The prompt is used to clearly indicate that the role of the prompt template generation function is a content review expert. The task options based on the attribute set include: user profile task options including gender, age, occupation, and education level; text feature task options including whether it violates regulations, violation category, rule violation, content text length, narrative perspective, and publishing platform; and avoidance strategy task options including avoidance methods and corresponding explanations. The generation requirements are specifically: to generate content text that conforms to the user profile task options and text feature task options, and to correctly apply the avoidance strategy task options when using them.
5. The content security reasoning and auditing method based on knowledge enhancement according to claim 1, characterized in that, The method of using synthetic datasets to guide supervised fine-tuning of student models to obtain content moderation category judgments and explanatory texts includes: Taking the text of the infringing content and the standardized knowledge rule base as input, the reasoning text and category labels generated by the teacher model are used as output to guide the student model to learn the joint distribution of candidate answers and the standardized knowledge rule base. The candidate answers contain implicit knowledge, and the standardized knowledge rule base contains explicit knowledge. A text sequence loss is established to supervise and fine-tune the student model, enabling the student model to have a preliminary semantic understanding and rule alignment ability, and output the category judgment and explanation text of content review.
6. The content security reasoning and auditing method based on knowledge enhancement according to claim 5, characterized in that, The text sequence loss is represented by the cross-entropy loss function, calculated as follows: , , In the formula, For text sequence loss, To infer the length of the text sequence, For the length of the category text sequence, This indicates the reasoning and interpretation of the text in the first part. Each word element, The first in the category text Each word element, Indicates that the attribute category belongs to the synthetic dataset. The Input of each sample, This represents the operation of the parameter value to obtain the minimum value of the function. The optimal student model parameter values are those that minimize the loss function. The number of attribute categories, For attribute categories The number of samples in This represents the parameters of the student model.
7. The content security reasoning and auditing method based on knowledge enhancement according to claim 6, characterized in that, Group-based relative policy optimization is adopted as a reinforcement learning optimization strategy to optimize student model parameters, specifically including: By setting a reward function Used to determine the student model's response to the input. The generated inference explanation text and category determination constitute a complete sequence. The accuracy and format of the judgment results are verified, and the final reward is output. ; For each input Sampling from the old strategy Candidate output sequences And calculate its corresponding reward group. If all rewards in this group If all are the same, discard the data in that group and repeat the sampling process until a reward group with a valid gradient signal is obtained. Based on the reward group with the effective gradient signal The objective function is defined as follows by optimizing based on the group relative strategy: , in, Indicates sample input, This indicates the corresponding answer. To follow the old strategy obtained from sampling One of the candidate output sequences; For the student model parameters of the old strategy; probability ratio For the current strategy Generate sequence The Middle The ratio of the probability of a word to the probability of that word being generated by the old strategy; advantage function estimation. For sequence Rewards Subtract groups The mean of the rewards for that group is then divided by the standard deviation of the rewards for that group. The lower bound clipping factor for decoupling. The upper limit pruning coefficient is used to decouple the policy and limit the step size of policy updates.
8. The content security reasoning and auditing method based on knowledge enhancement according to claim 1, characterized in that, After deploying the trained and optimized student model to a Chinese internet content moderation scenario, it uses any content as input and employs a greedy decoding strategy for inference. The inference results are all words corresponding to the current maximum probability, and the corresponding violation category label is output.
9. A knowledge-enhanced content security reasoning and auditing device, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that, When the one or more processors execute the executable code, they are used to implement the knowledge-enhanced content security reasoning auditing method according to any one of claims 1-8.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is used with a computer, it is used to implement the knowledge-enhanced content security reasoning and auditing method according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-mode reinforcement learning driven PM2.5 chemical component vertical profile inversion system and method
CN120123697A
Text-to-SQL (Structured Query Language) generation method and system based on large language model fine tuning
CN120144614A
Reinforcement learning method for improving mathematical ability of large language model and related device
CN120832930A
Cited By
DRG risk identification method and device based on collaborative learning and medium
CN121938542A