Training system of large language model

By employing a phased training and evaluation update method, the problem of rule knowledge and case samples sharing the same gradient update target in the training of large language models is solved, which improves the model's ability to recognize implicit and evolutionary content and achieves a more stable and structured semantic representation of rules.

CN122153439APending Publication Date: 2026-06-05BEIJING QIYI CENTURY SCI & TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING QIYI CENTURY SCI & TECH CO LTD
Filing Date
2026-02-10
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Because rule knowledge and case samples share the same gradient update objective during the training of a large language model, it is difficult for the large language model to form a stable and structured rule semantic representation in the parameters, resulting in insufficient ability of the trained large language model to recognize implicit and evolutionary content.

Method used

A phased training method is adopted. First, the initial large language model is trained based on the rule training set to generate a rule-based large language model. Then, the rule-based large language model is further trained based on the case training set to generate a target large language model. Finally, the training set is updated through the evaluation module to improve the model's recognition ability.

Benefits of technology

The trained large language model significantly improves the ability to recognize implicit and evolutionary content, solves the inherent competition problem in existing technologies, and achieves a more stable and structured rule semantic representation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122153439A_ABST
    Figure CN122153439A_ABST
Patent Text Reader

Abstract

The application provides a large language model training system. The large language model training system comprises a knowledge management module, a training data processing module, a first stage training module, and a second stage training module. The knowledge management module is configured to obtain an audit rule and generate a rule training set based on the audit rule. The training data processing module is configured to obtain a case data set and construct a case training set based on the case data set. The first stage training module is configured to train an initial large language model based on the rule training set to generate a rule large language model. The second stage training module is configured to train the rule large language model based on the case training set to obtain a target large language model. The technical problem that a large language model is difficult to form stable and structured rule semantic representation in parameters in the prior art, and the recognition ability of the trained large language model for implicit and evolving content is insufficient can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a training system for a large language model. Background Technology

[0002] With the in-depth application of artificial intelligence technology, using large language models for automated content security review has become a key trend for internet platforms to improve governance efficiency.

[0003] In related technologies, existing large language models are mainly trained through knowledge augmentation, introducing rule-based knowledge during the training process. This involves mixing rule-based knowledge as structured knowledge with manually labeled case samples for single-stage end-to-end training of the large language model. However, because rule-based knowledge and case samples share the same gradient update objective during the training of the large language model, they inherently compete at the semantic representation and task fitting levels. This makes it difficult for the large language model to form stable, structured rule-based semantic representations in its parameters, resulting in insufficient ability to recognize implicit and evolutionary content. Summary of the Invention

[0004] This application provides a training system for a large language model to address the technical problem that, since rule knowledge and case samples share the same gradient update target during the training process of a large language model, there is an inherent competition between the two in terms of semantic representation and task fitting, which makes it difficult for the large language model to form a stable and structured rule semantic representation in the parameters, resulting in insufficient recognition ability of the trained large language model for implicit and evolutionary content.

[0005] In a first aspect of the embodiments of this application, a training system for a large language model is provided, the training system for the large language model comprising: a knowledge management module, a training data processing module, a first-stage training module, and a second-stage training module; The knowledge management module is used to acquire the review rules and generate a rule training set based on the review rules. The training data processing module is used to acquire the case dataset and construct a case training set based on the case dataset. The first-stage training module is used to train the initial large language model based on the rule training set to generate a rule-based large language model. The second-stage training module is used to train the rule-based large language model based on the case training set to obtain the target large language model.

[0006] In an optional implementation, the training system for the large language model further includes: an evaluation module; The evaluation module is used to evaluate the performance of the target large language model and obtain the evaluation results.

[0007] In an optional implementation, the evaluation module is further configured to update the rule training set and / or the case training set based on the evaluation results.

[0008] In one optional implementation, the knowledge management module is specifically used for: The audit rules are preprocessed to generate initial audit rules; The initial review rules are structurally transformed to generate the rule training set.

[0009] In an optional implementation, the training system for the large language model further includes a control module; The control module is used to control the first-stage training module to retrain or fine-tune the target large language model according to the updated rule training set after the rule training set and / or the case training set are updated based on the evaluation results. And / or, The second-stage training module is controlled to retrain or fine-tune the target large language model based on the updated case training set.

[0010] In an optional implementation, updating the rule training set and / or the case training set based on the evaluation results includes: Obtain feedback on audit deviations from the assessment results; The type of deviation is determined based on the aforementioned audit deviation feedback; The rule training set and / or the case training set are updated based on the deviation type.

[0011] In an optional implementation, updating the rule training set and / or the case training set based on the deviation type includes: Determine whether the deviation type is a regular deviation; If the deviation type is a rule deviation, then the rule training set is updated; If the deviation type is not the rule deviation, then the case training set is updated.

[0012] In an optional implementation, updating the rule training set includes: The feedback on the audit deviations is analyzed to determine the revision rules; Add the revised rule to the rule training set.

[0013] In an optional implementation, updating the case training set includes: The audit deviation feedback is analyzed to obtain the original deviation cases; The case samples in the case training set that are related to the original biased cases are corrected.

[0014] In an optional implementation, the evaluation module is further configured to: Obtain associated cases related to the original deviation case; Based on the aforementioned related cases, a set of related training samples is determined; Add the associated training sample set to the case training set.

[0015] The large language model training system provided in this application includes: a knowledge management module, a training data processing module, a first-stage training module, and a second-stage training module. The knowledge management module is used to acquire review rules and generate a rule training set based on these rules. The training data processing module is used to acquire a case dataset and construct a case training set based on it. The first-stage training module is used to train the initial large language model based on the rule training set to generate a rule-based large language model. The second-stage training module is used to train the rule-based large language model based on the case training set to obtain the target large language model. This phased training of the large language model based on the rule training set and the case training set can train a large language model with strong recognition capabilities for implicit and evolutionary content. This solves the technical problem in existing technologies where rule knowledge and case samples share the same gradient update target during the training of the large language model, leading to inherent competition between them at the semantic representation and task fitting levels. This makes it difficult for the large language model to form stable and structured rule semantic representations in the parameters, resulting in insufficient recognition capabilities for implicit and evolutionary content in the trained large language model. Attached Figure Description

[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A schematic diagram of the structure of a training system for a large language model provided in an embodiment of this application; Figure 2A schematic diagram of the structure of another large language model training system provided in this application embodiment; Figure 3 A schematic diagram illustrating the implementation process of a training set update method provided in this application embodiment; Figure 4 A schematic diagram illustrating the implementation process of another training set update method provided in this application embodiment; Figure 5 This is a schematic diagram illustrating the implementation process of a training method for a large language model provided in an embodiment of this application. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0021] To address the technical problem in existing technologies where rule knowledge and case samples share the same gradient update target during the training of a large language model, leading to inherent competition between them at the semantic representation and task fitting levels, making it difficult for the large language model to form stable and structured rule semantic representations in its parameters, and thus resulting in insufficient recognition ability of the trained large language model for implicit and evolutionary content, this application provides a training system for a large language model, comprising: a knowledge management module, a training data processing module, a first-stage training module, and a second-stage training module; the knowledge management module is used to acquire review rules and generate a rule training set based on the review rules; the training data processing module is used to acquire a case dataset and construct a case training set based on the case dataset; the first-stage training module is used to train an initial large language model based on the rule training set to generate a rule-based large language model; the second-stage training module is used to train the rule-based large language model based on the case training set to obtain a target large language model. This phased training of the large language model based on the rule training set and the case training set can produce a large language model with strong recognition ability for implicit and evolutionary content.

[0022] See Figure 1 This is a schematic diagram of the structure of a training system for a large language model provided in an embodiment of this application. Figure 1 As shown, the training system 10 for the large language model may include a knowledge management module 11, a training data processing module 12, a first-stage training module 13, and a second-stage training module 14.

[0023] The knowledge management module 11 is used to acquire audit rules and generate a rule training set based on those rules. Specifically, audit rules can be obtained from the company's internal audit rule library or related rule documents. Audit rules refer to a series of text rules and clauses defined in the company's internal audit rule library or related rule documents to determine whether text content is compliant, such as prohibiting personal attacks and insults against others, and prohibiting the dissemination of unverified information.

[0024] The training data processing module 12 is used to acquire case datasets and construct case training sets based on these datasets. Specifically, case datasets can be obtained from historical review rule logs and manual annotation platforms. The case dataset refers to the collection of specific text content and its corresponding tags that has been manually or pre-determined as "violation" or "compliance."

[0025] The first-stage training module 13 is used to train the initial large language model based on the rule training set to generate a rule-based large language model. The initial large language model refers to a general-purpose large language model that has not been trained for a content moderation task, such as open-source or self-developed models like Qwen, ChatGLM, and LLaMA. This embodiment of the application does not limit this type of model.

[0026] The second-stage training module 14 is used to train the rule-based large language model based on the case training set to obtain the target large language model.

[0027] In the large language model training system 10 provided in this application embodiment, the knowledge management module 11 is used to obtain review rules and generate a rule training set based on the review rules. Specifically, the review rules can be preprocessed to generate initial review rules, and the initial review rules can be structurally transformed to generate a rule training set.

[0028] Preprocessing involves cleaning and standardizing the original review rule texts from diverse sources and with varying formats, transforming them into clean, uniform, and unambiguous standardized text (i.e., initial review rules). Specifically, rule engines or regular expressions can be used to remove descriptive text, version numbers, revision history, and other metadata irrelevant to the review standards. A synonym and near-synonym mapping table is established to standardize terms with the same meaning but different expressions within the rules. For rules that are too brief or contain implicit premises, necessary supplementary explanations are provided through knowledge base association or manual review.

[0029] Structured transformation refers to converting preprocessed, unstructured natural language rules into formatted training samples (i.e., rule training sets) suitable for large language models to learn rules.

[0030] The process of structurally transforming the initial review rules to generate a rule training set can include: using natural language processing tools or large language model interfaces to parse each preprocessed rule and identify its core elements. These core elements can include the violating subject (i.e., the prohibited content or behavior), constraints (i.e., the degree or scope of the prohibition, such as "prohibited," "must not," etc.), scenario description (the applicable context, such as "in the comments section"), and examples or feature descriptions (such as "including but not limited to the use of insulting words, malicious denigration"). Based on the parsed core elements, diverse question-and-answer pairs are generated and organized according to a specified fine-tuned data format (such as instruction-input-output JSONL format) to form the final rule training set. The rule training set can include rule questions and answers, violation type explanation questions and answers, feature summary and identification questions and answers, and compliance judgment and reasoning questions and answers.

[0031] For example, the rule Q&A states: "Q: According to the content security guidelines, under what circumstances does 'personal attack' constitute? A: According to the guidelines, using insulting or abusive language to maliciously belittle or attack another person's personality or dignity constitutes a personal attack."; the violation type explanation Q&A states: "Q: What specific characteristics does 'non-compliant content' refer to? A: Non-compliant content refers to content that uses vulgar language to gain attention." Feature summary and identification states: "Q: What keywords or patterns should be paid attention to when identifying 'leading' content? A: Be wary of enticing language such as promises of 'high returns' or 'scan the code to join the group and receive materials,' as well as unqualified recommendations for returns."; the compliance judgment reasoning Q&A states: "Q: Is a user comment 'You are extremely stupid' a violation? Why? A: Yes, it is a violation. This comment uses the insulting term 'extremely stupid,' maliciously belittling another person's intelligence, which meets the rule definition of 'personal attack.'"

[0032] In the training system 10 of the large language model provided by the embodiments of this application, the training data processing module 12 is used to obtain a case data set and construct a case training set based on the case data set.

[0033] For constructing a case training set based on the case data set, specifically, the following steps may be included: 1. Data cleaning and denoising, that is, format standardization, text normalization, label verification and correction. For format standardization, it means to统一文本编码(如UTF-8),去除无关字符、乱码、超长无意义字符串等。For text normalization, it means to perform a certain degree of normalization on Internet terms, pinyin, homophony, emoticons, etc. (for example, convert "NV" to "female"). For label verification and correction, it means to use a rule engine or a high-precision small model to perform consistency checking and conflict detection on the original labels. For samples in the historical logs where the system's automatic judgment and the manual review results are inconsistent, the manual review results shall be taken as the final standard; for obviously incorrect annotations (such as marking obviously illegal content as compliant), filter them or submit them for re-annotation.

[0034] 2. Sample screening and balancing, that is, quality filtering and class balancing. For quality filtering, it means to remove samples with too short text length (such as less than 2 valid characters), meaningless content or too high repetition rate. For class balancing, it means to address the common class imbalance problem in content review (such as far more compliant samples than illegal samples, or few samples in certain illegal categories), and use methods such as strategic sampling (such as oversampling the minority class, undersampling the majority class) or synthesizing minority class samples (care should be taken to ensure no bias is introduced) to construct a training set with a relatively balanced class distribution, so as to improve the overall recognition ability of the model for various illegal contents and avoid bias towards the majority class.

[0035] 3. Data desensitization and privacy protection: Identify and mask personal sensitive information (such as mobile phone numbers, ID numbers, bank card numbers, specific addresses, etc.) in the text (such as replace with generalization marks such as [PHONE], [ID], etc.), ensuring that the training process meets the requirements of data security and privacy protection.

[0036] 4. Construct the supervised fine-tuning format: Convert the (text, label) pairs after cleaning and screening into the format required for the large language model to perform instruction tuning or classification tuning.

[0037] For example, construct a question-and-answer instruction pair: The input (Instruction / Input) is: "Please determine whether the following text content is illegal and give the type of violation: [text to be reviewed]". The output (Output) is the standard answer that the model should learn, such as: "This text is illegal and belongs to the category of 'abusive personal attack'." or "This text is compliant." It should be noted that there are some inaccuracies in the Chinese text you provided. For example, "统一文本编码(如UTF-8),去除无关字符、乱码、超长无意义字符串等。" is an incomplete expression in Chinese. I have tried my best to translate it according to the context. If you have any further questions, please feel free to let me know.For example, a classification task format: take text as input and require the model to output the corresponding category label or judgment result.

[0038] 5. Dataset partitioning: The final processed case training set is divided into training set, validation set and test set according to a preset ratio (e.g., 8:1:1) for training, parameter tuning and final performance evaluation of the large language model.

[0039] A large language model training system 10 provided in this application embodiment includes: a knowledge management module 11, a training data processing module 12, a first-stage training module 13, and a second-stage training module 14. The knowledge management module 11 is used to acquire review rules and generate a rule training set based on these rules; the training data processing module 12 is used to acquire a case dataset and construct a case training set based on the case dataset; the first-stage training module 13 is used to train an initial large language model based on the rule training set to generate a rule-based large language model; and the second-stage training module 14 is used to train the rule-based large language model based on the case training set to obtain a target large language model. By training the large language model in stages based on rule training sets and case training sets, a large language model with strong recognition ability for implicit and evolutionary content can be trained. This solves the technical problem in existing technologies where rule knowledge and case samples share the same gradient update target during the training of the large language model. As a result, the two have an inherent competition in terms of semantic representation and task fitting, making it difficult for the large language model to form a stable and structured rule semantic representation in the parameters. Consequently, the trained large language model has insufficient recognition ability for implicit and evolutionary content.

[0040] See Figure 2 This is a schematic diagram of the structure of another large language model training system provided in an embodiment of this application. Figure 2 As shown, the large language model training system 10 provided in this application embodiment further includes: an evaluation module 15.

[0041] The evaluation module 15 in the training system of the large language model provided in this application embodiment can be used to evaluate the performance of the target large language model, obtain evaluation results, and update the rule training set and / or case training set based on the evaluation results. Performance evaluation refers to the process of comprehensively measuring the effectiveness of the target large language model in content moderation tasks according to preset quantitative indicators and test datasets. Evaluation results refer to the structured report output after performance evaluation, which may include the target large language model's scores on various indicators, the set of sample cases with identification errors (Bad Cases), and specific cases and preliminary analysis that deviate from human review standards.

[0042] For details on how to update the rule training set and / or case training set based on the evaluation results, please refer to... Figure 3 The method shown. (As illustrated) Figure 3 The diagram shown illustrates the implementation flow of a training set update method provided in this application embodiment, which may specifically include the following steps: S301, Obtain audit deviation feedback from assessment results.

[0043] In this embodiment, review deviation feedback can be obtained from the evaluation results. This feedback refers to detailed information regarding errors made by the target language model in its review judgment. Specifically, this may include: the original text content that caused the error, the target language model's incorrect judgment (e.g., classifying the incorrect text as compliant, or a classification error), the correct judgment given by the human reviewer, and a preliminary description of the error scenario. For example, the review deviation feedback might be: "Sample ID: 12345, Text: 'Add me for benefits', Target language model judgment: Compliant (advertising category), Human judgment: Illegal (traffic redirection type)."

[0044] S302, Determine the type of deviation based on audit deviation feedback.

[0045] In this embodiment, the deviation type can be determined based on the audit deviation feedback. The deviation type refers to the qualitative classification of the error root cause, used to distinguish whether the problem corresponding to the audit deviation feedback stems from imperfections or omissions in the audit rules themselves, or from insufficient application capabilities of the target large language model to existing rules. Deviation types can be divided into rule deviations and case deviations. Rule deviations indicate that the problem corresponding to the audit deviation feedback stems from imperfections or omissions in the audit rules themselves, while case deviations indicate that the problem corresponding to the audit deviation feedback stems from insufficient application capabilities of the target large language model to existing rules.

[0046] Specifically, the violation category or logic used for "human judgment" in the review deviation feedback can be semantically similar to all review rules currently maintained in Knowledge Management Module 11, and rule correlation analysis can be performed. If no review rule can be found that clearly and directly supports the result of "human judgment," it is initially marked as "rule deviation." For example, if the human judgment is "malicious hype," but the rule base only contains the clause "spreading false information," there is a semantic gap between the two in terms of malicious intent and boundaries. If one or more highly relevant existing rules can be found, it can be initially marked as "case deviation." Further analysis of the error reasons of the target large language model is needed, such as whether the rule matching failure is caused by complex linguistic phenomena such as homophones and metaphors in the text, or whether the target large language model has failed to learn the application pattern of the rule due to insufficient similar cases in the training data.

[0047] In addition, cases marked as "rule deviation" or difficult to judge can be submitted to the corresponding content security personnel for final adjudication.

[0048] For example, the text content in the audit deviation feedback is: "This acting is really 'gǎn (gān) rén zhì shēn'." The target large language model determines: compliant (ordinary comment). Manual determination: non-compliant (negative ridicule). Determination analysis: There is a clear rule in the existing rule library that "it is prohibited to use irony, homophony, etc. to conduct personal attacks or maliciously belittle". The target large language model fails to correctly recognize that 'gǎn (gān) rén zhì shēn' is a homophonic irony of 'gǎn rén zhì shēn', indicating that the target large language model lacks the ability to understand the negative intention implied by this homophone, which belongs to the problem of the application ability of existing rules. Therefore, it belongs to a case deviation. The solution is to add more negative comment samples containing homophonic irony to the case training set to enhance the training of the target large language model.

[0049] S303, update the rule training set and / or case training set based on the deviation type.

[0050] In the embodiments of the present application, the rule training set and / or case training set can be updated based on the deviation type.

[0051] And how to update the rule training set and / or case training set based on the deviation type can specifically refer to the Figure 4 method shown. As Figure 4 shown, it is a schematic diagram of the implementation process of another training set update method provided by the embodiments of the present application, which can specifically include the following steps: S401, determine whether the deviation type is a rule deviation.

[0052] In the embodiments of the present application, it can be determined whether the deviation type is a rule deviation. Among them, rule deviation refers to the situation where the existing audit rule library fails to cover the non-compliant situations of the current case, or the rule description is vague, ambiguous, or has unclear boundaries, resulting in the target large language model lacking a basis for correct judgment.

[0053] Specifically, the error samples in the audit deviation feedback can be matched with all the existing audit rules in the knowledge management module 11 in terms of semantic similarity. If no rule can be found to clearly support the correct determination given manually, it is preliminarily determined as a rule deviation. Manual review can also be carried out, such as submitting boundary cases that are difficult to judge through automatic matching to content security experts for review. If experts determine that new rules need to be added, refined, or modified to solve such problems, they are classified as rule deviations.

[0054] S402, if the deviation type is a rule deviation, update the rule training set.

[0055] In this embodiment of the application, if the deviation type is a rule deviation, the rule training set is updated.

[0056] Specifically, updating the rule training set may include: analyzing the feedback on audit deviations, determining the revised rules, and adding the revised rules to the rule training set.

[0057] S403 If the deviation type is not a regular deviation, then update the case training set.

[0058] In this embodiment of the application, if the deviation type is not a rule deviation, the case training set is updated.

[0059] Specifically, updating the case training set may include: analyzing the audit deviation feedback to obtain the original deviation cases, and correcting the case samples in the case training set that are related to the original deviation cases.

[0060] In addition, the training system 10 for the large language model may also include a control module 16.

[0061] In the large language model training system 10 provided in this application embodiment, the control module 16 is used to control the first-stage training module 13 to retrain or fine-tune the target large language model according to the updated rule training set after updating the rule training set and / or case training set based on the evaluation results.

[0062] And / or, The second-stage training module 14 controls the retraining or fine-tuning of the target large language model based on the updated case training set.

[0063] For example, when the rule training set is updated due to the addition of new rules, the control module 16 can trigger a complete first-stage incremental training, using the new rule training set to fine-tune the current target large language model (or the most recent rule large language model) to learn the new rules, and then decide whether to perform a second-stage lightweight fine-tuning to adapt to the case judgment under the new rules.

[0064] Once the training set of cases is updated, control module 16 can trigger the second phase of incremental training, using the newly added or corrected case samples to fine-tune the current target large language model to quickly improve its performance on relevant cases without having to relearn all the rules.

[0065] In addition, the evaluation module 15 in the training system 10 of the large language model is also used to: obtain related cases related to the original biased cases; determine the related training sample set based on the related cases; and add the related training sample set to the case training set.

[0066] Specifically, semantic similarity calculation, retrieval of the same non-compliant category, or associative generation techniques based on the target large language model can be used to search for text content similar to the "original deviation case" in intent, expression, and core vocabulary from historical case databases or publicly available internet data. The obtained related cases are then manually or automatically (using a calibrated model) labeled to determine their compliance / non-compliance status and specific category, forming a new batch of training samples. These new samples are added to the case training set, thereby enhancing the model's ability to identify such violation patterns or linguistic phenomena during iterative training, achieving a "learning by analogy" effect, preventing similar deviations from recurring, and improving the generalization and robustness of the target large language model.

[0067] See Figure 5 This is a schematic diagram illustrating the implementation process of a training method for a large language model provided in this application embodiment. The training system 10 applied to the aforementioned large language model may specifically include the following steps: S501, Obtain the audit rules and generate a rule training set based on the audit rules.

[0068] In this embodiment, audit rules can be obtained, and a rule training set can be generated based on these rules. Specifically, audit rules can be obtained from an enterprise's internal audit rule library or related rule documents. Audit rules refer to a series of text rules and clauses defined in an enterprise's internal audit rule library or related rule documents to determine whether text content is compliant, such as prohibiting personal attacks and insults against others, and prohibiting the dissemination of unverified information. The audit rules can be preprocessed to generate initial audit rules, and then structurally transformed to generate a rule training set.

[0069] Preprocessing involves cleaning and standardizing the original review rule texts from diverse sources and with varying formats, transforming them into clean, uniform, and unambiguous standardized text (i.e., initial review rules). Specifically, rule engines or regular expressions can be used to remove descriptive text, version numbers, revision history, and other metadata irrelevant to the review standards. A synonym and near-synonym mapping table is established to standardize terms with the same meaning but different expressions within the rules. For rules that are too brief or contain implicit premises, necessary supplementary explanations are provided through knowledge base association or manual review.

[0070] Structured transformation refers to converting preprocessed, unstructured natural language rules into formatted training samples (i.e., rule training sets) suitable for large language models to learn rules.

[0071] The process of structurally transforming the initial review rules to generate a rule training set can include: using natural language processing tools or large language model interfaces to parse each preprocessed rule and identify its core elements. These core elements can include the violating subject (i.e., the prohibited content or behavior), constraints (i.e., the degree or scope of the prohibition, such as "prohibited," "must not," etc.), scenario description (the applicable context, such as "in the comments section"), and examples or feature descriptions (such as "including but not limited to the use of insulting words, malicious denigration"). Based on the parsed core elements, diverse question-and-answer pairs are generated and organized according to a specified fine-tuned data format (such as instruction-input-output JSONL format) to form the final rule training set. The rule training set can include rule questions and answers, violation type explanation questions and answers, feature summary and identification questions and answers, and compliance judgment and reasoning questions and answers.

[0072] For example, the rule Q&A states: "Q: According to the content security guidelines, under what circumstances does 'personal attack' constitute? A: According to the guidelines, using insulting or abusive language to maliciously belittle or attack another person's personality or dignity constitutes a personal attack."; the violation type explanation Q&A states: "Q: What specific characteristics does 'non-compliant content' refer to? A: Non-compliant content refers to content that uses vulgar language to gain attention." Feature summary and identification states: "Q: What keywords or patterns should be paid attention to when identifying 'leading' content? A: Be wary of enticing language such as promises of 'high returns' or 'scan the code to join the group and receive materials,' as well as unqualified recommendations for returns."; the compliance judgment reasoning Q&A states: "Q: Is a user comment 'You are extremely stupid' a violation? Why? A: Yes, it is a violation. This comment uses the insulting term 'extremely stupid,' maliciously belittling another person's intelligence, which meets the rule definition of 'personal attack.'"

[0073] S502, Obtain the case dataset and construct a case training set based on the case dataset.

[0074] In this embodiment of the application, a case dataset can be obtained, and a case training set can be constructed based on the case dataset. Specifically, the case dataset can be obtained from historical review rule logs or manual annotation platforms. The case dataset refers to the collection of specific text content and its corresponding tags that have been manually or pre-determined as "violation" or "compliance".

[0075] S503, the initial large language model is trained based on the rule training set to generate a rule-based large language model.

[0076] In this embodiment, the initial large language model can be trained based on the rule training set to generate a rule-based large language model. The initial large language model refers to a general-purpose large language model that has not been trained for a content moderation task, such as open-source or self-developed models like Qwen, ChatGLM, and LLaMA; this embodiment does not limit this type of model.

[0077] S504, The rule-based large language model is trained based on the case training set to obtain the target large language model.

[0078] In this embodiment of the application, the rule-based large language model can be trained based on the case training set to obtain the target large language model.

[0079] In the large language model training method provided in this application embodiment, the following steps are taken: First, review rules are obtained, and a rule training set is generated based on these rules. Second, a case dataset is obtained, and a case training set is constructed based on the case dataset. Third, the initial large language model is trained using the rule training set to generate a rule-based large language model. Finally, the rule-based large language model is trained using the case training set to obtain the target large language model. This phased training of the large language model based on the rule training set and the case training set can produce a large language model with strong recognition capabilities for implicit and evolutionary content. This addresses the technical problem in existing technologies where rule knowledge and case samples share the same gradient update target during the training process of the large language model. The inherent competition between the two at the semantic representation and task fitting levels makes it difficult for the large language model to form stable and structured rule semantic representations in the parameters, resulting in insufficient recognition capabilities for implicit and evolutionary content.

[0080] In another embodiment provided in this application, a computer storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform the training method of any of the large language models described in the above embodiments.

[0081] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute the training method of any of the large language models described in the above embodiments.

[0082] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0083] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0084] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0085] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A training system for a large language model, characterized in that, The training system for the large language model includes: a knowledge management module, a training data processing module, a first-stage training module, and a second-stage training module; The knowledge management module is used to acquire the review rules and generate a rule training set based on the review rules. The training data processing module is used to acquire the case dataset and construct a case training set based on the case dataset. The first-stage training module is used to train the initial large language model based on the rule training set to generate a rule-based large language model. The second-stage training module is used to train the rule-based large language model based on the case training set to obtain the target large language model.

2. The system according to claim 1, characterized in that, The training system for the large language model also includes: an evaluation module; The evaluation module is used to evaluate the performance of the target large language model and obtain the evaluation results.

3. The system according to claim 2, characterized in that, The evaluation module is also used to update the rule training set and / or the case training set based on the evaluation results.

4. The system according to claim 1, characterized in that, The knowledge management module is specifically used for: The audit rules are preprocessed to generate initial audit rules; The initial review rules are structurally transformed to generate the rule training set.

5. The system according to claim 3, characterized in that, The training system for the large language model also includes a control module; The control module is used to control the first-stage training module to retrain or fine-tune the target large language model according to the updated rule training set after the rule training set and / or the case training set are updated based on the evaluation results. And / or, The second-stage training module is controlled to retrain or fine-tune the target large language model based on the updated case training set.

6. The system according to claim 3, characterized in that, The step of updating the rule training set and / or the case training set based on the evaluation results includes: Obtain feedback on audit deviations from the assessment results; The type of deviation is determined based on the aforementioned audit deviation feedback; The rule training set and / or the case training set are updated based on the deviation type.

7. The system according to claim 6, characterized in that, The step of updating the rule training set and / or the case training set based on the deviation type includes: Determine whether the deviation type is a regular deviation; If the deviation type is a rule deviation, then the rule training set is updated; If the deviation type is not the rule deviation, then the case training set is updated.

8. The system according to claim 7, characterized in that, The update of the rule training set includes: The feedback on the audit deviations is analyzed to determine the revision rules; Add the revised rule to the rule training set.

9. The system according to claim 7, characterized in that, The update of the case training set includes: The audit deviation feedback is analyzed to obtain the original deviation cases; The case samples in the case training set that are related to the original biased cases are corrected.

10. The system according to claim 9, characterized in that, The evaluation module is also used for: Obtain associated cases related to the original deviation case; Based on the aforementioned related cases, a set of related training samples is determined; Add the associated training sample set to the case training set.