Large model value alignment reinforcement method in educational scene

By constructing educational scenario information, a dynamic risk rule dictionary and an evaluation mechanism that simulates human metacognitive ability in the big model, and combining reinforcement learning to internalize the metacognitive thinking chain, the problem of implicit risk identification in educational scenarios of big models is solved, and the improvement of safety and positive values ​​is achieved. It is applicable to a variety of mainstream big models.

CN120671819APending Publication Date: 2025-09-19EAST CHINA NORMAL UNIV

Patent Information

Application Number
CN202510734757.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing large models are difficult to effectively identify and filter hidden risks in educational scenarios, especially for the bad behavioral tendencies and sensitive issues of young users, which may lead to misleading and health hazards. Existing technologies such as supervised fine-tuning, human feedback reinforcement learning and underlying constitutions have limitations in this regard.

Method used

The framework of value alignment reinforcement of large educational models is adopted. By adding scenario information, building a dynamic risk rule dictionary, simulating the automatic risk assessment of human metacognitive ability, and an iterative reinforcement strategy based on the metacognitive thinking chain, including a static multi-level rule tree, a dynamic patch rule pool, and an evaluation mechanism that simulates human metacognitive ability, combined with reinforcement learning to internalize the metacognitive thinking chain, the safety of the model's answers and the positivity of its values ​​are improved.

Benefits of technology

It significantly reduces the risk answer rate of large models in educational scenarios, improves the security and value positivity of the model, has good scalability and sustainability, does not require additional training costs, and ensures the healthy growth of young users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671819A_ABST
    Figure CN120671819A_ABST
Patent Text Reader

Abstract

The invention discloses a large model value alignment reinforcement method in an educational scene, which is characterized in that information such as user learning segments, time and space, function roles and the like is used as scene information to be added, a dynamic risk rule dictionary and an evaluation mechanism are adopted to carry out risk identification and correction, and finally a thinking chain is internalized into model deep thinking. The method specifically comprises the steps of constructing a scene layer embedded education scene information, constructing a dynamic risk rule dictionary, constructing a value risk automatic evaluation mechanism, constructing an iterative reinforcement mechanism, and reinforcing a model through reinforcement learning internalization meta-cognitive thinking chain. Compared with the prior art, the method has good expandability and sustainability, effectively reduces the risk response rate of the large model in an education scene, improves the safety and the value forward performance of the model, does not need extra training cost, can be directly applied to various mainstream large models, and has a wide application prospect. And a powerful safety guarantee is provided for large model application in the education field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large-scale model security value alignment, in particular to a large-scale model value alignment reinforcement method for educational scenes. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, the application of big models in education has become increasingly widespread, bringing profound changes to educational methods. Big models can provide personalized learning guidance, intelligent Q&A, and creative inspiration, greatly enriching educational methods and resources. However, the unique characteristics of the adolescent user population also present many unique and severe challenges for big models in educational settings. Adolescence is a critical period for the formation of a person's values ​​and perspectives. During this period, users' minds are not yet fully mature, and their ability to discern information and self-control is relatively weak, making them more susceptible to negative information.

[0003] While mainstream big data models have to some extent addressed explicit risks such as violence and illegal activities, a significant number of hidden risks still exist in educational settings. These risks are often more subtle and long-term, making them difficult for existing technologies to effectively identify and filter. For example, some teenagers may be influenced by negative aesthetic values ​​and imitate the self-harm behaviors depicted in films and television, believing facial scarring is fashionable. A junior high school student once asked a big data model for advice on how to alleviate the pain of a self-injured wound, but the model directly provided information on medications, failing to identify the user's hidden self-injury risk. In such cases, the big data model's responses may inadvertently reinforce the teenager's tendency toward negative behavior, causing serious harm to their physical and mental health. For another example, big data models sometimes provide inappropriate responses to questions about illegal or inappropriate behaviors such as smoking and drinking. When asked by minors, big data models fail to fully consider the user's identity and the legality of their behavior, failing to provide proper guidance. Such responses not only violate the principles of value guidance in educational settings but also may mislead teenagers and encourage negative behavior. Furthermore, there are risks involving sensitive issues such as national sovereignty and political security. Some users may intentionally or unintentionally misrepresent or misunderstand national territories, political systems, and other aspects of their questions. For example, when a user asked about dorm decorating and mentioned the use of maps, the large-scale model failed to identify the potential national territorial security risks involved and provided incorrect advice. These types of questions are highly insidious and require the large-scale model to possess a high degree of political sensitivity and risk identification capabilities.

[0004] In summary, existing large-model value reinforcement methods, such as supervised fine-tuning (SFT), reinforcement learning with human feedback (RLHF), and underlying constitutions, have limitations when addressing hidden risks in educational settings. Supervised fine-tuning relies primarily on large amounts of labeled data for model training, but covering all hidden risk scenarios in educational settings requires significant human and material resources. While reinforcement learning with human feedback can optimize models based on human feedback, the subjectivity and uncertainty of these feedback signals can lead to unstable model performance. Underlying constitutions, on the other hand, constrain model outputs by defining fundamental principles, struggle to adapt to the diverse and complex demands of educational settings. Therefore, to address the hidden risk issues of large models in educational settings, a more effective, adaptable, and training-free value alignment reinforcement framework is urgently needed to ensure the safe application of large models in educational settings and promote the healthy development of adolescents. Summary of the Invention

[0005] The present invention addresses the shortcomings of existing technologies by providing a method for strengthening the alignment of values ​​in large models for educational scenarios. This method utilizes a framework for strengthening the alignment of values ​​in educational large models. By adding contextual information, constructing a dynamic risk rule dictionary, automatically assessing risks by simulating human metacognitive abilities, reinforcing model risk responses based on metacognitive thought chains, and internalizing these chains through reinforcement learning, the method aims to improve the security and value positivity of large models' responses in educational settings. This framework applies preliminary constraints to model inputs by adding information such as the user's academic stage, time, space, and functional role at the scenario level. It then pre-constrains the inputs using a dynamic risk rule dictionary, performs risk assessment on the model output using an automatic assessment mechanism that simulates human metacognitive abilities, and corrects risk responses using a reinforcement strategy based on metacognitive thought chains. Metacognitive assessment results are converted into patch rules and stored in a dynamic patch rule pool, providing constraints for subsequent similar questions. Experiments on existing mainstream large models have demonstrated that this method can effectively reduce the risk response rate of large models in educational settings, improve model security and value positivity, and eliminate the need for additional training costs, demonstrating good scalability and sustainability.

[0006] The specific technical solution to achieve the purpose of this invention is: a method for aligning and reinforcing values ​​in a large model in an educational scenario. Its characteristics are to attach information such as the user's learning stage, time and space, and functional role as scenario information, and to use a dynamic risk rule dictionary and an evaluation mechanism that simulates human metacognitive ability to identify and correct risks. Finally, the thought chain is internalized into the deep thinking of the model to obtain an efficient safety model. This method includes:

[0007] Step 1: Build a scene layer to embed educational scene information

[0008] In educational scenarios, models cannot effectively determine the questioner's intent when answering questions. However, contextual information can help the model discern the questioner's intent, thereby increasing the likelihood of providing a safe answer. Before the user enters a question, additional contextual information is provided, including user account information, large model functional role, and the user's location and time. This additional contextual information can initially constrain the model's input and provide basic value guidance.

[0009] Step 2: Build a dynamic risk rule dictionary with autonomous development capabilities

[0010] The dynamic risk rule dictionary constructed in this step is mainly divided into a static multi-level rule tree and a dynamic patch rule pool, as well as a search header for retrieving rules, as follows:

[0011] 2-1: Static multi-level rule tree

[0012] It includes a macro rule area and a response style rule area. The macro rule area is composed of the country, society, and individuals as the first-level nodes, and splits out into second- and third-level sub-nodes. Each node is mounted with macro rules corresponding to major or minor categories; the response style rule area is divided into three colors: red line, yellow line, and green line according to the different risk levels of the problem. The response strategies corresponding to different risk levels are: red line problems need to be clearly prohibited, yellow line problems need to correct deviations, and green line problems need to be actively guided.

[0013] 2-2: Dynamic patch rule pool

[0014] Store patch rules converted from metacognitive evaluation results in the form of label-rule pairs exists in the form of: It is a patch tag, which is used to uniquely identify the patch rule; It is the patch rule content that defines specific constraints or adjustment strategies.

[0015] 2-3: Search Header

[0016] Responsible for retrieving the rules applicable to the current input from the static multi-level rule tree and the dynamic patch rule pool. The retrieval head uses the embedded matching mechanism to retrieve from the rule base and is trained by the prototype contrast learning method. , the search head first performs pruning matching as follows:

[0017] ;

[0018] in, represents the candidate rule set, Represents the static rule tree A subtree. is the coarse-grained matching function, is the pruning threshold.

[0019] The rule tree greatly shortens the time to retrieve the corresponding rule, and then the semantic similarity between it and all the rules on the subdivided branches is calculated by the following formula:

[0020] ;

[0021] .

[0022] in, is the embedding vector of the problem, For rules The embedding vector of is the cosine similarity function, and finally returns High scoring rules.

[0023] 2-4: Sustainable Development Mechanism

[0024] After a period of time, when a certain number of rules appear in the dynamic patch rule pool, certain patch rules are summarized and selected through human-machine interaction and mounted on the static multi-level rule tree, which is specifically expressed as follows:

[0025] .

[0026] in, It is a rule obtained by selecting patch rules from the dynamic patch rule pool using a large model or human experts.

[0027] Step 3: Build an automatic value risk assessment mechanism that simulates human metacognitive ability

[0028] The metacognitive evaluator designed in this step adopts the mechanism expressed as follows:

[0029] .

[0030] in, A large model for metacognitive assessment, is the model answer for round i, This is the evaluation report of the metacognitive evaluator. The risk assessment result is an integer between 1 and 5, where 1 represents a very high risk and 5 represents no risk, and the risk decreases gradually from small to large. For the Wheel of Answers;

[0031] The metacognitive evaluator Select a subset of thinking tools from the metacognitive toolbox for assessment. These tools mainly include the following options:

[0032] 3-1: Semantic deconstruction based on context

[0033] Strip away the surface semantics and combine the implicit context of the answer to determine whether the answer in that context contains implicit malicious positions, prejudices, etc.

[0034] 3-2: Questioning Assumptions

[0035] Through questioning thinking, think about what consequences will occur if the user takes the next step according to the behavior, words and other prompts in the answer.

[0036] 3-3: Replacement Verification

[0037] Determine whether the answers can remain consistent if the roles and scenarios change.

[0038] 3-4: Rule Introspection

[0039] Reflect on whether the answer is contrary to the mainstream values ​​of society, and whether it violates morality, laws, mainstream regulations, systems, etc.

[0040] 3-5: Origin of Values

[0041] Trace the values ​​and ideologies revealed in the answers and determine whether they conflict with or violate positive educational values.

[0042] The metacognitive evaluator is used in the selection tool When , the following formula is used to update the dynamic weight according to different scenarios :

[0043] .

[0044] in, Display Tools The utility ratio, is the learning rate. In different scenarios, the utility ratios of different tools vary depending on the actual situation of the scenario, so adopting this solution will make the evaluator more flexible.

[0045] Step 4: Build an iterative reinforcement mechanism through metacognitive thinking chains

[0046] This step adopts an iterative improvement strategy: If (Risky), the system will force the model to use the answers from the previous round , Metacognitive Assessment Report And the user's original question , generating an improved answer , and again conduct metacognitive assessment. The cycle continues until one of the following conditions is met:

[0047] 1) Risk assessment results show no risk: ;

[0048] 2) Maximum number of retries reached: .

[0049] After this step is completed, the patch rule summarizer will generate a patch tag for the meta-cognitive report in the improvement success case and generate a corresponding patch rule content. If the rule pool has already found the existence of the tag when generating the patch tag, only the existing patch rule content will be updated, which is specifically expressed as follows:

[0050] ;

[0051] .

[0052] in, Represents the final metacognitive assessment report, The final answer, represents a function that generates patch labels from metacognitive evaluation reports, Represents a function that generates patch rule content from the final answer and metacognitive evaluation report. By converting metacognitive evaluation results into patch rules and storing them in a dynamic patch rule pool, the model can provide constraints for subsequent similar questions.

[0053] Step 5: Internalize the metacognitive thinking chain through reinforcement learning to strengthen the model

[0054] This step uses a reinforcement learning framework based on policy optimization to internalize the metacognitive thinking chain into the model's deep reasoning ability, enabling it to directly output high-quality answers that meet safety thresholds during the generation phase. The specific technical implementation is as follows:

[0055] 1) Fine-tuning stage

[0056] During the fine-tuning phase, the PPO algorithm is mainly used to optimize the policy network. The PPO algorithm updates the policy network parameters by clipping the objective function to ensure the stability of the policy update. The loss function is expressed as follows:

[0057] .

[0058] Where, Represents a state-action pair expectations, among which The training dataset obtained in the previous steps contains questions in educational scenarios, answers generated by the model, and their corresponding metacognitive evaluation reports. Indicates that the policy network is in the current parameters Next, for the status (including user questions and metacognitive thinking chains) Generate actions the probability of (i.e., the model's answer); Indicates that the old policy network generates an action in state s The probability of , used to compare the differences between the new and old strategies; is the estimated value of the advantage function, which measures the Next action Compared to the additional benefits that the average strategy can bring, this is related to the risk assessment results given by the metacognitive evaluator. If the answer is safe and consistent with the values, the advantage value is high.

[0059] is the clipping function.

[0060] 2) Internalization stage

[0061] The goal of the internalization phase is to internalize the metacognitive thinking chain into the model's deep reasoning ability, so that the model can autonomously perform metacognitive reasoning when generating answers. At this point, the model has been fine-tuned by PPO and can directly output safe answers during the deep thinking phase. This is mainly achieved by optimizing the objective function expressed in the following formula: To achieve:

[0062] .

[0063] in, Indicates status The expected value of sampling from the data distribution D. Represents KL divergence, which is used to measure the difference between two probability distributions; It is the distribution of target strategies implicit in the metacognitive thinking chain; Is the model's current policy network in state This objective function encourages the model’s strategy distribution to move closer to the target strategy distribution in the metacognitive thinking chain, thereby achieving the internalization of metacognitive ability.

[0064] Compared with the prior art, the present invention has the following beneficial technical effects and significant technical progress:

[0065] 1) This invention is highly interpretable and effective. By simulating human metacognitive abilities, it enables in-depth risk assessment and iterative optimization of large-scale model responses in educational scenarios, ensuring that responses not only align with positive educational values ​​but also employ appropriate response strategies for different risk levels. Experiments have shown that the application of this invention to multiple educational datasets significantly reduces the model's jailbreak rate and improves the security and compliance of responses.

[0066] 2) The dynamic patch rule pool mechanism proposed in this paper can convert metacognitive evaluation results into reusable patch rules, realize the self-evolution and continuous improvement of the model, and has good scalability and sustainability.

[0067] 3) This invention can effectively reduce the risky answer rate of large models in educational scenarios, improve the security and value positivity of the model, and does not require additional training costs. It has good scalability and sustainability and can be directly applied to a variety of mainstream large models, providing strong security guarantees for the application of large models in the field of education. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0069] The present invention is further described in detail with reference to the following specific examples and accompanying drawings. The processes, conditions, experimental methods, etc. for implementing the present invention, except for those specifically mentioned below, are common knowledge and common common sense in the art and are not particularly limited by the present invention.

[0070] See Figure 1 The present invention adds information such as the user's learning stage, time and space, and functional role as scenario information, and uses a dynamic risk rule dictionary and an evaluation mechanism that simulates human metacognitive ability to identify and correct risks. Finally, the thought chain is internalized into the model's deep thinking to obtain an efficient safety model. The method specifically includes:

[0071] Step 1: Build a scene layer to embed educational scene information

[0072] Before the user enters the question, the system automatically adds scene information , specifically: .in, Represents user account information, such as academic level information: "junior high school student", "senior high school student", etc. Represents the functional role of the big model, such as "subject tutoring big model." The corresponding functional description is, for example, "In class, students access the education big model through smart devices and input their questions or content they want to learn more about. The model provides personalized knowledge explanations and supplementary materials based on students' learning progress and comprehension ability." Indicates the user's location in time and space, such as "in class at 10 am", "in the dormitory at 9 pm", etc.

[0073] Step 2: Build a dynamic risk rule dictionary with autonomous development capabilities

[0074] Construct a dynamic risk rule dictionary including a static multi-level rule tree and a dynamic patch rule pool, which is specifically expressed as follows:

[0075] ;

[0076] ;

[0077] ;

[0078] .

[0079] , It is an integrated rule of three rules; For macro rules; In response to style rules; For patch rules; The risk levels are represented by the red, yellow and green lines respectively;

[0080] Based on risk issues and scene information , the search head searches from the macro rule area in the static multi-level rule tree, and the macro rule , search from the response style rule area in the static multi-level rule tree to get the response style rule , retrieve from the dynamic patch rule pool to get the patch rules .

[0081] Step 3: Build an automatic value risk assessment mechanism that simulates human metacognitive ability

[0082] The metacognitive evaluator uses the metacognitive thinking chain tool to conduct the assessment. The assessment process generates a detailed assessment report and provides a risk assessment result (1-5 points). The risk scoring criteria are as follows:

[0083] 1 point: Extremely risky, containing content that seriously violates values;

[0084] 2 points: High risk, there is a great risk, and the answer needs to be improved and corrected;

[0085] 3 points: There is a slight risk and positive guidance is needed;

[0086] 4 points: Low risk, the content is generally safe but there is room for improvement;

[0087] 5 points: No risk, the content fully complies with the value requirements.

[0088] The metacognitive evaluator is used in the selection tool When , the dynamic weight is updated according to different scenarios by the following formula :

[0089] .

[0090] in, is the tool utility ratio; is the learning rate parameter.

[0091] Assume an educational scenario and the answers generated by the model The metacognitive evaluator uses the tools in the toolbox to evaluate the answers and generate an evaluation report and risk assessment results. The tool utility ratio The calculation can be performed according to the following steps:

[0092] 1) Calculate contribution: ;

[0093] 2) Update tool utility ratio: ;

[0094] 3) Normalized utility ratio: .

[0095] in, Is to use tools The risk score obtained by the assessment; It is the final risk score obtained by combining the evaluation of all tools; is the learning rate, which controls the magnitude of utility ratio updates.

[0096] Step 4: Build an iterative reinforcement mechanism through metacognitive thinking chains

[0097] This step adopts an iterative improvement strategy, and generally an acceptable risk threshold can be set. The maximum number of retries is 4 or 5. .when When the original question 、Current answer and assessment reports Input the model together. Ask the model to generate improved answers based on the problems pointed out in the evaluation report. . Use the metacognitive evaluator to evaluate the new answer again and get . In this way, the above process is repeated until or .

[0098] Step 5: Internalize the metacognitive thinking chain through reinforcement learning to strengthen the model

[0099] This step uses the PPO algorithm to internalize the metacognitive thinking chain. In the training phase, the training data is used to iteratively optimize the evaluation model in step three. , and then through the constructed Further . Specifically, it can be divided into two stages: fine-tuning and internalization.

[0100] The fine-tuning phase adopts the following loss update strategy network parameters:

[0101] .

[0102] The internalization stage uses the objective function expressed as follows to make the strategy distribution of the model closer to the target strategy distribution in the metacognitive thinking chain:

[0103] .

[0104] The protection content of the present invention is not limited to the above embodiments. Without departing from the spirit and scope of the inventive concept, changes and advantages that can be thought of by those skilled in the art are included in the present invention and are protected by the appended claims.

Claims

1. A large-scale model value alignment reinforcement method in an educational scenario, characterized by: The user's learning stage, time, space, and functional role are added as scenario information. A dynamic risk rule dictionary and an evaluation mechanism that simulates human metacognitive ability are used to identify and correct risks. Finally, the thought chain is internalized into the model's deep thinking to obtain an efficient safety model. The method specifically includes: Step 1: Build a scene layer to embed educational scene information Construct the scenario-level educational scenario information represented by the following formula: ; in, User account information; Functional role for large models; The user's spatial and temporal information; The user account includes the student's academic stage and student learning information; the functional role of the large model Including classroom learning assistants and library learning partners, each role has the following functional description set : ; Step 2: Build a dynamic risk rule dictionary Construct a dynamic risk rule dictionary with autonomous development capabilities, including a static multi-level rule tree, a dynamic patch rule pool, and a search head, and satisfy the integrated rules expressed in the following formula: : ; ; , For macro rules; In response to style rules; For patch rules; The risk levels are represented by the red, yellow and green lines respectively; The search head is based on the risk question and scene information , retrieve from the macro rule area and response style rule area in the static multi-level rule tree to obtain the macro rule and response style rules ; The search head is based on risk issues Retrieve from the dynamic patch rule pool to obtain patch rules ; Step 3: Build an automatic value risk assessment mechanism Construct an automatic value risk assessment mechanism that simulates human metacognitive ability expressed by the following formula: ; ; in, Large models for evaluation; Provide evaluation reports for the metacognitive evaluator; is the evaluation result of the metacognitive evaluator; For the Wheel of Answers; ; The metacognitive evaluator is based on risk questions From the Metacognition Toolbox , select the corresponding subset of thinking tools , and according to Wheel's answer , and obtain the risk assessment results and risk assessment reports ; Step 4: Build an iterative reinforcement mechanism Through the metacognitive thinking chain, that is, risk assessment report , construct an iterative reinforcement mechanism represented by the following formula: ; in, is the maximum number of retries; is the number of retries; when When the risk assessment results , then the large model Leverage risk answers from the previous round , risk assessment report and risk issues Perform iterative optimization to get a new round of answers , until or , ends the iterative optimization, where is the risk assessment threshold; Step 5: Internalize the metacognitive thinking chain through reinforcement learning to strengthen the model Using the PPO algorithm to internalize the metacognitive thinking chain into the big model Deep thinking, thus outputting a large model of safe answers ; The PPO algorithm links metacognitive thinking As the state input of the policy network, define the action space to generate answers for the model The semantic distribution space of , then the reward function It is composed of the following elements: ; in, and They are security rewards, consistency rewards and efficiency rewards is the weight coefficient.

2. The method for aligning and reinforcing large-scale models in educational scenarios according to claim 1 is characterized in that: The macro-rule area consists of the country, society and individuals as the first-level nodes, and splits out into second- and third-level nodes. Each node is equipped with macro-rules corresponding to major or minor categories.

3. The method for aligning and reinforcing large-scale models' values ​​in educational scenarios according to claim 1 is characterized in that: The retrieval head is responsible for finding the patch labels in the dynamic patch rule pool and the patch rule content corresponding to the patch labels. The retrieval head uses the prototype contrast learning strategy to train an embedding model. The training loss function is It is expressed by the following formula: ; in, is the number of risk issues in a mini-batch; Embedding for anchoring questions; is the embedding of the positive semantic prototype; Embedding for irrelevant prototypes; is the temperature coefficient.

4. The method for aligning and reinforcing large-scale models in educational scenarios according to claim 1 is characterized in that: The evaluation result is represented by one of five integers between 1 for extremely high risk and 5 for no risk.

5. The method for aligning and reinforcing large-scale models' values ​​in educational scenarios according to claim 1 is characterized in that: The metacognitive assessment results Dynamic calculation, when Take the maximum value when .

6. The method for aligning and reinforcing large-scale models' values ​​in educational scenarios according to claim 1 is characterized in that: The five steps of metacognitive thinking chain are internalized using a dual-model collaborative mechanism, that is, evaluating the big model Through metacognitive thinking chain Build training data set and pre-trained policy network Under the PPO framework, the policy network is fine-tuned by interacting with the dynamic risk rule dictionary and the metacognitive evaluator. Directly internalize metacognitive reasoning skills to generate responses that meet safety thresholds , forming a large model of the final security answer .

7. The method for aligning and reinforcing large-scale models in educational scenarios according to claim 1 or claim 6, characterized in that: The PPO algorithm is used to internalize the metacognitive thinking chain into the big model In the in-depth thinking, its specific implementation includes: 1) Fine-tuning stage The PPO algorithm is used to optimize the policy network. The PPO algorithm updates the policy network parameters by clipping the objective function. Its loss function is expressed as follows: ; Where, State-action pair expectations, is the training data set; The policy network is at the current parameter Next, for the status Generate Action probability; Generate an action for the old policy network in state s probability; is the estimated value of the advantage function; is the clipping function; 2) Internalization stage By optimizing the objective function expressed in the following formula, we can internalize the metacognitive thinking chain into the model's deep reasoning ability, enabling the model to autonomously perform metacognitive reasoning when generating answers: ; in, Status The expected value of sampling from the data distribution D; is the KL divergence; The distribution of target strategies implicit in the metacognitive thinking chain; The current policy network of the model is in state The distribution of actions under .

8. The method for aligning and reinforcing large-scale models' values ​​in educational scenarios according to claim 1 is characterized in that: The rules in the dynamic patch rule pool are derived by the experience summarizer. The rules in the rule pool exist in the form of label-rule pairs and are expressed as follows: ; in, for the patch label; This is the patch rule content.

9. The method for aligning and reinforcing large-scale models' values ​​in educational scenarios according to claim 1 is characterized in that: The response strategy for the risk levels is: red line issues need to be explicitly prohibited, yellow line issues need to correct deviations, and green line issues need to be actively guided.

Citation Information

Patent Citations

  • Education large model tower type construction method based on multilevel experience learning

    CN119202200A

  • Mobile application large model risk automatic assessment method based on improved BERT

    CN119248643A

  • Value identification data enhancement method based on large language model

    CN119415962A

  • Metacognitive level prediction method based on sounding thinking driven retrieval enhancement

    CN119557848A

Cited By

  • Method for generating safe learning resources based on individual endogenous state variables

    CN122693064A