A Design Method of Reinforcement Learning Reward and Punishment Mechanism for Agent's Chain of Thought
By designing a multi-dimensional instant reward and punishment mechanism in the agent's thinking chain, the existing agent's inaccurate decision-making in complex problems and is limited to a single solution, achieving higher task execution accuracy and diversity.
Patent Information
- Application Number
- CN202510283750.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-11
AI Technical Summary
When existing agents face complex problems, the reward and punishment mechanism of the thinking chain is incomplete and it is difficult to accurately guide decisions at each step, resulting in low accuracy in task execution, and the agents are limited to a single solution and lack the ability to explore diverse paths.
Design a reinforced learning reward and punishment mechanism for the intellectual thinking chain. By setting up an instant reward and punishment mechanism at each thinking chain step, including correct reward and punishment, relevance reward and punishment and information density reward and punishment, calculate the instant reward and punishment and evaluate it through a multi-dimensional reward and punishment mechanism to ensure the accuracy of decision-making at each step, and provide an overall evaluation after the task is over.
Through the instant reward and punishment mechanism, agents can perform tasks more accurately when facing complex problems, improve the accuracy and diversity of task execution, avoid being limited to a single solution, and enhance their ability to respond to different situations and mutated problems.
Smart Images

Figure CN119783760B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method for designing a reinforcement learning reward and punishment mechanism for an agent's chain of thought. Background Art
[0002] Current agents mostly rely on models in natural language processing for language understanding and generation. In dialogue agents, the model performs semantic parsing and context understanding of language through training on a large corpus of texts, and generates answers based on the input content. When facing complex problems, current agents use the "chain of thought" multi-step reasoning method. The result of each step of reasoning serves as the input for the next step. The agent generates a series of state-action pairs based on historical interactions and problem contexts, aiming to simulate the logical reasoning process of human thinking, so as to achieve better performance in complex tasks. With the development of technology, reinforcement learning is being used to train the agent to generate the chain of thought. By providing the agent with an environment and a reference path, the agent learns the strategy for generating the chain of thought path through trial and error. Reinforcement learning mostly conducts reward and punishment evaluation based on reward and punishment signals. The goal of the agent is to maximize the obtained reward and punishment signals and gradually optimize its generation strategy, that is, to select the optimal actions to complete tasks in complex tasks. Many existing systems build agents based on such algorithms, enabling them to complete specific tasks in a given environment.
[0003] However, there are several deficiencies:
[0004] Traditional agents have single actions and lack semantic interaction, so it is necessary to define the reinforcement training of the agent's chain of thought.
[0005] Regarding the imperfect reward and punishment mechanism for the agent's chain of thought, traditional reward and punishment signals rely on the overall evaluation at the end of the task, which is difficult to accurately guide each step of decision-making, and no reward and punishment mechanism is proposed for the chain of thought path planning.
[0006] In the existing reinforcement learning and chain of thought construction methods, agents are often limited to a single solution for specific problems and lack the ability to explore diverse paths, resulting in obvious limitations when dealing with different situations or variant problems.
[0007] Chinese patent document with publication number CN118331199A and publication date July 12, 2024 discloses a multi-task collaborative processing method based on an artificial intelligence agent and an event chain, including the following steps:
[0008] Step S1, using generative artificial intelligence to construct an artificial intelligence agent system to simulate, predict, and optimize the multi-task autonomous scheduling of complex industrial process scenarios;
[0009] Step S2: By performing fusion learning on multi-source and multi-modal data in the industrial process, establish various association models between complex tasks, and use the comprehensive learning ability of the generative pre-training model to comprehensively model multiple tasks in the industrial process;
[0010] Step S3: The artificial intelligence agent system establishes an event trigger and response mechanism between multiple tasks by learning the relevance and interaction effects between complex tasks in the industrial process, and captures the key event process chain and thinking chain in the industrial process;
[0011] Step S4: Through in-depth understanding of the multi-task scenarios in the industrial process by the artificial intelligence agent system, realize the collaborative processing of multiple tasks in complex industrial scenarios.
[0012] The multi-task collaborative processing method based on artificial intelligence agents and event chains disclosed in this patent document proposes a multi-task autonomous scheduling framework for complex scenarios based on the thinking chain. Based on the artificial intelligence agent system, it analyzes the operation requirements of the industrial process and decomposes complex manufacturing tasks into multiple scenario tasks; trained based on specific scenario industrial data sets and domain expert knowledge, a scenario-based industrial large model is obtained to support real-time analysis and intelligent decision-making of scenario tasks; mining the interdependencies between multiple tasks, comprehensively considering the time sensitivity, resource requirements, and priority factors of multiple tasks, can improve the operation efficiency and resource utilization rate of the industrial process. However, each step in the thinking chain cannot receive immediate rewards and punishments, which cannot guide precise decision-making for each step and affects the accuracy of the agent's task execution when facing complex problems. Summary of the Invention
[0013] In order to overcome the above-mentioned defects of the prior art, the present invention provides a method for designing a reinforcement learning reward and punishment mechanism for an agent's thinking chain. Each step in the thinking chain of the present invention can receive immediate rewards and punishments to accurately guide the decision-making of each step. At the same time, after the task is completed, an overall evaluation is provided through the reward and punishment mechanism, enabling the agent to greatly improve the task execution accuracy when facing complex problems.
[0014] The present invention is realized through the following technical solutions:
[0015] A method for designing a reinforcement learning reward and punishment mechanism for an agent's thinking chain, including the following steps:
[0016] S1: State and Action Definition
[0017] Define the current state and executable actions of the agent;
[0018] S2: Sub-task State and Action Planning
[0019] Plan the status and actions of the next subtask based on the execution result of the previous subtask, and use the execution result of each subtask as the input for the next task;
[0020] S3. Construction of the thought chain path
[0021] Define the thought chain path, which consists of multiple state-action pairs. The path goes from the initial state to the target state, and the agent advances the task progress through reasoning and decision-making at each step;
[0022] S4. Reward and punishment mechanism for thought chain steps
[0023] Set up a reward and punishment mechanism for each step in the thought chain path, calculate the immediate reward and punishment, evaluate through a multi-dimensional reward and punishment mechanism, and calculate the total score of the thought chain steps;
[0024] S5. Reward and punishment mechanism for thought chain paths
[0025] Design multiple path reward and punishment dimensions to evaluate the path nature, calculate the total score of the thought chain path nature, and combine it with the total score of the thought chain steps to obtain the overall reward and punishment function.
[0026] In S1, the current state includes the user's question , context information and historical interaction records , and the current state . The executable actions include asking questions , searching , reasoning and generating answers , and the executable actions .
[0027] In S2, the status of the next subtask , where is the execution result of the previous subtask; the action of the next subtask , where is the action plan for the next subtask.
[0028] In S3, the thought chain path is defined as:
[0029] ;
[0030] where is the thought chain path, is the initial state, is the initial action, is the state of the second subtask, is the action of the second subtask, is the end state, It is the end action of the subtask. In step S4, setting a reward and punishment mechanism for each step in the thought chain path means setting a correctness reward and punishment, a relevance reward and punishment, and an information density reward and punishment.
[0031] In step S4, the immediate reward and punishment is calculated by Equation 1;
[0032] Equation 1;
[0033] Where, is the immediate reward and punishment, is the weight of the correctness reward and punishment, is the correctness reward and punishment, is the weight of the relevance reward and punishment, is the relevance reward and punishment, is the weight of the information density reward and punishment, is the information density reward and punishment.
[0034] In step S4, evaluating through a multi-dimensional reward and punishment mechanism means evaluating the correctness score of the answer for each step of the thought chain through the correctness reward and punishment, evaluating the relevance of the actions, states, and dialogue context for each step of the thought chain through the relevance reward and punishment, and quantifying the useful information in the answer for each step of the thought chain through the information density reward and punishment.
[0035] In step S4, the total score of the thought chain steps is calculated by Equation 2;
[0036] Equation 2;
[0037] Where, is the total score of the thought chain steps, is the termination time step of the path.
[0038] In step S5, the total score of the thought chain path nature is calculated by Equation 3;
[0039] Equation 3;
[0040] Where, is the total score of the thought chain path nature, is the linear penalty, is the path diversity reward and punishment, is the path harmony evaluation degree.
[0041] In step S5, the total reward and punishment function is calculated by Equation 4;
[0042] Equation 4;
[0043] Where, is the total reward and punishment function.
[0044] The agent described in the present invention refers to an agent that can perceive the environment and take actions to achieve specific goals, and has autonomy, adaptability, and interaction capabilities; the agent perceives changes in the environment, such as through sensors or data input, makes judgments and decisions based on the knowledge and algorithms learned by itself, and then executes actions to affect the environment or achieve a predetermined goal.
[0045] The thinking chain described in the present invention refers to the process of disassembling a problem with relatively complex logic and forming a complete thinking process through a series of logically related thoughts.
[0046] The policy network described in the present invention is a neural network, and its core function is to input the current environmental state and output the probability distribution of taking each possible action in this state or directly output a specific action; when processing image data, the front end of the policy network usually includes convolutional layers, and these convolutional layers can extract key features of the environmental state.
[0047] The BM25 described in the present invention is an information retrieval algorithm used to evaluate the relevance between a document and a query, and evaluates the relevance of the document by calculating the weighted sum of the term frequency and the inverse document frequency.
[0048] The beneficial effects of the present invention are mainly manifested in the following aspects:
[0049] 1. In the present invention, compared with the prior art, each step in the thinking chain can obtain immediate rewards and punishments to accurately guide the decision-making of each step. At the same time, after the task is completed, an overall evaluation is provided through the reward and punishment mechanism, enabling the agent to greatly improve the task execution accuracy when facing complex problems.
[0050] 2. In the present invention, a mathematical definition is given to the existing agent thinking chain, and the inference path evaluation of the agent is optimized through a multi-dimensional reward and punishment mechanism, solving the problem of insufficient reward and punishment mechanism in existing reinforcement learning.
[0051] 3. In the present invention, a diversity reward and punishment mechanism is used to guide the agent to explore more paths, avoiding the agent being limited to a single solution method, and the solution methods are more diverse.
[0052] 4. In the present invention, the matching degree of the actions, states, and dialogue contexts of each step of the thinking chain can be quantified through BM25.
[0053] 5. In the present invention, the balance of different paths generated by the question can be quantified, and the coverage rate and accuracy rate of key points in the path are evaluated. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The present invention will be further specifically described below in conjunction with the accompanying drawings of the specification and the specific implementation manners:
[0055] Figure 1This is the flowchart of the present invention. Detailed implementation manners
[0056] Example 1
[0057] Refer to Figure 1 , a method for designing a reinforcement learning reward and punishment mechanism for an agent's thought chain, comprising the following steps:
[0058] S1. State and action definition
[0059] Define the current state and executable actions of the agent;
[0060] S2. Sub-task state and action planning
[0061] According to the execution result of the previous sub-task, plan the state of the next sub-task and the action of the next sub-task, and use the execution result of each sub-task as the input of the next task;
[0062] S3. Thought chain path construction
[0063] Define the thought chain path, which consists of multiple groups of state-action pairs. The path goes from the initial state to the target state, and the agent advances the task progress through reasoning and decision-making at each step;
[0064] S4. Thought chain step reward and punishment mechanism
[0065] Set up a reward and punishment mechanism at each step in the thought chain path, calculate the immediate reward and punishment, evaluate through a multi-dimensional reward and punishment mechanism, and calculate the total score of the thought chain steps;
[0066] S5. Thought chain path reward and punishment mechanism
[0067] Design multiple path reward and punishment dimensions to evaluate the path nature, calculate the total score of the thought chain path nature, and combine it with the total score of the thought chain steps to obtain the overall reward and punishment function.
[0068] This embodiment is the most basic implementation manner. Compared with the prior art, each step in the thought chain can obtain immediate rewards and punishments to accurately guide the decision-making at each step. At the same time, after the task is completed, an overall evaluation is provided through the reward and punishment mechanism, enabling the agent to greatly improve the task execution accuracy when facing complex problems.
[0069] Example 2
[0070] Refer to Figure 1 , a method for designing a reinforcement learning reward and punishment mechanism for an agent's thought chain, comprising the following steps:
[0071] S1. State and action definition
[0072] Define the current state and executable actions of the agent;
[0073] S2, Sub - task Status and Action Planning
[0074] Based on the execution result of the previous sub - task, plan the status of the next sub - task and the action of the next sub - task, and use the execution result of each sub - task as the input for the next task;
[0075] S3, Thought Chain Path Construction
[0076] Define the thought chain path, which consists of multiple state - action pairs. The path goes from the initial state to the target state, and the agent advances the task progress through reasoning and decision - making at each step;
[0077] S4, Thought Chain Step Reward and Punishment Mechanism
[0078] Set up a reward and punishment mechanism for each step in the thought chain path, calculate the immediate reward and punishment, evaluate through a multi - dimensional reward and punishment mechanism, and calculate the total score of the thought chain steps;
[0079] S5, Thought Chain Path Reward and Punishment Mechanism
[0080] Design multiple path reward and punishment dimensions to evaluate the path nature, calculate the total score of the thought chain path nature, and combine it with the total score of the thought chain steps to obtain the overall reward and punishment function.
[0081] In the above - mentioned S1, the current state includes the user's question , context information and historical interaction records , the current state , and the executable actions include asking questions , searching , reasoning and generating answers , the executable actions .
[0082] In the above - mentioned S2, the status of the next sub - task , where is the execution result of the previous sub - task; the action of the next sub - task , where is the action plan of the next sub - task.
[0083] This embodiment is a preferred embodiment, which mathematically defines the existing thought chain of the agent, optimizes the evaluation of the agent's reasoning path through a multi - dimensional reward and punishment mechanism, and solves the problem of insufficient reward and punishment mechanism in existing reinforcement learning.
[0084] Embodiment 3
[0085] See Figure 1, A method for designing a reinforcement learning reward and punishment mechanism for an agent's thought chain, comprising the following steps:
[0086] S1. State and action definition
[0087] Define the current state and executable actions of the agent;
[0088] S2. Sub-task state and action planning
[0089] According to the execution result of the previous sub-task, plan the state of the next sub-task and the action of the next sub-task, and use the execution result of each sub-task as the input for the next task;
[0090] S3. Thought chain path construction
[0091] Define the thought chain path, which consists of multiple groups of state-action pairs. The path goes from the initial state to the target state, and the agent advances the task progress through reasoning and decision-making at each step;
[0092] S4. Thought chain step reward and punishment mechanism
[0093] Set up a reward and punishment mechanism for each step in the thought chain path, calculate the immediate reward and punishment, evaluate through a multi-dimensional reward and punishment mechanism, and calculate the total score of the thought chain steps;
[0094] S5. Thought chain path reward and punishment mechanism
[0095] Design multiple path reward and punishment dimensions to evaluate the path nature, calculate the total score of the thought chain path nature, and combine it with the total score of the thought chain steps to obtain the overall reward and punishment function.
[0096] In S1, the current state includes the user's question , context information and historical interaction records , the current state , and the executable actions include asking questions , searching , reasoning and generating answers , the executable actions .
[0097] In S2, the state of the next sub-task , where is the execution result of the previous sub-task; the action of the next sub-task , where is the action plan of the next sub-task.
[0098] In S3, the thought chain path is defined as:
[0099] ;
[0100] Among them, is the thought chain path, is the initial state, is the initial action, is the state of the second subtask, is the action of the second subtask, is the end state, is the end action of the subtask. In S4, setting a reward and punishment mechanism for each step in the thought chain path means setting a correctness reward and punishment, a relevance reward and punishment, and an information density reward and punishment.
[0101] In S4, the immediate reward and punishment is calculated by Equation 1;
[0102] Equation 1;
[0103] Among them, is the immediate reward and punishment, is the weight of the correctness reward and punishment, is the correctness reward and punishment, is the weight of the relevance reward and punishment, is the relevance reward and punishment, is the weight of the information density reward and punishment, is the information density reward and punishment.
[0104] In S4, evaluating through a multi-dimensional reward and punishment mechanism means evaluating the correctness score of the answer for each step of the thought chain through the correctness reward and punishment, evaluating the relevance of the action, state, and dialogue context for each step of the thought chain through the relevance reward and punishment, and quantifying the useful information in the answer for each step of the thought chain through the information density reward and punishment.
[0105] In S4, the total score of the thought chain steps is calculated by Equation 2;
[0106] Equation 2;
[0107] Among them, is the total score of the thought chain steps, is the termination time step of the path.
[0108] This embodiment is another preferred embodiment. By using a diversity reward and punishment mechanism to guide the agent to explore more paths, it avoids the agent being limited to a single solution and makes the solutions more diverse.
[0109] Embodiment 4
[0110] See Figure 1 , a method for designing a reinforcement learning reward and punishment mechanism for an agent's thought chain, comprising the following steps:
[0111] S1. State and Action Definition
[0112] Define the current state and executable actions of the agent;
[0113] S2. Sub - task State and Action Planning
[0114] According to the execution result of the previous sub - task, plan the state of the next sub - task and the actions of the next sub - task, and use the execution result of each sub - task as the input for the next task;
[0115] S3. Thought Chain Path Construction
[0116] Define the thought chain path, which consists of multiple state - action pairs. The path goes from the initial state to the target state, and the agent advances the task progress through reasoning and decision - making at each step;
[0117] S4. Thought Chain Step Reward and Punishment Mechanism
[0118] Set up a reward and punishment mechanism for each step in the thought chain path, calculate the immediate reward and punishment, evaluate through a multi - dimensional reward and punishment mechanism, and calculate the total score of the thought chain steps;
[0119] S5. Thought Chain Path Reward and Punishment Mechanism
[0120] Design multiple path reward and punishment dimensions to evaluate the path nature, calculate the total score of the thought chain path nature, and combine it with the total score of the thought chain steps to obtain the overall reward and punishment function.
[0121] In the above - mentioned S1, the current state includes the user's question , context information and historical interaction records , the current state , and the executable actions include asking questions , searching , reasoning and generating answers , the executable actions .
[0122] In the above - mentioned S2, the state of the next sub - task , where is the execution result of the previous sub - task; the actions of the next sub - task , where is the action plan of the next sub - task.
[0123] In the above - mentioned S3, the thought chain path is defined as:
[0124] ;
[0125] Among them, is the thought chain path, is the initial state, is the initial action, is the state of the second subtask, is the action of the second subtask, is the end state, is the end action of the subtask. In S4, setting a reward and punishment mechanism for each step in the thought chain path means setting correctness rewards and punishments, relevance rewards and punishments, and information density rewards and punishments.
[0126] In S4, the immediate reward and punishment is calculated by Equation 1;
[0127] Equation 1;
[0128] Among them, is the immediate reward and punishment, is the weight of the correctness reward and punishment, is the correctness reward and punishment, is the weight of the relevance reward and punishment, is the relevance reward and punishment, is the weight of the information density reward and punishment, is the information density reward and punishment.
[0129] In S4, evaluating through a multi-dimensional reward and punishment mechanism means evaluating the correctness score of the answer for each step of the thought chain through correctness rewards and punishments, evaluating the relevance of the actions, states, and dialogue context for each step of the thought chain through relevance rewards and punishments, and quantifying the useful information in the answer for each step of the thought chain through information density rewards and punishments.
[0130] In S4, the total score of the thought chain steps is calculated by Equation 2;
[0131] Equation 2;
[0132] Among them, is the total score of the thought chain steps, is the termination time step of the path.
[0133] In S5, the total score of the thought chain path properties is calculated by Equation 3;
[0134] Equation 3;
[0135] Among them, is the total score of the thought chain path properties, is the linear penalty, is the path diversity reward and punishment, is the path harmony evaluation degree.
[0136] In S5, the total reward and penalty function is calculated by Equation 4;
[0137] Equation 4;
[0138] Among them, is the total reward and penalty function.
[0139] This embodiment is the best implementation mode, which can quantify the matching degree of the actions, states, and dialogue contexts of each step of the thought chain through BM25.
[0140] It can quantify the balance of different paths generated by the question, and evaluate the coverage rate and accuracy of key points in the path.
[0141] The thought chain step reward and penalty mechanism includes correctness reward and penalty , relevance reward and penalty and information density reward and penalty .
[0142] Correctness reward and penalty : If a certain step of the system gives a wrong answer or inappropriate reasoning, the score should be negative or zero. By evaluating the correctness of each step, positive and negative rewards and penalties are given:
[0143]
[0144] Relevance reward and penalty : Based on the BM25 method of matching metrics, it measures the matching degree between the answer and the question or context.
[0145]
[0146] BM25 is a ranking function in information retrieval, used to evaluate the relevance between a document and a query. BM25 scores documents based on word frequency and document length, and is suitable for evaluating the relevance between an answer and a question and context.
[0147]
[0148] Among them, is the word frequency of the word in the answer ; is the length of the answer ; is the average length of all answers and is a constant; is a tuning parameter, with a value of 1.2; is a tuning parameter, with a value of 0.75; is the word Inverse document frequency.
[0149]
[0150] Among them, is the total number of the document set, is the number of documents containing the word .
[0151] Information density reward and punishment :
[0152] Information density measures the ratio of useful information to total information in an answer.
[0153] Information density is defined by measuring the ratio of the key information volume in an answer to the total length of the answer:
[0154]
[0155] Among them, is the length of the key information volume in the answer, is the total length of the answer.
[0156] The thought chain path reward and punishment mechanism includes linear punishment , path diversity reward and punishment and path harmonic evaluation degree .
[0157] Path diversity reward and punishment : Path diversity reward and punishment is used to encourage the system to generate more diverse problem-solving methods and avoid falling into local optimal solutions. High-diversity paths can cover more possible solution spaces, improving the flexibility of task completion and the robustness of the system.
[0158] A thought chain path is a sequence composed of a series of subtasks, expressed as:
[0159]
[0160] Among them, is the subtask of the thought chain path at time step t. The subtask is a different action decision, reasoning step or problem decomposition, expressed as , is the state-action pair.
[0161] Path diversity definition: For the same problem, set a group of thought chain paths , each path represents a decision sequence or chain of thought of the system for a certain task, and the diversity of the paths is reflected in the differences between the paths. The diversity of the paths is judged by measuring the similarity between different paths. The less similar two paths are, the higher the diversity. The difference between paths is measured by the edit distance or the hash distance.
[0162] Coverage Reward and Penalty : Coverage Reward and Penalty is used to evaluate whether the generated chain of thought paths cover all the key points in the question or meet the key requirements of the task goal. The coverage rate is calculated by comparing with a predefined reference answer or task goal. The coverage rate is defined as the degree of overlap between the system output and the target content.
[0163] Precision : Precision is used to measure how many relevant key points there are in the content generated by the system; it is defined as the proportion of key points that truly belong to the reference answer in the content generated by the system:
[0164]
[0165] Among them, is the number of key steps included in the intersection of the system-generated path and the reference path, is the number of all steps of the generated path.
[0166] Precision considers the accuracy of the generated steps and avoids the system generating a large number of irrelevant or redundant steps. The value range of precision is [0,1], where 1 means that all the generated steps hit the key information, and 0 means that no key information is hit.
[0167] Path Harmony Evaluation Degree : By introducing a weighting coefficient to adjust the Coverage Reward and Penalty and Precision of the relative importance:
[0168]
[0169] Among them, is the weighting coefficient;
[0170] When > 1, the weight of the Coverage Reward and Penalty is larger, emphasizing finding more key information;
[0171] When < 1, the weight of Precision is larger, emphasizing that the generated path is more accurate and contains less irrelevant information.
[0172] To prevent the situation of Coverage Reward and Penalty or extreme Precision from occurring, a penalty function is introduced for dynamic adjustment.
[0173]
[0174] Among them, is a penalty coefficient used to control the intensity of the penalty;
[0175] Path harmonic evaluation degree .
Claims
1. A method for designing a reinforcement learning reward and punishment mechanism for an intelligent agent thinking chain, characterized in that: The following steps are involved: S1. State and action definition Define the agent's current state and executable actions; S2. Subtask status and action planning According to the execution result of the previous subtask, plan the state and action of the next subtask, and use the execution result of each subtask as the input of the next task; S3. Construction of thinking chain path Define a thought chain path, which consists of multiple sets of state-action pairs. The path goes from the initial state to the target state. The agent advances the task through reasoning and decision-making at each step. S4. Reward and Penalty Mechanism for Thinking Chain Steps Set up a reward and penalty mechanism for each step in the thinking chain path, and calculate the immediate reward and penalty. Evaluate through the multi-dimensional reward and penalty mechanism to calculate the total score of the thinking chain steps; S5. Thinking chain path reward and punishment mechanism Design multiple path reward and punishment dimensions to evaluate the path properties, calculate the total score of the thinking chain path properties, and combine the total score of the thinking chain steps to obtain the total reward and punishment function; In S1, the current state includes user questions , context information Interaction history , current status , executable actions include asking questions ,search ,reasoning and generate answers , executable actions ; In S4, setting a reward and penalty mechanism for each step in the thinking chain path refers to setting a correctness reward and penalty, a relevance reward and penalty, and an information density reward and penalty; In said S4, the instant reward and penalty are calculated by formula 1; Formula 1; in, For immediate rewards and punishments, is the weight of the correctness reward and penalty, Rewards and penalties for correctness, is the weight of the relevance reward and penalty, Rewards and penalties for relevance, is the weight of information density reward and penalty, rewards and penalties for information density; In S4, the total score of the thought chain steps is calculated by formula 2; Formula 2; in, Give a total score for the thought chain steps. is the termination time step of the path; in S5, the total score of the thinking chain path property is calculated by formula 3; Formula 3; in, is the total score of the path nature of the thinking chain, is a linear penalty, Rewards and penalties for path diversity, is the path reconciliation evaluation degree.
2. The method for designing a reinforcement learning reward and punishment mechanism for an intelligent agent thinking chain according to claim 1, characterized in that: In S2, the state of the next subtask ,in The execution result of the previous subtask; the action of the next subtask ,in Action planning for the next subtask.
3. The method for designing a reinforcement learning reward and punishment mechanism for an intelligent agent thinking chain according to claim 1, characterized in that: In S3, the thought chain path is defined as: ; in, For the thinking chain path, is the initial state, For the initial action, is the status of the second subtask, is the action of the second subtask, For the end state, The end action of the subtask.
4. The method for designing a reinforcement learning reward and punishment mechanism for an intelligent agent thinking chain according to claim 1, characterized in that: In S4, the evaluation through a multi-dimensional reward and penalty mechanism refers to evaluating the correctness score of the answer to each step of the thinking chain through correctness rewards and penalties, evaluating the relevance of the action, state and dialogue context of each step of the thinking chain through relevance rewards and penalties, and quantifying the useful information in the answer to each step of the thinking chain through information density rewards and penalties.
5. The method for designing a reinforcement learning reward and punishment mechanism for an intelligent agent thinking chain according to claim 1, characterized in that: In S5, the total reward and penalty function is calculated by formula 4; Formula 4; in, is the total reward and penalty function.
Citation Information
Patent Citations
Multi-task cooperative processing method based on artificial intelligence agent and event chain
CN118331199A
Heterogeneous network multi-dimensional resource collaborative optimization method based on deep reinforcement learning
CN116132304A
Systems and methods for training an autonomous machine to perform an operation
US20240419977A1