Navigation method, device and equipment based on GRPO algorithm and medium
By dynamically adjusting the policy update step size factor and gradient estimation correction term, the problems of excessive policy update magnitude and data distribution changes in the GRPO algorithm are solved, improving training stability and agent navigation accuracy. It is suitable for reinforcement learning tasks such as robot navigation, autonomous driving and game AI.
Patent Information
- Application Number
- CN202511123956.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-12
- Publication Date
- 2025-11-18
AI Technical Summary
The GRPO algorithm suffers from problems such as excessive policy update magnitude, data distribution changes, and gradient estimation bias during training, leading to training instability and the agent's inability to effectively complete the task.
The step size factor is dynamically adjusted by calculating the similarity between the current policy and historical policies and the rate of change of reward. The objective function of the GRPO algorithm is modified by using gradient estimation correction terms and importance weight pruning. Combined with priority storage and sampling of historical data, the policy update is optimized.
The training stability and convergence speed of the GRPO algorithm have been improved, and its generalization ability in different environments and tasks has been enhanced, ensuring that the agent can quickly and accurately find the maze exit.
Smart Images

Figure CN120970648A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of reinforcement learning technology, and in particular to a navigation method, apparatus, device, and medium based on the GRPO algorithm. Background Technology
[0002] Reinforcement learning, an important branch of machine learning, aims to enable agents to learn optimal policies through interaction with the environment to achieve specific goals. The GRPO (Group Relative Policy Optimization) algorithm, as an advanced achievement in reinforcement learning, has shown certain advantages in handling complex tasks. However, in practical applications, the training process of the GRPO algorithm often faces instability issues, for the following reasons:
[0003] 1. Excessive Policy Update Amplitude: When updating the policy, if the GRPO algorithm updates the policy by too large an amplitude, the policy may quickly deviate from the current optimal policy region. Taking a reinforcement learning task of training a robot to walk as an example, when the policy update amplitude is too large, the robot may suddenly change its walking mode, such as gait or stride size, causing it to fall and be unable to continue completing the task. Consequently, it will not be able to obtain subsequent reward signals, which will seriously affect the stability of training.
[0004] 2. Changes in Data Distribution: As the strategy is continuously updated, the distribution of data generated by the agent's interaction with the environment will also change. The GRPO algorithm typically assumes a relatively stable data distribution during training, but in reality, rapid changes in data distribution can make the algorithm difficult to adapt. For example, in a game scenario, the agent's initial strategy is to explore the map, at which point the data is mainly exploration data; as the strategy is updated, it gradually shifts to attacking enemies, at which point the data distribution changes from being dominated by exploration data to being dominated by combat data. The algorithm may not be able to adjust in time to adapt to this change, leading to unstable training.
[0005] 3. Gradient estimation bias: When calculating the policy gradient, the GRPO algorithm may be affected by gradient estimation bias. This bias may be caused by sampling errors, model approximation errors, etc., and it can cause the policy to be updated in the wrong direction, thus affecting the stability of training. For example, in an autonomous driving task, due to sampling errors, the algorithm may incorrectly estimate the gradient for taking an acceleration policy under a certain road condition, leading to dangerous behavior by the agent in actual driving. Summary of the Invention
[0006] In view of this, the purpose of this invention is to provide a navigation method, apparatus, device, and medium based on the GRPO algorithm, which can solve the problems of excessive policy update amplitude, data distribution variation, and gradient estimation bias in the training process of existing GRPO algorithms, thereby improving the performance and reliability of the algorithm. The specific solution is as follows:
[0007] In a first aspect, this application discloses a navigation method based on the GRPO algorithm, comprising:
[0008] The KL divergence is determined by the state space and action space, and the average similarity between the current policy and the policies of the previous several iterations is calculated based on the KL divergence; the state space includes each position in the target maze environment; the action space includes the movement direction of the agent;
[0009] The reward obtained by the agent in each iteration for successfully finding the exit is determined, and the updated step size factor is obtained by updating the step size factor based on the average reward change rate determined by the rewards of several iterations and the average similarity.
[0010] The gradient estimation correction term is determined based on the gradient estimation of the sampled trajectory. The target gradient estimation is determined based on the gradient estimation correction term and the original gradient estimation of the current policy. The current policy is updated using the target gradient estimation and the updated step size factor. The sampled trajectory is the trajectory obtained by trajectory sampling during the interaction between the agent and the target maze environment.
[0011] When updating the current strategy, the importance weights are calculated, the importance weights are pruned to obtain the pruned weights, and the objective function of the GRPO algorithm is modified according to the pruned weights to obtain the modified function.
[0012] The GRPO algorithm is trained based on the modified function and the corresponding updated policy, so that the agent learns the optimal policy based on the trained GRPO algorithm and determines the exit of the target maze environment according to the optimal policy.
[0013] Optionally, determining the KL divergence through the state space and action space, and calculating the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence, includes:
[0014] Sampling was performed in the target maze environment to count the access frequency of each state under each strategy;
[0015] The KL divergence is determined using the state space, action space, and access frequency according to the KL divergence calculation formula; the KL divergence calculation formula is:
[0016] ;
[0017] in, Let KL divergence be the value of KL. In strategy The access frequency of the lower state s; S is the state space; A is the action space; This is the current strategy;
[0018] Calculate the average KL divergence of the strategy in the first few iterations, and determine the average similarity as the average similarity.
[0019] Optionally, the step-size factor is updated by using the average rate of change of reward determined based on rewards over several iterations and the average similarity to update the step-size factor, resulting in the updated step-size factor, including:
[0020] The average reward change rate is determined based on the rewards from the aforementioned several iterations using the average reward change rate calculation formula; the average reward change rate calculation formula is as follows:
[0021] ;
[0022] Where R is the average rate of change of reward; m is the number of iterations; This is the reward for the i-th iteration; This is the reward for the (i-1)th iteration;
[0023] If the average similarity is less than the first threshold or the average reward change rate is greater than the second threshold, then the product between the initial step size factor and the first value is determined as the updated step size factor.
[0024] If the average similarity is greater than or equal to the first threshold or the average reward change rate is less than or equal to the second threshold, then the product between the initial step size factor and the second value is determined as the updated step size factor.
[0025] Wherein, the first value is less than 1; the second value is greater than 1.
[0026] Optionally, the gradient estimation based on the sampled trajectory determines the gradient estimation correction term, including:
[0027] The gradient estimation correction term is determined based on the gradient estimation of the sampled trajectory using the gradient estimation correction term calculation formula; the gradient estimation correction term calculation formula is as follows:
[0028] ;
[0029] in, The gradient estimation correction term is denoted by N; N is the number of sampling trajectories. This is the original gradient estimate for the current policy; The dominant function; The gradient of the logarithmic probability of the action; This is for gradient estimation of each sampling trajectory.
[0030] Optionally, determining the target gradient estimate based on the gradient estimation correction term and the original gradient estimate of the current policy includes:
[0031] The gradient estimation correction term is weighted and summed with the original gradient estimate of the current policy to obtain the corresponding weighted summation result;
[0032] The weighted summation result is determined as the target gradient estimate.
[0033] Optionally, the calculation of importance weights, and the pruning of the importance weights to obtain pruned weights, includes:
[0034] The importance weight is determined by the ratio between the probability of taking the target action under the current policy in the target state and the probability of taking the target action under the target state using the policy used when generating the data.
[0035] The importance weights are pruned by setting a preset pruning threshold to obtain the pruned weights;
[0036] Accordingly, the step of modifying the objective function of the GRPO algorithm based on the pruned weights to obtain the modified function includes:
[0037] The objective function of the GRPO algorithm is modified based on the pruned weights using the modified function determination formula; the modified function determination formula is as follows:
[0038] ;
[0039] in, The corrected function; The weights after clipping; The dominant function; For strategy The expectation of taking action a in state s.
[0040] Optionally, the method further includes:
[0041] The priority of historical data is determined by the sum of the absolute value of the product between the time difference error and the pruned weights and the target constant.
[0042] The historical data is stored in a pre-built data buffer according to its priority.
[0043] During the training of the GRPO algorithm, historical data on the number of targets are sampled from the data buffer according to the cropped weights, and the GRPO algorithm is trained based on the historical data on the number of targets.
[0044] Secondly, this application discloses a navigation device based on the GRPO algorithm, comprising:
[0045] The average similarity calculation module is used to determine the KL divergence through the state space and action space, and calculate the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence; the state space includes each position in the target maze environment; the action space includes the movement direction of the agent;
[0046] The step size factor update module is used to determine the reward obtained by the agent in each iteration for successfully finding the exit. The step size factor is updated by the average reward change rate determined based on the rewards of several iterations and the average similarity.
[0047] The policy update module is used to determine a gradient estimation correction term based on the gradient estimation of the sampled trajectory, determine a target gradient estimation based on the gradient estimation correction term and the original gradient estimation of the current policy, and update the current policy using the target gradient estimation and the updated step size factor; the sampled trajectory is the trajectory obtained by trajectory sampling during the interaction between the agent and the target maze environment;
[0048] The function correction module is used to calculate importance weights when updating the current strategy, prune the importance weights to obtain pruned weights, and correct the objective function of the GRPO algorithm based on the pruned weights to obtain the corrected function.
[0049] The exit determination module is used to train the GRPO algorithm based on the modified function and the corresponding updated policy, so that the agent learns the optimal policy based on the corresponding trained GRPO algorithm and determines the exit of the target maze environment according to the optimal policy.
[0050] Thirdly, this application discloses an electronic device, comprising:
[0051] Memory, used to store computer programs;
[0052] A processor is used to execute computer programs to implement navigation methods based on the GRPO algorithm as described above.
[0053] Fourthly, this application discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the navigation method based on the GRPO algorithm as described above.
[0054] This application first determines the KL divergence through the state space and action space, and calculates the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence; the state space includes each position in the target maze environment; the action space includes the movement direction of the agent; the reward obtained by the agent in successfully finding the exit in each iteration is determined, and the step size factor is updated by the average reward change rate determined based on the rewards of several iterations and the average similarity, resulting in the updated step size factor; a gradient estimation correction term is determined based on the gradient estimation of the sampled trajectory, and the target gradient estimation is determined based on the gradient estimation correction term and the original gradient estimation of the current policy, through... The target gradient estimation and the updated step size factor are used to update the current policy; the sampled trajectory is the trajectory obtained through trajectory sampling during the interaction between the agent and the target maze environment; when updating the current policy, importance weights are calculated, the importance weights are pruned to obtain pruned weights, and the objective function of the GRPO algorithm is modified according to the pruned weights to obtain the modified function; the GRPO algorithm is trained based on the modified function and the corresponding updated policy so that the agent learns the optimal policy based on the corresponding trained GRPO algorithm and determines the exit of the target maze environment according to the optimal policy. It can be seen that this application effectively solves the problems of excessive policy update magnitude, data distribution changes, and gradient estimation bias in the GRPO algorithm training process by dynamically adjusting the policy update step size factor and utilizing historical data and corrected gradient estimation. Furthermore, during the navigation process, this invention avoids the problem of the agent getting lost in the maze due to excessively large policy updates. It also avoids the problem of the algorithm being unable to adapt to the task requirements at different stages of the agent's navigation due to changes in data distribution, and the problem of gradient estimation bias affecting the agent's ability to learn the correct policy while searching for the exit. This allows the agent to accurately and quickly find the maze exit. Simultaneously, this invention significantly improves the algorithm's training stability, accelerates convergence speed, and enhances its performance and reliability. It also enhances the algorithm's generalization ability in different environments and tasks, providing a more reliable and efficient solution for complex reinforcement learning tasks. Attached Figure Description
[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0056] Figure 1 This is a flowchart of a navigation method based on the GRPO algorithm disclosed in this application;
[0057] Figure 2 This is a schematic diagram of a navigation device based on the GRPO algorithm disclosed in this application;
[0058] Figure 3 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation
[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] Current GRPO algorithm training processes suffer from problems such as excessive policy update magnitude, data distribution variations, and gradient estimation bias. To address these issues, this application discloses a navigation method, apparatus, device, and medium based on the GRPO algorithm. This method resolves the problems of excessive policy update magnitude, data distribution variations, and gradient estimation bias in existing GRPO algorithm training processes, thereby improving the algorithm's performance and reliability.
[0061] See Figure 1 As shown, this embodiment of the invention discloses a navigation method based on the GRPO algorithm, including:
[0062] Step S11: Determine the KL divergence through the state space and action space, and calculate the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence; the state space includes each position in the target maze environment; the action space includes the movement direction of the agent.
[0063] In this embodiment, during the reinforcement learning task of training a robot to walk, when the policy update magnitude is too large, the robot may suddenly change its walking pattern, such as gait or stride length, leading to a fall and inability to continue the task. Consequently, it will be unable to obtain subsequent reward signals, severely affecting the stability of the training. Therefore, this application introduces a dynamically adjusted policy update step size factor α, which is adjusted based on the similarity between the current policy and historical policies, as well as the reward changes during the training process. First, the current policy is calculated. The average similarity S with the strategies of the previous k iterations is used, and the similarity can be measured by metrics such as KL divergence (KL divergence). Sampling is performed in the target maze environment to statistically analyze the access frequency of each state under each strategy. The KL divergence is determined using the state space, action space, and access frequency according to the KL divergence calculation formula. The KL divergence calculation formula is as follows:
[0064] ;
[0065] in, Let KL divergence be the value of KL. In strategy The access frequency of the lower state s; S is the state space; A is the action space; Given the current strategy; calculate the average KL divergence of the strategy over the previous several iterations, and determine the average similarity as the average similarity. To calculate the average similarity S, average the KL divergence of the strategy over the previous k iterations.
[0066] This application uses the task of training an agent to find an exit in a maze environment as an example. The maze environment has multiple rooms and passages, and the agent needs to learn the optimal policy by interacting with the environment (moving and observing the surroundings) to find the exit as quickly as possible. During training, the GRPO algorithm is used. At the beginning of training, the policy update step size factor α=1 is initialized, k=5 is set (i.e., the average similarity of the policies in the first 5 iterations is calculated), and a threshold is set. =0.1. The current policy is calculated after each policy iteration. The average KL divergence S with the policies of the previous 5 iterations. Assume the current state space S contains all positions in the maze, and the action space A contains the agent's movement directions (up, down, left, right). By sampling in the maze environment, the frequency of state visits under each policy is statistically analyzed. Then, the KL divergence between the strategies in each iteration is calculated according to the KL divergence formula. The average similarity is obtained by averaging the results. It should be noted that an intelligent agent is a proxy capable of perceiving its environment and taking actions to achieve a specific goal. It can be software, hardware, or a system, possessing autonomy, adaptability, and interactivity. The intelligent agent in this application can be a hardware entity in the form of a robot.
[0067] Step S12: Determine the reward obtained by the agent in each iteration for successfully finding the exit, and update the step size factor by using the average reward change rate determined based on the rewards of several iterations and the average similarity.
[0068] In this embodiment, after calculating the similarity, the average reward change rate R over the most recent m iterations is calculated. The average reward change rate is determined based on the rewards from the aforementioned iterations using the average reward change rate calculation formula:
[0069] ;
[0070] Where R is the average rate of change of reward; m is the number of iterations; This is the reward for the i-th iteration; This is the reward for the (i-1)th iteration.
[0071] Then, the policy update step size factor α is dynamically adjusted based on the similarity S0 and the reward change rate R. When the similarity is low or the reward change rate is large, the value of α is decreased to limit the policy update magnitude; when the similarity is high and the reward change rate is small, the value of α is appropriately increased to accelerate the policy update speed. For example, a threshold is set. and When similarity S0 < or R> When α = α × β (β < 1); when S0 ≥ Sth and R ≤ When α = α × γ (γ > 1), where β < 1 and γ > 1. For example, β = 0.9 and γ = 1.1 can be set. and This can be determined through grid search. Specifically, if the average similarity is less than a first threshold or the average reward change rate is greater than a second threshold, the product of the initial step size factor and the first value is determined as the updated step size factor; if the average similarity is greater than or equal to the first threshold or the average reward change rate is less than or equal to the second threshold, the product of the initial step size factor and the second value is determined as the updated step size factor; wherein the first value is less than 1; and the second value is greater than 1.
[0072] In one specific embodiment, m is set to 6 (i.e., the average rate of change of reward over the most recent 6 iterations is calculated), and a threshold is set. =0.2. At the end of each iteration, record the reward the agent receives for successfully finding the exit in that iteration. (If no exit is found, the reward is negative or zero). Calculate the average reward change rate R over the last 10 iterations using the formula for the average reward change rate. Adjust the policy update step size factor α based on the similarity S0 and the reward change rate R. Set β=0.9 and γ=1.1. If S0<0.1 or R>0.2, then α=α×0.9; if S0≥0.1 and R≤0.2, then α=α×1.1. In subsequent policy updates, use the adjusted α to control the magnitude of the policy update.
[0073] Step S13: Determine the gradient estimation correction term based on the gradient estimation of the sampled trajectory; determine the target gradient estimation based on the gradient estimation correction term and the original gradient estimation of the current policy; update the current policy using the target gradient estimation and the updated step size factor; the sampled trajectory is the trajectory obtained by trajectory sampling during the interaction between the agent and the target maze environment.
[0074] In this embodiment, the GRPO algorithm may be affected by gradient estimation bias when calculating the policy gradient. This bias may be caused by sampling errors, model approximation errors, etc., which can lead the policy to be updated in the wrong direction, thus affecting the stability of training. For example, in an autonomous driving task, due to sampling errors, the algorithm may incorrectly estimate the gradient of the acceleration policy under a certain road condition, causing the agent to exhibit dangerous behavior in actual driving. Therefore, this application introduces a gradient estimation correction term to reduce gradient estimation bias. When calculating the policy gradient, in addition to the original gradient estimation... In addition, a correction term is calculated. The correction term can be obtained by averaging the gradient estimates of multiple sampled trajectories or by using other correction methods. The gradient estimation correction term is determined based on the gradient estimates of the sampled trajectories using the following formula:
[0075] ;
[0076] in, The gradient estimation correction term is denoted by N; N is the number of sampling trajectories. This is the original gradient estimate for the current policy; The dominant function; The gradient of the logarithmic probability of the action; This is for gradient estimation of each sampling trajectory.
[0077] Then, the gradient estimation correction term is weighted and summed with the original gradient estimate of the current policy to obtain the corresponding weighted summation result; the weighted summation result is determined as the target gradient estimate. :
[0078] ;
[0079] Where λ is a weighting coefficient, which can be adjusted according to the actual situation:
[0080] ;
[0081] In one specific embodiment, an averaging method is used to calculate the gradient estimates of multiple sampled trajectories, with T = 5 trajectories sampled. For each trajectory, its gradient estimate is calculated. Then, calculate the correction term according to the correction term formula: Set the weight coefficient λ = 0.1 (in the early stages of training, rely more on the original gradient estimate). According to the final gradient estimation formula: The final gradient estimate is calculated and used to update the policy parameters θ. As training progresses, the value of λ is gradually increased; for example, λ is increased by 0.05 every 10 iterations to better utilize the correction term and reduce gradient estimation bias. In this way, gradient estimation bias can be effectively reduced, improving the accuracy of policy updates. Finally, the current policy is updated using the target gradient estimate and the updated step size factor.
[0082] Step S14: When updating the current strategy, calculate the importance weights, prune the importance weights to obtain the pruned weights, and modify the objective function of the GRPO algorithm according to the pruned weights to obtain the modified function.
[0083] In this embodiment, as the strategy is continuously updated, the data distribution generated by the agent's interaction with the environment also changes. The GRPO algorithm typically assumes a relatively stable data distribution during training, but in reality, rapid changes in data distribution can make the algorithm difficult to adapt. For example, in a game scenario, the agent's initial strategy is to explore the map, at which point the data is mainly exploration data; as the strategy is updated, it gradually shifts to attacking enemies, at which point the data distribution changes from being dominated by exploration data to being dominated by combat data. The algorithm may not be able to adjust in time to adapt to this change, leading to training instability. Therefore, this application employs importance sampling technology to address the problem of data distribution changes. At each policy update, the importance weight w between the current policy and the policy used when generating the data is calculated, i.e. ,in It is the probability that the current policy will take action a in state s. This represents the probability that the strategy used to generate the data will take action 'a' in state 's'. To avoid excessively large importance weights 'w' leading to variance explosion, weight pruning is employed. The objective function is then adjusted based on the pruned weights. The original objective function of GRPO is:
[0084] ;
[0085] The revised version incorporates importance weight pruning. Specifically, the importance weight is determined by the ratio of the probability of taking the target action under the current policy in the target state to the probability of taking the target action under the target state using the policy used when generating the data. This importance weight is then pruned using a preset pruning threshold to obtain the pruned weights. Finally, the objective function of the GRPO algorithm is modified based on these pruned weights using a modified function determination formula. The modified function determination formula is as follows:
[0086] ;
[0087] in, The corrected function; The weights after clipping; The dominant function; For strategy The expectation of taking action a in state s.
[0088] Step S15: Train the GRPO algorithm based on the corrected function and the corresponding updated policy, so that the agent learns the optimal policy based on the corresponding trained GRPO algorithm and determines the exit of the target maze environment according to the optimal policy.
[0089] In this embodiment, a data buffer is established to store data generated by the agent's interaction with the maze environment, including information such as state, action, reward, and next state. Specifically, during each policy update, a certain number of data points are sampled from the buffer according to their importance weights for training. Weighted random sampling can be used to achieve this weighted sampling. Let the buffer contain N data points, and the importance weight of the i-th data point be... The probability that the i-th data item is sampled is:
[0090] ;
[0091] Where N is the number of data entries in the buffer. Each time the strategy is updated, a certain number of data entries are sampled from the buffer according to these probabilities for training. For example, 100 data entries are sampled each time to ensure that the algorithm can still obtain effective training signals when the data distribution changes. In this way, it is guaranteed that the algorithm can still obtain effective training signals when the data distribution changes. When constructing the buffer, Priority Experience Playback (PER) is used, and the priority is determined by TD (Temporal Difference) error or importance weight:
[0092] ;
[0093] in, For TD error, It is a small constant; Priority ratio. Sampling is performed according to priority ratio to avoid outdated data dominating updates.
[0094] Finally, the GRPO algorithm is trained based on the corrected function and the corresponding updated policy, so that the agent learns the optimal policy based on the trained GRPO algorithm. Specifically, the priority of historical data is determined by the sum of the absolute value of the product between the temporal difference error and the pruned weights and the target constant; the historical data is stored in a pre-constructed data buffer according to the priority of the historical data; during the training of the GRPO algorithm, historical data of the target quantity is sampled from the data buffer according to the pruned weights, and the GRPO algorithm is trained based on the historical data of the target quantity. Finally, the exit of the target maze environment is determined according to the optimal policy. It should also be noted that the method of this application is not only applicable to navigation, but also to reinforcement learning tasks such as robot control, autonomous driving, and game AI (Artificial Intelligence), which can significantly improve training stability and convergence speed.
[0095] In summary, this application first determines the KL divergence through the state space and action space, and calculates the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence; the state space includes each position in the target maze environment; the action space includes the movement direction of the agent; the reward obtained by the agent in successfully finding the exit in each iteration is determined, and the step size factor is updated by the average reward change rate determined based on the rewards of several iterations and the average similarity, resulting in the updated step size factor; a gradient estimation correction term is determined based on the gradient estimation of the sampled trajectory, and the target gradient estimate is determined based on the gradient estimation correction term and the original gradient estimate of the current policy. The current policy is updated using the target gradient estimation and the updated step size factor. The sampled trajectory is obtained through trajectory sampling during the interaction between the agent and the target maze environment. When updating the current policy, importance weights are calculated, and these importance weights are pruned to obtain pruned weights. The objective function of the GRPO algorithm is then modified based on these pruned weights to obtain a modified function. The GRPO algorithm is trained based on the modified function and the corresponding updated policy so that the agent can learn the optimal policy based on the trained GRPO algorithm and determine the exit of the target maze environment based on the optimal policy. Therefore, this application effectively solves the problems of excessive policy update magnitude, data distribution changes, and gradient estimation bias in the GRPO algorithm training process by dynamically adjusting the policy update step size factor and utilizing historical data and corrected gradient estimation. Furthermore, during the navigation process, this invention avoids the problem of the agent getting lost in the maze due to excessively large policy updates. It also avoids the problem of the algorithm being unable to adapt to the task requirements at different stages of the agent's navigation due to changes in data distribution, and the problem of gradient estimation bias affecting the agent's ability to learn the correct policy while searching for the exit. This allows the agent to accurately and quickly find the maze exit. Simultaneously, this invention significantly improves the algorithm's training stability, accelerates convergence speed, and enhances its performance and reliability. It also enhances the algorithm's generalization ability in different environments and tasks, providing a more reliable and efficient solution for complex reinforcement learning tasks.
[0096] See Figure 2 As shown, this embodiment of the invention discloses a navigation device based on the GRPO algorithm, comprising:
[0097] The average similarity calculation module 11 is used to determine the KL divergence through the state space and action space, and calculate the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence; the state space includes each position in the target maze environment; the action space includes the movement direction of the agent;
[0098] The step size factor update module 12 is used to determine the reward obtained by the agent in each iteration for successfully finding the exit, and to update the step size factor by the average reward change rate determined based on the rewards of several iterations and the average similarity, so as to obtain the updated step size factor.
[0099] The policy update module 13 is used to determine a gradient estimation correction term based on the gradient estimation of the sampled trajectory, determine a target gradient estimation based on the gradient estimation correction term and the original gradient estimation of the current policy, and update the current policy through the target gradient estimation and the updated step size factor; the sampled trajectory is the trajectory obtained by trajectory sampling during the interaction between the agent and the target maze environment.
[0100] The function correction module 14 is used to calculate the importance weights when updating the current strategy, prune the importance weights to obtain the pruned weights, and correct the objective function of the GRPO algorithm according to the pruned weights to obtain the corrected function.
[0101] The exit determination module 15 is used to train the GRPO algorithm based on the modified function and the corresponding updated policy, so that the agent learns the optimal policy based on the corresponding trained GRPO algorithm and determines the exit of the target maze environment according to the optimal policy.
[0102] As can be seen, this application effectively solves the problems of excessive policy update magnitude, data distribution variation, and gradient estimation bias in the GRPO algorithm training process by dynamically adjusting the policy update step size factor and utilizing historical data and corrective gradient estimation. Therefore, during agent navigation, it avoids the problem of the agent getting lost in the maze due to excessive policy update magnitude, and also avoids the problem of the algorithm being unable to adapt to the task requirements at different stages of the agent's navigation due to data distribution variation, as well as the problem of gradient estimation bias affecting the agent's ability to learn the correct policy in the process of finding the exit. This allows the agent to accurately and quickly find the maze exit. At the same time, this invention significantly improves the algorithm's training stability, accelerates the convergence speed, and enhances the algorithm's performance and reliability. It also enhances the algorithm's generalization ability in different environments and tasks, providing a more reliable and efficient solution for complex reinforcement learning tasks.
[0103] In some specific embodiments, the average similarity calculation module 11 can be used to sample the target maze environment and count the access frequency of states under each strategy; determine the KL divergence using the state space, action space, and access frequency according to the KL divergence calculation formula; the KL divergence calculation formula is:
[0104] ;
[0105] in, Let KL divergence be the value of KL. In strategy The access frequency of the lower state s; S is the state space; A is the action space; Given the current strategy; calculate the average KL divergence of the strategies in the previous several iterations, and determine the average similarity as the average similarity.
[0106] In some specific embodiments, the step size factor update module 12 can be used to determine the average reward change rate based on the rewards of the several iterations using an average reward change rate calculation formula; the average reward change rate calculation formula is:
[0107] ;
[0108] Where R is the average rate of change of reward; m is the number of iterations; This is the reward for the i-th iteration; The reward for the (i-1)th iteration is given. If the average similarity is less than a first threshold or the average reward change rate is greater than a second threshold, the product of the initial step size factor and the first value is determined as the updated step size factor. If the average similarity is greater than or equal to the first threshold or the average reward change rate is less than or equal to the second threshold, the product of the initial step size factor and the second value is determined as the updated step size factor. Wherein, the first value is less than 1, and the second value is greater than 1.
[0109] In some specific embodiments, the policy update module 13 can be used to determine a gradient estimation correction term based on the gradient estimation of the sampled trajectory using a gradient estimation correction term calculation formula; the gradient estimation correction term calculation formula is:
[0110] ;
[0111] in, The gradient estimation correction term is denoted by N; N is the number of sampling trajectories. This is the original gradient estimate for the current policy; The dominant function; The gradient of the logarithmic probability of the action; This is for gradient estimation of each sampling trajectory.
[0112] In some specific embodiments, the policy update module 13 can be used to perform a weighted summation of the gradient estimation correction term and the original gradient estimation of the current policy to obtain a corresponding weighted summation result; and to determine the weighted summation result as the target gradient estimation.
[0113] In some specific embodiments, the function correction module 14 can be used to determine the importance weight by the ratio between the probability of taking the target action under the current policy in the target state and the probability of taking the target action under the target state using the policy used when generating the data; to prune the importance weight by a preset pruning threshold to obtain the pruned weight; and to correct the objective function of the GRPO algorithm according to the pruned weight using the corrected function determination formula; the corrected function determination formula is as follows:
[0114] ;
[0115] in, The corrected function; The weights after clipping; The dominant function; For strategy The expectation of taking action a in state s.
[0116] In some specific embodiments, the apparatus can also be used to determine the priority of historical data based on the sum of the absolute value of the product between the time difference error and the pruned weights and the target constant; store the historical data in a pre-constructed data buffer according to the priority of the historical data; and, during the training of the GRPO algorithm, sample a target number of historical data from the data buffer according to the pruned weights, and train the GRPO algorithm based on the target number of historical data.
[0117] Furthermore, embodiments of this application also disclose an electronic device, Figure 3 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0118] Figure 3 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the navigation method based on the GRPO algorithm disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0119] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0120] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0121] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the GRPO algorithm-based navigation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0122] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned navigation method based on the GRPO algorithm. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0123] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0124] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0125] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0126] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0127] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A navigation method based on the GRPO algorithm, characterized in that, include: The KL divergence is determined by the state space and action space, and the average similarity between the current policy and the previous several iteration policies is calculated based on the KL divergence. The state space includes various locations within the target maze environment; the action space includes the agent's movement directions. The reward obtained by the agent in each iteration for successfully finding the exit is determined, and the updated step size factor is obtained by updating the step size factor based on the average reward change rate determined by the rewards of several iterations and the average similarity. The gradient estimation correction term is determined based on the gradient estimation of the sampled trajectory. The target gradient estimation is determined based on the gradient estimation correction term and the original gradient estimation of the current policy. The current policy is updated using the target gradient estimation and the updated step size factor. The sampling trajectory is the trajectory obtained by trajectory sampling during the interaction between the intelligent agent and the target maze environment; When updating the current strategy, the importance weights are calculated, the importance weights are pruned to obtain the pruned weights, and the objective function of the GRPO algorithm is modified according to the pruned weights to obtain the modified function. The GRPO algorithm is trained based on the modified function and the corresponding updated policy, so that the agent learns the optimal policy based on the trained GRPO algorithm and determines the exit of the target maze environment according to the optimal policy.
2. The navigation method based on the GRPO algorithm according to claim 1, characterized in that, The process of determining the KL divergence through the state space and action space, and calculating the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence, includes: Sampling was performed in the target maze environment to count the access frequency of each state under each strategy; The KL divergence is determined using the state space, action space, and access frequency according to the KL divergence calculation formula; the KL divergence calculation formula is: ; in, Let KL divergence be the value of KL. In strategy The access frequency of the lower state s; S is the state space; A is the action space; This is the current strategy; Calculate the average KL divergence of the strategy in the first few iterations, and determine the average similarity as the average similarity.
3. The navigation method based on the GRPO algorithm according to claim 1, characterized in that, The step-size factor is updated by using the average reward change rate determined based on rewards from several iterations and the average similarity to update the step-size factor, resulting in the updated step-size factor. The average reward change rate is determined based on the rewards from the aforementioned several iterations using the average reward change rate calculation formula; the average reward change rate calculation formula is as follows: ; Where R is the average rate of change of reward; m is the number of iterations; This is the reward for the i-th iteration; This is the reward for the (i-1)th iteration; If the average similarity is less than the first threshold or the average reward change rate is greater than the second threshold, then the product between the initial step size factor and the first value is determined as the updated step size factor. If the average similarity is greater than or equal to the first threshold or the average reward change rate is less than or equal to the second threshold, then the product between the initial step size factor and the second value is determined as the updated step size factor. Wherein, the first value is less than 1; the second value is greater than 1.
4. The navigation method based on the GRPO algorithm according to claim 1, characterized in that, The gradient estimation based on the sampling trajectory determines the gradient estimation correction term, including: The gradient estimation correction term is determined based on the gradient estimation of the sampled trajectory using the gradient estimation correction term calculation formula; the gradient estimation correction term calculation formula is as follows: ; in, The gradient estimation correction term is denoted by N; N is the number of sampling trajectories. This is the original gradient estimate for the current policy; The dominant function; The gradient of the logarithmic probability of the action; This is for gradient estimation of each sampling trajectory.
5. The navigation method based on the GRPO algorithm according to claim 1, characterized in that, The step of determining the target gradient estimate based on the gradient estimation correction term and the original gradient estimate of the current policy includes: The gradient estimation correction term is weighted and summed with the original gradient estimate of the current policy to obtain the corresponding weighted summation result; The weighted summation result is determined as the target gradient estimate.
6. The navigation method based on the GRPO algorithm according to claim 1, characterized in that, The calculation of importance weights, and the pruning of these importance weights to obtain pruned weights, includes: The importance weight is determined by the ratio between the probability of taking the target action under the current policy in the target state and the probability of taking the target action under the target state using the policy used when generating the data. The importance weights are pruned by setting a preset pruning threshold to obtain the pruned weights; Accordingly, the step of modifying the objective function of the GRPO algorithm based on the pruned weights to obtain the modified function includes: The objective function of the GRPO algorithm is modified based on the pruned weights using the modified function determination formula; the modified function determination formula is as follows: ; in, The corrected function; The weights after clipping; The dominant function; For strategy The expectation of taking action a in state s.
7. The navigation method based on the GRPO algorithm according to any one of claims 1 to 6, characterized in that, Also includes: The priority of historical data is determined by the sum of the absolute value of the product between the time difference error and the pruned weights and the target constant. The historical data is stored in a pre-built data buffer according to its priority. During the training of the GRPO algorithm, historical data on the number of targets are sampled from the data buffer according to the cropped weights, and the GRPO algorithm is trained based on the historical data on the number of targets.
8. A navigation device based on the GRPO algorithm, characterized in that, include: The average similarity calculation module is used to determine the KL divergence through the state space and action space, and calculate the average similarity between the current policy and the policies of the previous several iterations based on the KL divergence; the state space includes each position in the target maze environment; the action space includes the movement direction of the agent; The step size factor update module is used to determine the reward obtained by the agent in each iteration for successfully finding the exit. The step size factor is updated by the average reward change rate determined based on the rewards of several iterations and the average similarity. The policy update module is used to determine a gradient estimation correction term based on the gradient estimation of the sampled trajectory, determine a target gradient estimation based on the gradient estimation correction term and the original gradient estimation of the current policy, and update the current policy using the target gradient estimation and the updated step size factor. The sampling trajectory is the trajectory obtained by trajectory sampling during the interaction between the intelligent agent and the target maze environment; The function correction module is used to calculate importance weights when updating the current strategy, prune the importance weights to obtain pruned weights, and correct the objective function of the GRPO algorithm based on the pruned weights to obtain the corrected function. The exit determination module is used to train the GRPO algorithm based on the modified function and the corresponding updated policy, so that the agent learns the optimal policy based on the corresponding trained GRPO algorithm and determines the exit of the target maze environment according to the optimal policy.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing a computer program to implement the steps of the navigation method based on the GRPO algorithm as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, A computer program is stored on a computer-readable storage medium, which, when executed by a processor, implements the steps of the navigation method based on the GRPO algorithm as described in any one of claims 1 to 7.