Mobile robot path planning method based on IDCAC
By combining DRL and Deep Diffusion Contrastive Architecture to generate virtual data and introduce intrinsic rewards, the problem of unstable Agent decision-making in complex environments is solved, and more efficient path planning is achieved.
Patent Information
- Application Number
- CN202510483466.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-29
AI Technical Summary
In complex environments, existing path planning algorithms have rapidly increased data demand due to high-dimensional state and action space, making it difficult for Agents to obtain complete environmental information, affecting decision-making effects, and traditional dimensionality reduction techniques lead to data loss and increase learning difficulty.
Combining DRL training and Deep Diffusion Contrastive Architecture, virtual state action pairs are generated and cross-match prediction rewards are performed. Collaborative navigation architecture is built through auxiliary learning and strategy learning, and Agent's own line speed and angular speed are introduced as intrinsic rewards to optimize path planning.
Improve the decision stability and adaptability of Agent in complex environments, reduce dependence on a single environment, enhance learning ability in high-dimensional space, and achieve more efficient path planning.
Smart Images

Figure CN120386353A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of path planning, and particularly to a mobile robot path planning method based on IDCAC. Background Art
[0002] SLAM can perform environment construction and target positioning navigation in an unknown environment; however, the efficiency of SLAM is affected by various factors such as sensor accuracy and environmental dynamic changes; introducing the DRL algorithm can further improve the navigation efficiency and promote the solution of the path planning problem; the application of DRL in robot autonomous navigation allows the Agent to find the optimal or sub-optimal navigation strategy by directly interacting with the environment; but as the environment becomes more and more complex and the dimensions of the state and action spaces continue to increase, the Curse of Dimensionality will lead to a rapid growth in data requirements.
[0003] Mobile robot navigation mainly consists of point-to-point P2P motion and obstacle avoidance. The P2P task aims to move efficiently and safely; obstacles can include static and dynamic types; SAC introduces entropy regularization on the basis of TD3 to encourage the Agent to explore more random actions; but the coexistence of complex P2P tasks and various obstacles increases the dimension of the environment, making the state and action spaces more and more complex; common dimensionality reduction techniques will inevitably cause data loss while reducing the dimension, and the partially observable nature makes the Agent unable to obtain complete environmental information during decision-making, further increasing the learning difficulty and making it more difficult for the Agent to make the optimal decision. Summary of the Invention
[0004] Aiming at the deficiencies of the existing methods, the present invention seamlessly integrates DRL training with Deep Diffusion Contrastive Architecture, which is called collaborative navigation training; virtual state-action pairs are generated by Deep Diffusion Contrastive Architecture in the auxiliary training, and cross-matching is performed on them to predict the reward result as diverse simulation data; contrastive learning is respectively performed on the real data and the simulation data, and the contrastive learning features are embedded into the policy learning as a kind of "enhanced encoding"; in addition, an intrinsic reward element (linear velocity and angular velocity) including the structural properties of the Agent itself is introduced into the traditional goal-driven reward function.
[0005] The technical solution adopted by the present invention is: a mobile robot path planning method based on IDCAC includes the following steps:
[0006] Step 1, construct an MDP model of reinforcement learning RL;
[0007] Step 2, use the Agent to obtain the linear velocity, angular velocity, and obstacle distribution;
[0008] Step 3: Set the linear velocity reward value and angular velocity reward value of the Agent respectively, and use the maximum value and the current value of the linear velocity to suppress the deviation of the Agent; use the maximum value and the current value of the angular velocity to suppress the angular velocity deviation of the Agent;
[0009] As a preferred embodiment of the present invention, the formula for the linear velocity reward value is:
[0010]
[0011] where, ɑ1 is the correction coefficient; ν max is the maximum linear velocity; ν t is the current linear velocity.
[0012] As a preferred embodiment of the present invention, the formula for the angular velocity reward value is:
[0013]
[0014] where, ɑ2 is the correction coefficient; ω max is the maximum angular velocity; ω t is the current angular velocity.
[0015] Step 4: Set the external reward value; and establish a hybrid reward mechanism by combining the external reward value with the linear velocity and angular velocity reward values;
[0016] As a preferred embodiment of the present invention, the formula for the hybrid reward value is:
[0017] reward = w ν *r ν +w ω *r ω +w dir *r dir +w dis *r dis +w obs *r obs (7)
[0018] where, w ν 、w ω 、w dir 、w dis 、w obs are the weighted coefficients of linear velocity reward, angular velocity reward, direction reward, distance reward and obstacle reward respectively.
[0019] Step 5: Construct a cooperative navigation architecture based on auxiliary learning and policy learning; use the DDPM model and contrastive learning to generate virtual data and perform cross-matching reward prediction on it to simulate diverse data, and embed the contrastive learning features into the policy learning as auxiliary training;
[0020] As a preferred embodiment of the present invention, the in - data feature extraction for auxiliary learning includes:
[0021] First, extract N episodes from the RBuffer of policy learning, and sort the data in each episode according to the reward value;
[0022] Secondly, use the data s r_h -a r_h corresponding to each maximum reward value as an anchor point, and the data s r_l -a r_l corresponding to the minimum reward value as a negative sample, and randomly extract s r_m -a r_m as a positive sample; perform transformation on the selected data to obtain Anchor, Positive, and Negative, which are respectively represented as v θ (s r_h -a r_h ), and
[0023] Finally, maximize the cosine similarity between Anchor and Positive, and minimize the cosine similarity between Anchor and Negative through contrastive learning.
[0024] As a preferred embodiment of the present invention, the formula for cosine similarity is:
[0025]
[0026] where Neg csi represents the cosine similarity between the i - th sample Anchor and Negative, Pos csi represents the cosine similarity between the i - th sample Anchor and Positive, and margin represents the similarity difference between positive and negative samples.
[0027] As a preferred embodiment of the present invention, the cross - data feature prediction for auxiliary learning includes:
[0028] First, take a randomly generated sample x0 as the starting point, and calculate the noisy sample x t at each time step t;
[0029] Secondly, through the reverse diffusion loop, gradually denoise x T ;
[0030] Subsequently, separate x T into the simulated state s sim and the simulated action a sim; Randomly select a pair of s sim and a sim and merge them into the array s sim -a sim as the input of DDPM; Divide the prediction result into the reward r sim and the next state s' sim , and save them in the VBuffer;
[0031] Finally, extract data from the RBuffer to maximize the similarity Pos between the anchor point and the positive sample csi and minimize the similarity Neg between the anchor point and the negative sample csi .
[0032] As a preferred embodiment of the present invention, the noisy sample is calculated for x0 using the cumulative scaling factor and random sampling noise.
[0033] As a preferred embodiment of the present invention, the mobile robot path planning system based on IDCAC includes: a memory for storing instructions executable by a processor; a processor for executing the instructions to implement the mobile robot path planning method based on IDCAC.
[0034] As a preferred embodiment of the present invention, a computer-readable medium storing computer program code, the computer program code implementing the mobile robot path planning method based on IDCAC when executed by a processor.
[0035] Advantages of the present invention:
[0036] 1. Incorporate auxiliary training based on the DRL algorithm, including policy learning and auxiliary training. In the auxiliary training part, DDPM and contrastive learning are adopted to alleviate problems such as data sparsity and low sample quality of the Agent in complex dynamic environments;
[0037] 2. The structure of the Agent itself determines how it perceives the environment and how to process information to make decisions. Different Agent structures may lead to different learning results and strategies; Incorporate the Agent structure information used as an intrinsic reward element into the traditional goal-driven reward function to achieve effective policy learning;
[0038] 3. The present invention simulates multiple environments with different complexities for the high-dimensional complex space. On this basis, large-scale algorithm comparison experiments are carried out to evaluate the navigation performance of the latest deep reinforcement learning algorithms including SAC and TD3 in environments with different complexities. Description of the Drawings
[0039] Figure 1It is the flowchart of the mobile robot path planning method based on IDCAC (Intrinsic and Diffusion Contrastive Architectures SAC) of the present invention;
[0040] Figure 2 It is the perspective view of 3 angles at the top of the lidar of the Agent;
[0041] Figure 3 It is the collaborative navigation architecture diagram;
[0042] Figure 4 It is the description of the experimental environment;
[0043] Figure 5 It is the comparison chart of rewards in the static environment;
[0044] Figure 6 It is the comparison chart of success rates in the static environment;
[0045] Figure 7 It is the comparison chart of rewards in the dynamic environment;
[0046] Figure 8 It is the comparison chart of success rates in the dynamic environment;
[0047] Figure 9 It is the comparison chart of rewards in the static and dynamic environments;
[0048] Figure 10 It is the comparison chart of success rates in the static and dynamic environments. Detailed implementation manners
[0049] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. This figure is a simplified schematic diagram, which only illustrates the basic structure of the present invention in a schematic manner. Therefore, it only shows the components related to the present invention.
[0050] As Figure 1 shown, a mobile robot path planning method based on IDCAC includes the following steps:
[0051] Step 1: Construct an MDP model of reinforcement learning RL;
[0052] Reinforcement learning RL focuses on the interaction between the agent and the environment to obtain the maximum reward return; the agent learns through trial and error, continuously optimizes the strategy according to the environmental feedback, and learns to select the best action in different environmental states.
[0053] RL has the Markov property: for a given state s t and action a t , the next state s t+1 is only determined by the current s t and a tDecisions are therefore usually made using the Markov Decision Process (MDP) as its model framework;
[0054] The MDP model consists of a five-tuple ; where is a finite set, and each element represents a state; is a finite set, and each element represents an action; is the state transition function, representing the probability of reaching state s t by executing action a t ; t+1 The probability; represents the reward function, representing the immediate reward obtained by executing action a in state s t ; the discount factor γ ∈ [0, 1]; the policy π can be regarded as the probability mapping of executing different actions a for the current state s t ; t The probability of executing action a in state s; each execution of the policy accumulates rewards in the direct interaction with the environment, forming a return ; The goal of RL is to find an optimal policy π such that the Agent can still obtain the maximum expected return without being affected by the initial state; *
[0055]
[0056] Step 2: Use the Agent to obtain the linear velocity, angular velocity, and obstacle distribution;
[0057] In a complex and changing environment, whether the Agent can safely and effectively avoid obstacles and reach the target point becomes a major challenge; to improve the Agent's environmental perception and path planning capabilities, Turtlebot3 with LiDAR can be used as the Agent model;
[0058] The LiDAR sensor first performs a complete 360° scan to provide the Agent with comprehensive environmental perception; as Figure 2 shown, the LiDAR collects data at 0°, 30°, and 330° respectively to detect whether there are obstacles at that angle, and decides its behavior based on the distance data in the front and on the sides; if there is no obstacle in the front, the robot moves straight ahead; when encountering an obstacle on the left, it turns right, and turns left when encountering an obstacle on the right; if the distance in the front is less than the set safety distance, it immediately turns right to avoid collision; given that the physical structure of the Agent will affect its behavior ability and learning process to a certain extent;
[0059] The present invention incorporates the linear velocity and angular velocity of the Agent into the reward and re - sets the reward mechanism;
[0060] Step 3: Set the linear velocity reward value and angular velocity reward value of the Agent respectively. Use the maximum value and the current value of the linear velocity to suppress the deviation of the Agent; use the maximum value and the current value of the angular velocity to suppress the angular velocity deviation of the Agent;
[0061] The motion state of the Agent, as an incentive factor, affects the decision - making and learning process of the Agent to a certain extent. Given the linear velocity ν and angular velocity ω, as important motion parameters of the robot, they are used as a quasi - intrinsic reward (a reward signal between extrinsic reward and intrinsic reward) to encourage the Agent to optimize its motion pattern;
[0062] The formula for the linear velocity reward value is:
[0063]
[0064] where ɑ1 is a correction coefficient; ν max is the maximum linear velocity, ν max = 0.22m / s; ν t is the current linear velocity.
[0065] First, calculate the deviation between ν t and ν max . By imposing a penalty on the behavior of deviating from ν max , suppress the behavior of deviating significantly from ν t , and encourage the robot to maintain an appropriate linear velocity to learn a stable and efficient motion pattern. max
[0066] The formula for the angular velocity reward value is:
[0067]
[0068] where ɑ2 is a correction coefficient; ω max is the maximum angular velocity, ω max = 2.0rad / s; ω t is the current angular velocity.
[0069] Formula (3) is a negative quadratic function. By squaring ω t , ensure the processing of positive and negative angles, suppress excessive angular velocity, improve the stability of the Agent's rotation, and enhance its adaptability to angular velocities of different scales.
[0070] The quasi-intrinsic reward is designed based on the physical structure and characteristics of the Agent itself. This structured intrinsic reward combines with DRL to form a feedback loop. As the robot learns and adapts, it can better utilize its own structural characteristics and at the same time have a positive promoting effect on the DRL algorithm.
[0071] Step 4: Set the extrinsic reward value; and establish a hybrid reward mechanism by combining the extrinsic reward value with the linear velocity and angular velocity reward values.
[0072] Extrinsic rewards are the most direct and commonly used reward type in DRL. Compared with intrinsic rewards, extrinsic rewards can directly reflect the completion degree of the target task. For example, in the autonomous driving task, the extrinsic reward is directly related to whether the Agent successfully reaches the destination. If it safely reaches the destination, a large amount of positive rewards will be given; on the contrary, if a collision occurs or it fails to arrive in time, negative rewards will be given. The main task of the Agent in this invention is to avoid obstacles and perform path planning in a dynamic environment to successfully reach the target point. Therefore, the extrinsic reward mainly consists of direction reward, distance reward, and obstacle reward.
[0073] The formula for the direction reward value is:
[0074]
[0075] where α t represents the current direction angle, α g represents the target point angle, and ɑ3 is the correction coefficient. By taking the absolute value of the difference, it ensures that the same-sized result will be generated regardless of the direction, thus avoiding unstable behaviors caused by changes in the angle sign.
[0076] The formula for the distance reward value is:
[0077]
[0078] where ɑ3, d ini represents the distance between the Agent and the target point at the initial moment, and d t represents the distance between the Agent and the target point at the current moment. This distance scale can effectively measure the change of the current distance d t relative to the initial distance d ini : When d t < d ini , r dis will be a positive value, indicating the reward for the Agent approaching the target point, thus encouraging the Agent to get closer to the target point; when d t > d ini , r dis will be a negative value, indicating a penalty.
[0079] The formula for the obstacle reward value is as follows:
[0080]
[0081] where d obs represents the distance between the Agent and the obstacle. Once the Agent gets too close to the obstacle, a penalty will be given to encourage the Agent to learn effective obstacle avoidance methods.
[0082] The formula for the mixed reward value Endogenous Incentives is as follows:
[0083] reward = w ν *r ν +w ω *r ω +w dir *r dir +w dis *r dis +w obs *r obs (7)
[0084] where w ν , w ω , w dir , w dis , w obs are the weighted coefficients of linear velocity reward, angular velocity reward, direction reward, distance reward, and obstacle reward respectively, and w ν +w ω +w dir +w dis +w obs = 1.
[0085] This invention focuses on obstacle reward and distance reward, encouraging the Agent to quickly learn the strategy to reach the target point; at the same time, considering the influence of the Agent's own speed and direction on the strategy, small weights are assigned to the linear / angular velocity and direction rewards, which can flexibly adjust the focus of the Agent training module and improve the adaptability and generalization ability of the Agent.
[0086] Based on the introduction of quasi-intrinsic reward elements into the extrinsic reward, the total reward is obtained by decentralizing and summing each reward module; firstly, this mixed reward combines the main task of the Agent and its own structural characteristics, emphasizing their correlation, which can effectively reduce unnecessary behavior deviations and optimize each function of the Agent, achieving the consistency of behavior and reward; secondly, the rewards are subdivided according to different functions and finally different weights are assigned for summation.
[0087] Step 5: Build a Collaborative Navigation Architecture based on Auxiliary Learning and Policy Learning; use the DDPM model and contrastive learning to generate virtual data and perform cross-matching reward prediction to simulate diverse data, and embed the contrastive learning features into policy learning as auxiliary training;
[0088] Flexible Decision Making is one of the important concepts in artificial intelligence, requiring the Agent to be able to adapt to environmental changes and maintain the stability of decision-making behavior; however, as the environmental complexity increases, it will lead to the Curse of Dimensionality; traditional dimensionality reduction methods will cause data loss, and the limitations of the Agent's sensors and the Partially Observable Environment caused by the inability to access the complete observations of all time steps further reduce the available data for the Agent; this will lead to the Agent's misunderstanding of the environment and affect the effectiveness and stability of its decision-making.
[0089] The present invention proposes a Deep Diffusion Contrastive Architecture, which combines with the DRL algorithm to achieve collaborative navigation training; as Figure 3 shown in the structure diagram of the Collaborative Navigation Architecture;
[0090] Auxiliary Learning learns feature extraction from the data RBuffer collected from multiple policy updates and the diverse simulation data VBuffer, and embeds it into Policy Learning; Auxiliary Learning consists of two key parts: Intra-Data feature extraction and Cross-Data feature prediction;
[0091] Intra-Data feature extraction aims to maximize the feature space similarity between the state-action pairs corresponding to the highest reward and the lowest reward through contrastive learning;
[0092] Specifically, first extract N episodes from the RBuffer (data storing multiple policy updates), sort the data in each episode according to the reward value, and the reward value can be the mixed reward value of formula (7); for the data s corresponding to each maximum reward valuer_h -a r_h As an anchor point, the minimum reward value s r_l -a r_l As a negative sample, randomly select s r_m -a r_m As a positive sample; the selected data is transformed through a preprocessing function to obtain the final Anchor, Positive, and Negative, which are respectively represented as v θ (s r_h -a r_h ), and Finally, through contrastive learning, maximize the cosine similarity between Anchor and Positive, while minimizing the cosine similarity between Anchor and Negative:
[0093]
[0094] where, Neg csi represents the cosine similarity between the i-th sample Anchor and Negative, and Pos csi represents the cosine similarity between the i-th sample Anchor and Positive, and margin represents the similarity difference between positive and negative samples;
[0095] When and only when the difference between Neg csi and Pos csi is as large as possible, the loss will increase, so as to improve the discriminative ability of the model learning.
[0096] Cross-Data feature prediction focuses on the prediction and simulation of virtual data; this simulated data is more diverse than the data obtained by policy learning, alleviating the environmental understanding deviation caused by insufficient data while reducing the model's dependence on a specific environment and enhancing the model's generalization ability; the present invention uses DDPM for data prediction;
[0097] First, use the randomly generated sample x0 as the starting point. In each subsequent time step t, calculate the noisy sample x t according to formula (9), and continuously perform this process until the time step reaches the maximum value T, and finally obtain the sample x T .
[0098] x t The formula for is:
[0099]
[0100] where, α bars[t] As the cumulative scaling factor, it represents the proportion of the original signal retained after t time steps; ∈ represents the noise randomly sampled from the standard normal distribution.
[0101] Then, through the reverse diffusion cycle, gradually denoise the initial noise x T to obtain a clear sample; it should be noted that when t > 0, it means that the current has not reached the last time step, and at this time, a certain proportion of random noise still needs to be added to prevent overfitting and generation instability in the subsequent denoising process. When t == 0, it means that the denoising is completed, and the sample obtained at this time is the final output x T .
[0102] Subsequently, separate x T into the simulated state s sim and the simulated action a sim ; randomly extract a pair of s sim and a sim from the data set VBuffer, and combine these two values into a single array s sim -a sim , as the input of the Deep Diffusion Probability Model (DDPM); separate the predicted result of the output, divide it into the reward r sim and the next state s' sim , and save this data to VBuffer for contrastive learning;
[0103] Contrastive learning is to extract data from RBuffer, maximize the similarity Pos csi between the anchor point and the positive sample, and minimize the similarity Neg csi between the anchor point and the negative sample; through the formula, optimize the feature representation of the high-reward state-action pair and the low-reward state-action pair by the model to improve the model performance.
[0104] The Collaborative Navigation Architecture proposed by the present invention consists of Policy Learning and Auxiliary Learning; Auxiliary Learning mainly uses the Deep Diffusion Contrastive Architecture, which can predict the reward of the cross-matched s sim _a sim while generating virtual data and update and improve VBuffer; mix the simulated data and the real data as the diverse data supplement of RBuffer; in addition, the present invention splices and fuses the corresponding embedding vectors obtained by the contrastive learning model in Auxiliary Learning with the original state-action pair to obtain st-emb and a t-emb , as the input of the Enhanced Encoder to more efficiently utilize feature information; these two fusion methods are jointly used for Policy Learning, so that SAC can more quickly adapt to complex dynamic environments and learn effective policies.
[0105] The IDCAC of the present invention consists of two learning processes, Policy Learning and Auxiliary Learning, which is an innovative integration of Endogenous Incentives and Deep Diffusion Contrastive Architecture with SAC; first, linear velocity and angular velocity are introduced on the basis of traditional externally driven rewards with goals; in the auxiliary learning process, the R(V)Buffer data is analyzed by comparison and the contrast learning features are embedded into the Actor and Critic of SAC to optimize the decision-making of the Agent.
[0106] The specific process is as follows:
[0107] 1. Initialization
[0108] 11. Replay Buffers;
[0109] RB (real experience replay pool): stores data (s, a, r, s') collected from the real environment;
[0110] VB (virtual experience replay pool): stores virtual data (s sim , a sim , r pre , s' pre ) generated by the diffusion model;
[0111] 12. Q-function:
[0112] Two Q functions and are used to estimate state-action values;
[0113] 13. Policy:
[0114] Policy π φ , used to select and execute actions;
[0115] 14. Target Networks:
[0116] Target networks θ1' and θ2', used to improve the stability of training;
[0117] 2. Outer loop: Each iteration represents a complete training cycle;
[0118] 3. Environmental Interaction:
[0119] 31. Select action a and execute it according to the current policy π φ (a|s), observe the reward r and the new state s';
[0120] 32. Generate virtual data using a deep diffusion model: Simulate the state s sim , simulate the action a sim , predict the reward r pre and the next state s' pre ;
[0121] 33. Store data:
[0122] · Store the real data in the real data replay pool RB;
[0123] · Store the generated virtual data in the virtual data replay pool VB;
[0124] 4. Loss Optimization:
[0125] 41. Optimize the Critic loss function using the data in the real data replay pool RB;
[0126] 42. Optimize the loss function using the data in the virtual data replay pool VB
[0127] 5. Inner Loop:
[0128] 51. Train the Q function using the Critic loss function and the number of gradient steps G Q ;
[0129] 52. Optimize using the Actor loss function and the number of gradient steps π φ ;
[0130] 53. Update the target network parameters θ1' and θ2' through soft update, where τ is the update coefficient;
[0131] The pseudocode is as follows:
[0132]
[0133] Experimental Environment and Configuration:
[0134] Conducted on the Windows 11 system, in the Pytorch 1.8 environment and the Gazebo simulation platform, using an i9-13900HX processor and 16GB of memory; Gazebo is a physical simulation platform model that can be used to simulate various robots, sensors, and environments; the robot model parameters remain consistent in all experiments without prior environmental knowledge; the experiments were carried out in static, dynamic, and static-dynamic environments respectively, using SAC, TD3, DDPG, and the IDCAC of the present invention to train for 2000 episodes and evaluate for 100 rounds, comparing their average reward values and obstacle avoidance success rates. The detailed parameters of the IDCAC model are referred to in Table 1.
[0135] Table 1 IDCAC Hyperparameters
[0136]
[0137]
[0138] As Figure 4 In the six environments (E1 - E6), the circles represent dynamic obstacles; E1 and E2 are pure static environments, and E2 has more static obstacles than E1; E3 and E4 are pure dynamic environments, and E4 has two more dynamic obstacles than E3; finally, E5 and E6 are static-dynamic combined environments, and E6 is the most complex, containing a large number of static obstacles and 6 dynamic obstacles; the present invention divides these six environments into three groups: the static group (E1, E2), the dynamic group (E3, E4), and the static-dynamic combined group (E5, E6), with each group containing a simple environment and a relatively complex environment; 4.3, 4.4, and 4.5 will analyze the experimental results of these three groups respectively.
[0139] Figure 5 Shows the reward results of the E1 (R1) and E2 (R2) environments. In the simple E1 environment, the performance of IDCAC is inferior to TD3 because the Deep Diffusion Contrastive Architecture in IDCAC increases diverse feature learning, and this additional data has little value for utilization in a simple environment like E1; however, on the contrary, in the complex E2 environment, the average reward of IDCAC is the highest, approximately 1500. Figure 6 Represents the success rate in the static environment; X1 represents the success rate of the experiment in the E1 environment, and X2 represents the success rate of the experiment in the E2 environment. Whether in X1 or X2, IDCAC clearly has the highest success rate compared to the other three algorithms, but in X2, the average value of IDCAC is higher than that in X1. Generally speaking, IDCAC performs better in terms of both average reward and navigation success rate in the more complex static environment E2 than in the simple environment E1.
[0140] Dynamic environments such as Figure 4E3 and E4 among them, both of which only contain dynamic columns, and the difference lies in the number of columns; Figure 7 and Figure 8 respectively show the average rewards and success rates of four algorithms in the dynamic environments E3 and E4; it can be seen that IDCAC outperforms other algorithms in both environments, but in the relatively more complex E4, the supplementation of diverse simulation data plays a greater advantage, and the gap between IDCAC and other algorithms is more significant in both average rewards and success rates.
[0141] The dynamic and static environment combines static and dynamic obstacles, corresponding to Figure 4 E5 and E6 in Figure 9 ; as shown in Figure 10 represents the comparison chart of the success rates of four algorithms in two dynamic and static environments; in the most complex environment E6, IDCAC fully demonstrates the advantage of supplementing diverse data and feature learning by embedding DeepDiffusion Contrastive Architecture as an auxiliary training into Policy Leaning; whether it is Figure 9 the average reward in Figure 10 or the success rate in
[0142] Table 2 shows the average rewards of 100 evaluation rewards of IDCAC, SAC, TD3, and DDPG algorithms in six environments; E1 and E2 are divided into the static group; E3 and E4 are divided into the dynamic group; E5 and E6 are divided into the static-dynamic combination group; E1, E3, and E5 respectively represent the relatively simple environments in each group; in these three environments, IDCAC does not show much advantage, and the average reward of IDCAC is the second highest among the four algorithms; on the contrary, in the complex environments E2, E4, and E6 of the three groups, IDCAC has the highest average reward, and as the environment becomes more and more complex, the gap between it and the other three algorithms becomes larger and larger; in E2, IDCAC is 247 higher than the worst-performing SAC; secondly, it is 227 higher than DDPG and 41 higher than TD3; in E4, IDCAC is 1204 higher than SAC, 181 higher than DDPG, and 148 higher than the relatively better TD3; in the most complex environment E6, the average reward of IDCAC is 1733 higher than the worst-performing SAC, 515 higher than DDPG, and 67 higher than TD3; IDCAC mainly makes an innovative embedding of Deep Diffusion Contrastive Architecture and SAC, alleviates the data loss problem brought by the high-dimensional space through auxiliary contrastive learning, generates diverse data at the same time, improves the Agent's ability to make flexible decisions in complex environments, reduces the dependence on a single environment, and improves the Agent's adaptability to new environments.
[0143] Average Rewards of Four Algorithms in Six Environments in Table 2
[0144]
[0145]
[0146] The present invention aims at the path planning problem of mobile robots in complex environments, and introduces an auxiliary training called Deep Diffusion Contrastive Architecture to supplement diverse data; this method first gradually denoises randomly initialized noise to generate virtual data, simulates diverse data through cross-matching state-action pairs, and finally integrates this contrastive learning feature into the policy update;
[0147] The present invention introduces the linear and angular velocities of the Agent with intrinsic reward elements into the extrinsic reward function driven by the goal, and tests SAC, TD3, DDPG, and IDCAC in six different difficulty environments. The experimental results show that IDCAC has more decision-making advantages in more complex environments by using diverse data.
[0148] Based on the above-mentioned ideal embodiments of the present invention as inspiration, through the above description, relevant staff can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.
Claims
1. A mobile robot path planning method based on IDCAC, characterized in that, It includes the following steps: Step 1, construct the MDP model of reinforcement learning RL; Step 2, use the Agent to obtain the linear velocity, angular velocity, and obstacle distribution; Step 3, set the linear velocity reward value and angular velocity reward value of the Agent, and use the maximum value and the current value of the linear velocity to suppress the deviation of the Agent; use the maximum value and the current value of the angular velocity to suppress the angular velocity deviation of the Agent; Step 4, set the extrinsic reward value; and establish a hybrid reward mechanism by combining the extrinsic reward value with the linear velocity and angular velocity reward values; Step 5, construct a collaborative navigation architecture based on auxiliary learning and policy learning; use the DDPM model and contrastive learning to generate virtual data and perform cross-matching reward prediction to simulate diverse data, and embed the contrastive learning features into the policy learning as auxiliary training.
2. The method for path planning of a mobile robot based on IDCAC according to claim 1, wherein The formula for the hybrid reward value is: reward = w ν *r ν +w ω *r ω +w dir *r dir +w dis *r dis +w obs *r obs (1) Among them, w ν , w ω , w dir , w dis , w obs are the weighted coefficients of linear velocity reward, angular velocity reward, direction reward, distance reward, and obstacle reward respectively.
3. The method for path planning of a mobile robot based on IDCAC according to claim 2, wherein The formula for the linear velocity reward value is: Among them, ɑ1 is the correction coefficient; ν max is the maximum linear velocity; ν t is the current linear velocity.
4. The method for path planning of a mobile robot based on IDCAC according to claim 2, characterized in that, The formula for the angular velocity reward value is: Among them, ɑ2 is the correction coefficient; ω max is the maximum angular velocity; ω t is the current angular velocity.
5. The method for path planning of a mobile robot based on IDCAC according to claim 1, wherein The extraction of in-data features for auxiliary learning includes: First, extract N episodes from the RBuffer of policy learning, and sort the data in each episode according to the reward value; Secondly, the data s corresponding to each maximum reward value r_h -a r_h is used as an anchor point, and the minimum reward value s r_l -a r_l is used as a negative sample, and s r_m -a r_m is randomly selected as a positive sample; the selected data is transformed to obtain Anchor, Positive, and Negative, which are respectively represented as v θ (s r_h -a r_h ), and Finally, maximize the cosine similarity between Anchor and Positive and minimize the cosine similarity between Anchor and Negative through contrastive learning.
6. The method for path planning of a mobile robot based on IDCAC according to claim 5, wherein The formula for the cosine similarity is: Among them, Neg csi represents the cosine similarity between the i-th sample Anchor and Negative, and Pos csi represents the cosine similarity between the i-th sample Anchor and Positive, and margin represents the similarity difference between positive and negative samples.
7. The method for path planning of a mobile robot based on IDCAC according to claim 5, wherein The cross-data feature prediction for auxiliary learning includes: First, take the randomly generated sample x0 as the starting point, and calculate the noisy sample x according to each time step t t ; Secondly, through the reverse diffusion loop, gradually denoise x T ; Subsequently, separate x T into the simulated state s sim and the simulated action a sim ; randomly select a pair of s sim and a sim , combine them into the array s sim -a sim as the input of DDPM; divide the prediction result into the reward r sim and the next state s' sim , and save them to the VBuffer; Finally, extract data from the RBuffer to maximize the similarity Pos between the anchor point and the positive sample csi , and minimize the similarity Neg between the anchor point and the negative sample csi .
8. The method for path planning of a mobile robot based on IDCAC according to claim 7, characterized in that, The noisy samples are calculated for x0 using the cumulative scaling factor and random sampling noise.
9. Mobile robot path planning system based on IDCAC, characterized in that, It includes: A memory for storing instructions executable by a processor; A processor for executing the instructions to implement the IDCAC-based mobile robot path planning method according to any one of claims 1-8.
10. A computer-readable medium storing computer program code, characterized in that, The computer program code implements the IDCAC-based mobile robot path planning method according to any one of claims 1-8 when executed by the processor.