Robot bidirectional motion system and method based on mirror image enhanced experience playback
Through mirroring, the robot's two-way motion system is enhanced by experience playback, and the reverse motion trajectory is generated and parameter optimization is performed, which solves the problem that the robot cannot retreat in complex scenarios, realizes the robot's flexible navigation in narrow spaces, and improves navigation capabilities and adaptability to autonomous navigation.
Patent Information
- Application Number
- CN202510379860.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-15
AI Technical Summary
Existing robot path planning methods cannot effectively deal with challenges in dynamic obstacles or narrow spaces when facing complex, dynamic or narrow scenarios, resulting in robots being trapped without flexibly performing backward operations.
A robot bidirectional motion system based on mirror-enhanced experience playback is adopted. Through the combination of motion control module, mirror-enhanced module and training module, reverse motion trajectory is generated and parameter optimization is performed. The flexible actor-critician algorithm is used to train the robot's bidirectional motion ability.
The robot shows excellent navigation capabilities in various scenarios, can flexibly retreat, avoid being trapped, improves navigation capabilities in changing scenarios, and broadens the application boundaries of deep reinforcement learning in complex scenarios.
Smart Images

Figure CN120491623A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robot control technology, and in particular to a robot bidirectional motion system and method based on mirror-enhanced experience playback. Background Art
[0002] In robotic navigation technology, deep reinforcement learning (DRL)-based methods have demonstrated excellent path planning and navigation capabilities in static scenes. However, existing deep reinforcement learning methods still have significant drawbacks when faced with complex, dynamic, or narrow scenes.
[0003] Traditional deep reinforcement learning methods often favor forward motion when learning navigation strategies, which prevents robots from responding effectively when they encounter situations where they need to retreat. For example, in confined spaces or complex obstacle scenarios, robots may become stuck because they are unable to perform a backward maneuver.
[0004] Although traditional path planning methods such as the dynamic window approach (DWA) and timed elastic band (TEB) consider the complete action space including forward and backward movement, these methods perform well in simple scenarios. However, they lack flexibility and adaptability when faced with complex and changing scenarios, and therefore cannot effectively cope with challenges in dynamic obstacles or narrow spaces. Summary of the Invention
[0005] The technical problem to be solved by the present invention is how to overcome the technical shortcomings of existing robotic path planning methods, which are unable to effectively cope with dynamic obstacles or challenges in confined spaces. To overcome these shortcomings of the existing technology, the present invention provides a robotic bidirectional motion system and method based on mirror-enhanced experience replay, comprising a robotic bidirectional motion system based on mirror-enhanced experience replay and a robotic bidirectional motion method based on mirror-enhanced experience replay.
[0006] The present invention provides a robot bidirectional motion system based on mirror-enhanced experience playback, comprising:
[0007] The motion control module is configured to drive the robot from a starting point to a target position;
[0008] a mirror enhancement module, electrically connected to the motion control module, configured to obtain a forward trajectory of the robot from a starting point to a target position, extract a forward success trajectory from the forward trajectory, then perform a geometric symmetry transformation on the forward success trajectory to generate a reverse motion trajectory opposite to the forward motion, then extract additional pose information from all forward trajectories, obtain a transformation result of the additional pose information using a geometric symmetry transformation, and fuse the transformation result with the reverse motion trajectory to obtain a robot state equation;
[0009] A training module is electrically connected to the mirror enhancement module and is configured to optimize parameters of the motion control module using the robot state equation and the forward success trajectory in multiple scenarios through a flexible actor-critic algorithm to train the robot's bidirectional motion.
[0010] The robot bidirectional motion system based on mirror-enhanced experience replay disclosed in the present invention sets a motion control module, a mirror enhancement module and a training module. The motion control module is used to drive the robot from a starting point to a target position, and the mirror enhancement module is used to extract a forward success trajectory from the forward trajectory. The forward success trajectory is then geometrically transformed to generate a reverse motion trajectory opposite to the forward direction. Additional posture information is extracted from all forward trajectories, and the transformation result of the additional posture information is obtained by geometrically transforming. The transformation result is fused with the reverse motion trajectory to obtain the robot state equation. Finally, the training module is used to implement parameter optimization of the motion control module using the robot state equation and the forward success trajectory using a flexible actor-critic algorithm, and the robot's bidirectional motion is trained, so that the robot can demonstrate excellent navigation capabilities in various scenarios, ensuring navigation capabilities in scenarios with high uncertainty and dense obstacles. At the same time, since the mirror enhancement module performs a geometric symmetric transformation on the forward successful trajectory to generate a reverse motion trajectory opposite to the forward movement, it has bidirectional operation, which enhances the robot's ability to avoid being trapped. It can ensure that the robot can flexibly retreat in narrow passages or complex terrain, and ensure that it can complete tasks stably and efficiently in different scenarios. It not only improves the robot's navigation ability in changing scenarios, but also broadens the boundaries of deep reinforcement learning in practical applications, and provides a new solution for the robot's autonomous navigation in complex scenarios.
[0011] In a possible implementation, the mirroring enhancement module is configured to perform the following steps:
[0012] A1: Set an initial strategy, call the motion control module to drive the robot from the starting point to the target position according to the initial strategy, and record the state, action and reward of each step of the robot's forward movement to obtain the forward trajectory;
[0013] A2: extracting a forward success trajectory from the forward trajectory based on a custom forward success criterion, and then performing a geometric symmetry transformation on the forward success trajectory to generate a reverse motion trajectory opposite to the forward motion;
[0014] A3: Extract additional pose information from all forward trajectories and obtain the transformation result of the additional pose information using geometric symmetry transformation;
[0015] A4: Fusing the mapping result with the inverse motion trajectory to obtain the robot state equation;
[0016] This solution uses geometric symmetry transformation to generate reverse motion without relying on failed trajectories or modifying the reward function. Instead, it directly generates learning samples for backward motion by converting the states and actions in the forward trajectory. This not only improves the robot's bidirectional motion capabilities in complex scenarios, but also enables the robot to effectively learn reverse navigation without increasing additional hardware costs, thereby demonstrating better adaptability in scenarios where retreat is required, such as narrow spaces.
[0017] In a possible implementation, the forward trajectory is expressed as follows:
[0018]
[0019] Where,
[0020] T+1 represents the total number of steps;
[0021] x t represents the initial state of the robot at step t;
[0022] a t represents the action performed by the robot in step t;
[0023] r t Represents the feedback reward received by the robot after performing the action in step t;
[0024] This solution ensures the generation of a forward trajectory, which records the robot's movement path from the starting point to the target, ensuring the efficiency of generating reverse motion using geometric symmetry transformation in the later stage.
[0025] In a possible implementation, the transformation formula of the geometric symmetry transformation is as follows:
[0026]
[0027] a t,mirror =-a t ,
[0028] Where,
[0029] a t,mirror Represents a t Perform the transformation result of reversing the direction of motion;
[0030] represents the target coordinates of the robot after transformation in the robot reference frame at step t;
[0031] L trepresents the minimum pooled LiDAR obstacle distance data of the robot at step t;
[0032] v t represents the linear velocity of the robot at step t;
[0033] ω t represents the angular velocity of the robot at step t;
[0034] The above transformation formula can be used to perform geometric symmetry transformation on the forward successful trajectory to generate a reverse motion trajectory opposite to the forward motion.
[0035] In a possible implementation, step A3 includes the following steps:
[0036] A31: Set up the standard experience replay buffer B as the classic buffer for reinforcement learning navigation tasks and set up a mirror enhancement buffer To capture additional pose information for each forward trajectory;
[0037] A32: Obtaining a transformation result of the additional pose information by using geometric symmetry transformation;
[0038] This solution balances the learning of forward and reverse motion by setting up a double-buffered experience storage mechanism, that is, combining a standard replay buffer and a mirror enhancement buffer. The standard replay buffer is used to store conventional training experience, while the mirror enhancement buffer is specifically used to store mirror experience generated by the reverse of the forward trajectory, thereby ensuring the accuracy and efficiency of the extraction of additional pose information.
[0039] In one possible implementation, the standard experience replay buffer The structure is shown below:
[0040]
[0041] Where,
[0042] d t Represents the termination code of the robot at step t;
[0043] The standard replay buffer with the above structure can be used to store regular training experience to ensure the efficiency of experience extraction.
[0044] In a possible implementation, the mirror enhancement buffer The structure is shown below:
[0045]
[0046] Where,
[0047] p tRepresents the additional pose information of the robot in step t;
[0048] The mirror enhancement buffer with the above structure can be used specifically to store mirror experiences generated in reverse from the forward trajectory. Through this design, the robot can learn forward and backward movement capabilities simultaneously during training, avoiding the bias caused by traditional deep reinforcement learning methods that only rely on forward trajectories.
[0049] In one possible implementation, the robot state equation is expressed as follows:
[0050]
[0051] In one possible implementation, the training module is configured to set an adaptive course when optimizing parameters of the motion control module, and lock a single scene according to a scenario selection probability calculation formula to start optimizing parameters of the motion control module in the single scene, and unlock the scene when the success rate in the single scene exceeds a specified value;
[0052] The scene selection probability calculation formula is:
[0053] p(env i )=(1-μ i ) / ∑ j (1-μ j ),
[0054] Where,
[0055] p(env i ) represents the probability of locking the i-th single scene;
[0056] μ i represents the average success rate of locking the i-th single scenario;
[0057] This solution ensures that the robot can retreat flexibly in narrow passages or complex terrain, ensuring that it can complete tasks stably and efficiently in different scenarios.
[0058] Another technical solution of the present invention is to provide a robot bidirectional motion method based on mirror-enhanced experience playback, comprising the following steps:
[0059] S1: Set the initial strategy to enable the motion control module to drive the robot from the starting point to the target position;
[0060] S2: obtaining a forward trajectory of the robot from a starting point to a target position through a mirror enhancement module, and extracting a forward successful trajectory from the forward trajectory;
[0061] S3: performing a geometric symmetric transformation on the forward successful trajectory through the mirror enhancement module to generate a reverse motion trajectory opposite to the forward motion, then extracting additional pose information from all forward trajectories, and obtaining a transformation result of the additional pose information by using the geometric symmetric transformation, and fusing the transformation result with the reverse motion trajectory to obtain a robot state equation;
[0062] S4: Optimizing parameters of the motion control module using the robot state equation and the forward success trajectory using a flexible actor-critic algorithm through the training module to train the bidirectional motion of the robot.
[0063] The method disclosed in the present application uses a motion control module to drive the robot from a starting point to a target position, uses a mirror enhancement module to extract a forward success trajectory from the forward trajectory, then performs a geometric symmetry transformation on the forward success trajectory to generate a reverse motion trajectory opposite to the forward direction, extracts additional posture information from all forward trajectories, and uses geometric symmetry transformation to obtain the transformation result of the additional posture information, fuses the transformation result with the reverse motion trajectory to obtain the robot state equation, and finally implements a flexible actor-critic algorithm through the set training module to optimize the parameters of the motion control module using the robot state equation and the forward success trajectory, trains the robot's bidirectional motion, and enables the robot to demonstrate excellent navigation capabilities in various scenarios, ensuring navigation capabilities in scenarios with high uncertainty and dense obstacles. At the same time, since the mirror enhancement module performs a geometric symmetric transformation on the forward successful trajectory to generate a reverse motion trajectory opposite to the forward movement, it has bidirectional operation, which enhances the robot's ability to avoid being trapped. It can ensure that the robot can flexibly retreat in narrow passages or complex terrain, and ensure that it can complete tasks stably and efficiently in different scenarios. It not only improves the robot's navigation ability in changing scenarios, but also broadens the boundaries of deep reinforcement learning in practical applications, and provides a new solution for the robot's autonomous navigation in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 Schematic diagram of the structure of a robot bidirectional motion system based on mirror-enhanced experience playback disclosed in an embodiment of the present application;
[0065] Figure 2 This is a flowchart of the operation of the mirror enhancement module disclosed in the embodiments of this application;
[0066] Figure 3 Schematic diagram of the method disclosed in the examples of this application. DETAILED DESCRIPTION
[0067] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.
[0068] In the embodiments of the present application, unless otherwise clearly specified and limited, the electrical connection between the first feature and the second feature means that there is transmission of electrical signals between the first feature and the second feature, that is, there is an electrical relationship, and the way to achieve the transmission of electrical signals can be electrical connection of wires, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication achieved by channels, etc.
[0069] In the embodiments of the present application, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, a first feature being "above," "above," and "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.
[0070] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0071] See also Figures 1 to 3 The embodiment of the present application discloses a robot bidirectional motion system based on mirror-enhanced experience playback. The structural diagram of the robot bidirectional motion system is shown in FIG. Figure 1 As shown, the robot bidirectional motion system includes a motion control module, a mirror enhancement module and a training module, wherein the mirror enhancement module is electrically connected to the motion control module, and the training module is electrically connected to the mirror enhancement module.
[0072] In the robot bidirectional motion system, the motion control module includes an actor network, which is configured to drive the robot from a starting point to a target position and can accept a defined operation strategy to drive the robot from its starting point to the target position.
[0073] In this robot bidirectional motion system, a mirror enhancement module is configured to obtain the robot's forward trajectory from the starting point to the target position, extract the forward successful trajectory from the forward trajectory, and then perform a geometric symmetry transformation on the forward successful trajectory to generate a reverse motion trajectory opposite to the forward direction. After that, additional pose information is extracted from all forward trajectories, and the transformation result of the additional pose information is obtained by using the geometric symmetry transformation. The transformation result is then fused with the reverse motion trajectory to obtain the robot state equation.
[0074] See also Figure 2 In this embodiment, the mirror enhancement module is configured to perform the following steps:
[0075] A1: Set the initial strategy, call the motion control module to drive the robot from the starting point to the target position according to the initial strategy, and record the state, action, and reward of each step of the robot's forward movement to obtain the forward trajectory. The expression of the forward trajectory is as follows:
[0076]
[0077] Where,
[0078] T+1 represents the total number of steps;
[0079] x t represents the initial state of the robot at step t;
[0080] a t represents the action performed by the robot in step t;
[0081] r t Represents the feedback reward received by the robot after performing the action in step t.
[0082] A2: Based on the customized forward success criteria, the forward successful trajectory is extracted from the forward trajectory, and then the forward successful trajectory is geometrically transformed to generate a reverse motion trajectory opposite to the forward trajectory. The transformation formula of the geometric symmetry transformation is as follows:
[0083]
[0084] a t,mirror =-a t ,
[0085] Where,
[0086] a t,mirror Represents a t Perform the transformation result of reversing the direction of motion;
[0087] represents the target coordinates of the robot after transformation in the robot reference frame at step t;
[0088] L t Represents the minimum pooled LiDAR obstacle distance data of the robot at step t;
[0089] v t represents the linear velocity of the robot at step t;
[0090] ω t represents the angular velocity of the robot at step t.
[0091] A3: Extract additional pose information from all forward trajectories and use geometric symmetry transformation to obtain the transformation results of the additional pose information.
[0092] In this embodiment, step A3 includes the following steps:
[0093] A31: Setting the standard experience replay buffer Setting up a mirrored augmented buffer as a classic buffer for reinforcement learning navigation tasks This captures additional pose information for each forward trajectory.
[0094] Standard experience replay buffer The structure is shown below:
[0095]
[0096] Where,
[0097] d t Represents the termination code of the robot at step t.
[0098] Mirror Enhanced Buffer The structure is shown below:
[0099]
[0100] Where,
[0101] p t Represents the additional pose information of the robot in step t.
[0102] A32: Transformation result using geometric symmetry transformation to obtain additional pose information.
[0103] A4: Combine the mapping result with the inverse motion trajectory to obtain the robot state equation. The robot state equation is calculated as follows:
[0104]
[0105] In this robot bidirectional motion system, the training module is configured to use the SoftActor-Critic (SAC) algorithm to optimize the parameters of the motion control module in various scenarios using the robot state equation and positive success trajectory to train the robot's bidirectional motion. In this embodiment, the training module sets adaptive courses when optimizing the parameters of the motion control module. The number of courses is controlled according to actual conditions, such as Figure 3 As shown, a single scene is locked according to the scene selection probability calculation formula to start parameter optimization of the motion control module in the single scene. When the success rate in the single scene exceeds a specified value (70% in this embodiment), it is unlocked. The scene selection probability calculation formula is:
[0106] p(env i )=(1-μ i ) / ∑ j (1-μ j ),
[0107] Where,
[0108] p(env i ) represents the probability of locking the i-th single scene;
[0109] μ i Represents the average success rate of locking the i-th single scene.
[0110] The following will further disclose the corresponding method of using the robot bidirectional motion system based on mirror-enhanced experience playback in this embodiment. The flow chart of this method is as follows: Figure 3 As shown, the method includes the following steps:
[0111] S1: Set the initial strategy to enable the motion control module to drive the robot from the starting point to the target position;
[0112] S2: Obtain the robot's forward trajectory from the starting point to the target position through the mirror enhancement module, and extract the forward successful trajectory from the forward trajectory;
[0113] S3: The forward successful trajectory is geometrically transformed through the mirror enhancement module to generate a reverse motion trajectory opposite to the forward trajectory. Then, additional pose information is extracted from all forward trajectories, and the transformation result of the additional pose information is obtained by using the geometric symmetry transformation. The transformation result is fused with the reverse motion trajectory to obtain the robot state equation;
[0114] S4: Optimizing the parameters of the motion control module using the robot state equation and the forward success trajectory through the training module to train the robot's bidirectional motion.
[0115] This example demonstrates the navigation performance of the robot's bidirectional motion system in multiple high-complexity simulation scenarios. Compared to baseline systems (DRL-DCLP and MAER-Raw), the robot's bidirectional motion system achieved a 100% success rate in crowded environment 1 (80% for DRL-DCLP and 70% for MAER-Raw), and the collision rate was reduced to 0% (16% for DRL-DCLP and 12% for MAER-Raw). In crowded environment 2, the robot's bidirectional motion system still maintained a 98% success rate (86% for MAER-Raw and 80% for DRL-DCLP), and the average number of steps (AES*) was significantly shortened by 201.31 steps (207.53 steps for DRL-DCLP), verifying the efficiency of its path planning. In addition, the maximum average navigation score (MANS) of the robot's bidirectional motion system is 0.29±0.28, far exceeding the comparison methods of 0.08±0.56 (DRL-DCLP) and 0.14±0.51 (MAER-Raw), indicating that it has both stability and efficiency in complex environments.
[0116] In extreme scenarios where backward escape is required (such as in a narrow corridor with the initial posture facing backward or against a wall), traditional 3D deep reinforcement learning methods tend to crash or get stuck due to their forward-biased action. However, the technical solution of the robot's bidirectional motion system successfully escapes backward through a mirror experience-enhanced replay mechanism. For example, in a test with the initial posture facing backward in a narrow corridor, the robot's bidirectional motion system escaped from the initial position and reached the target through reverse action, while DRL-DCLP directly collided with the obstacle due to its inability to adjust the direction. The improved success rate in such scenarios is directly attributed to the enhancement of experience data by the bidirectional action recovery mechanism, which solves the local optimality problem caused by the limited action space of traditional methods.
[0117] In real-world testing (such as reversing out of a garage and avoiding dynamic pedestrian obstacles), the robot's bidirectional motion system demonstrated remarkable adaptability. In a dynamic pedestrian scenario, when a pedestrian blocked the forward path, the robot's bidirectional motion system proactively triggered a backward motion strategy to escape the predicament, achieving a 100% success rate and avoiding collisions. In contrast, the DRL-DCLP system suffered from insufficient backward motion capability and lack of flexibility, leading to collisions. Furthermore, in a garage reversing task, the present invention rapidly completed navigation by synthesizing a backward trajectory, demonstrating its advantages.
[0118] In summary, the robot bidirectional motion system based on mirror-enhanced experience replay disclosed in this embodiment, by setting a motion control module, a mirror enhancement module and a training module, uses the motion control module to drive the robot from the starting point to the target position, uses the mirror enhancement module to extract the forward success trajectory from the forward trajectory, and then performs a geometric symmetry transformation on the forward success trajectory to generate a reverse motion trajectory opposite to the forward direction, and extracts additional posture information from all forward trajectories, and uses the geometric symmetry transformation to obtain the transformation result of the additional posture information, and merges the transformation result with the reverse motion trajectory to obtain the robot state equation, and finally, through the set training module, a flexible actor-critic algorithm is used to optimize the parameters of the motion control module using the robot state equation and the forward success trajectory, and train the robot's bidirectional motion, so that the robot can demonstrate excellent navigation capabilities in various scenarios, ensuring navigation capabilities in scenarios with high uncertainty and dense obstacles. At the same time, since the mirror enhancement module performs a geometric symmetric transformation on the forward successful trajectory to generate a reverse motion trajectory opposite to the forward movement, it has bidirectional operation, which enhances the robot's ability to avoid being trapped. It can ensure that the robot can flexibly retreat in narrow passages or complex terrain, and ensure that it can complete tasks stably and efficiently in different scenarios. It not only improves the robot's navigation ability in changing scenarios, but also broadens the boundaries of deep reinforcement learning in practical applications, and provides a new solution for the robot's autonomous navigation in complex scenarios.
[0119] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.
[0120] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are mutually inconsistent.
[0121] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A robot bidirectional motion system based on mirror-enhanced experience playback, characterized in that: include: The motion control module is configured to drive the robot from a starting point to a target position; a mirror enhancement module, electrically connected to the motion control module, configured to obtain a forward trajectory of the robot from a starting point to a target position, extract a forward success trajectory from the forward trajectory, then perform a geometric symmetry transformation on the forward success trajectory to generate a reverse motion trajectory opposite to the forward motion, then extract additional pose information from all forward trajectories, obtain a transformation result of the additional pose information using a geometric symmetry transformation, and fuse the transformation result with the reverse motion trajectory to obtain a robot state equation; A training module is electrically connected to the mirror enhancement module and is configured to optimize parameters of the motion control module using the robot state equation and the forward success trajectory in multiple scenarios through a flexible actor-critic algorithm to train the robot's bidirectional motion.
2. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 1, characterized in that: The mirror enhancement module is configured to perform the following steps: A1: Set an initial strategy, call the motion control module to drive the robot from the starting point to the target position according to the initial strategy, and record the state, action and reward of each step of the robot's forward movement to obtain the forward trajectory; A2: extracting a forward success trajectory from the forward trajectory based on a custom forward success criterion, and then performing a geometric symmetry transformation on the forward success trajectory to generate a reverse motion trajectory opposite to the forward motion; A3: Extract additional pose information from all forward trajectories and obtain the transformation result of the additional pose information using geometric symmetry transformation; A4: Fusing the mapping result with the inverse motion trajectory to obtain the robot state equation.
3. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 2, characterized in that: The expression of the forward trajectory is as follows: Where, T+1 represents the total number of steps; x t represents the initial state of the robot at step t; a t represents the action performed by the robot in step t; r t Represents the feedback reward received by the robot after performing the action in step t.
4. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 3, characterized in that: The transformation formula of the geometric symmetry transformation is as follows: a t,mirror =-a t , Where, a t,mirror Represents a t Perform the transformation result of reversing the direction of motion; represents the target coordinates of the robot after transformation in the robot reference frame at step t; L t represents the minimum pooled LiDAR obstacle distance data of the robot at step t; v t represents the linear velocity of the robot at step t; ω t represents the angular velocity of the robot at step t.
5. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 4, characterized in that: Step A3 includes the following steps: A31: Setting the standard experience replay buffer Setting up a mirrored augmented buffer as a classic buffer for reinforcement learning navigation tasks To capture additional pose information for each forward trajectory; A32: Obtain a transformation result of the additional pose information using geometric symmetry transformation.
6. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 5, characterized in that: The standard experience replay buffer The structure is shown below: Where, d t Represents the termination code of the robot at step t.
7. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 6, characterized in that: The mirror enhancement buffer The structure is shown below: Where, p t Represents the additional pose information of the robot in step t.
8. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 7, characterized in that: The robot state equation is as follows:
9. The robot bidirectional motion system based on mirror-enhanced experience playback according to claim 8, characterized in that: The training module is configured to set an adaptive course when optimizing parameters of the motion control module, and lock a single scene according to a scene selection probability calculation formula to start optimizing parameters of the motion control module in the single scene, and unlock the scene when the success rate in the single scene exceeds a specified value; The scene selection probability calculation formula is: p(env i )=(1-μ i ) / ∑ j (1-m j ), Where, p(env i ) represents the probability of locking the i-th single scene; μ i Represents the average success rate of locking the i-th single scene.
10. A robot bidirectional motion method based on mirror-enhanced experience playback, characterized in that: The robot bidirectional motion system based on mirror-enhanced experience replay according to any one of claims 1 to 7 comprises the following steps: S1: Set the initial strategy to enable the motion control module to drive the robot from the starting point to the target position; S2: obtaining a forward trajectory of the robot from a starting point to a target position through a mirror enhancement module, and extracting a forward successful trajectory from the forward trajectory; S3: performing a geometric symmetric transformation on the forward successful trajectory through the mirror enhancement module to generate a reverse motion trajectory opposite to the forward motion, then extracting additional pose information from all forward trajectories, and obtaining a transformation result of the additional pose information by using the geometric symmetric transformation, and fusing the transformation result with the reverse motion trajectory to obtain a robot state equation; S4: Optimizing parameters of the motion control module using the robot state equation and the forward success trajectory using a flexible actor-critic algorithm through the training module to train the bidirectional motion of the robot.