Control method and control device of mobile platform, equipment and storage medium
By constructing a digital twin platform for mobile platforms and an optimization strategy network for simulation environments, the problems of high testing costs and poor adaptability of navigation algorithms on mobile platforms are solved, achieving efficient optimization and stability improvement of navigation algorithms.
Patent Information
- Application Number
- CN202511426038.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-13
AI Technical Summary
In existing technologies, the setup cost of testing sites for navigation algorithms on mobile platforms is high and the testing efficiency is low, resulting in poor adaptability to various environments.
By constructing a digital twin platform for the mobile platform, optimizing the policy network using a simulation environment, and verifying the optimized policy network in a real environment, the reliability and stability of the network under different conditions are ensured.
This reduces the optimization cost of the policy network, improves its adaptability and reliability in real-world environments, and ensures the stability and efficiency of the navigation algorithm.
Smart Images

Figure CN121325699A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent driving technology, and in particular to a control method, control device, equipment and storage medium for a mobile platform. Background Technology
[0002] Autonomous navigation capability refers to the ability of a mobile platform to perform tasks such as positioning, movement, path planning, and obstacle avoidance in unknown or dynamic environments without relying on external manual remote control or pre-laid guide rails, magnetic strips, or other guiding tracks, and by relying on sensors and algorithms integrated within the mobile platform.
[0003] The realization of automatic navigation capabilities depends on the reliability and stability of the navigation algorithms integrated into the mobile platform. To ensure the reliability and stability of navigation algorithms in actual use, related technologies typically require setting up various test sites with different environments to test the environmental adaptability of the navigation algorithms integrated into the mobile platform.
[0004] However, the above methods have high setup costs for test sites, low testing efficiency, and the limited availability of test sites can lead to poor environmental adaptability of the navigation algorithms integrated into the mobile platform. Summary of the Invention
[0005] This application provides a control method, control device, equipment, and storage medium for a mobile platform to solve the technical problems existing in related technologies. Specifically, it includes the following technical solutions.
[0006] In a first aspect, this application provides a control method for a mobile platform, the method comprising: optimizing a policy network based on a first task result after a digital twin platform of the mobile platform performs a first navigation task, wherein the digital twin platform and the mobile platform output an action vector less than an error threshold when the input state vectors are the same, the policy network being used to map the state vectors to action vectors, the first navigation task being a task in which the digital twin platform moves toward a first target position in a simulation environment based on the policy network; obtaining a second task result after the mobile platform performs a second navigation task, the second navigation task being a task in which the mobile platform moves toward a second target position in a real environment based on the optimized policy network; and controlling the mobile platform to perform navigation operations according to the optimized policy network when the second task result meets preset conditions.
[0007] In some possible implementations, there are multiple real environments, the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments, the second task result includes a second success rate of the multiple second navigation tasks, and the second task result also includes a first collision count for indicating the occurrence of collision events in the multiple second navigation tasks; the preset conditions include the first collision count being less than a first collision count threshold, and the second success rate being greater than a success rate threshold.
[0008] In some possible implementations, there are multiple simulation environments, and the first navigation task includes multiple first navigation tasks corresponding to the multiple simulation environments. The first task result includes a first success rate of the multiple first navigation tasks. There are multiple real environments, and the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments. The second task result includes a second success rate of the multiple second navigation tasks. The second task result also includes a first number of collisions to indicate the occurrence of collision events in the multiple second navigation tasks. The preset conditions include that the first number of collisions is less than a second number threshold, and the difference between the first success rate and the second success rate is less than a first difference threshold.
[0009] In some possible implementations, the method further includes: if the result of the second task does not meet the preset condition, and if the difference between the first success rate and the second success rate is less than a second difference threshold, then the optimized policy network is further optimized based on the result of the second task; wherein the second difference threshold is greater than the first difference threshold.
[0010] In some possible implementations, the method further includes: obtaining a third task result after the mobile platform executes multiple third navigation tasks, wherein the multiple third navigation tasks are tasks in which the mobile platform moves towards a third target location in multiple real environments based on a re-optimized policy network; and controlling the mobile platform to execute navigation operations according to the re-optimized policy network when the difference between the first success rate and the third success rate in the third task result is less than the first difference threshold, and the second collision count in the third task result is less than the second collision count threshold; wherein the third success rate is used to indicate the success rate of the multiple third navigation tasks, and the second collision count is used to indicate the number of collision events that occur in the multiple third navigation tasks.
[0011] In some possible implementations, the simulation environment includes environmental parameters for describing the physical quantities of the interaction between the digital twin platform and the simulation environment; the method further includes: if the result of the second task does not meet the preset conditions, and if the difference between the first success rate and the second success rate is greater than or equal to the second difference threshold, adjusting the environmental parameters to increase the randomness of the simulation environment.
[0012] In some possible implementations, optimizing the policy network based on the first task result after the mobile platform's digital twin platform performs the first navigation task in a simulation environment includes: correcting the digital twin platform based on the hardware errors of the mobile platform; and optimizing the policy network based on the first task result after the corrected digital twin platform performs the first navigation task.
[0013] In a second aspect, this application provides a control device for a mobile platform, the device comprising an optimization module, an acquisition module, and a navigation module; the optimization module is configured to optimize a policy network based on a first task result after a digital twin platform of the mobile platform performs a first navigation task, wherein the digital twin platform and the mobile platform, when having the same input state vector, output an action vector less than an error threshold, the policy network being used to map the state vector to an action vector, and the first navigation task being a task in which the digital twin platform moves towards a first target position in a simulation environment based on the policy network; the acquisition module is configured to acquire a second task result after the mobile platform performs a second navigation task, the second navigation task being a task in which the mobile platform moves towards a second target position in a real environment based on the optimized policy network; the navigation module is configured to control the mobile platform to perform navigation operations according to the optimized policy network when the second task result meets preset conditions.
[0014] In some possible implementations, there are multiple real environments, the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments, the second task result includes a second success rate of the multiple second navigation tasks, and the second task result also includes a first collision count for indicating the occurrence of collision events in the multiple second navigation tasks; the preset conditions include the first collision count being less than a first collision count threshold, and the second success rate being greater than a success rate threshold.
[0015] In some possible implementations, there are multiple simulation environments, and the first navigation task includes multiple first navigation tasks corresponding to the multiple simulation environments. The first task result includes a first success rate of the multiple first navigation tasks. There are multiple real environments, and the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments. The second task result includes a second success rate of the multiple second navigation tasks. The second task result also includes a first number of collisions to indicate the occurrence of collision events in the multiple second navigation tasks. The preset conditions include that the first number of collisions is less than a second number threshold, and the difference between the first success rate and the second success rate is less than a first difference threshold.
[0016] In some possible implementations, the optimization module is further configured to, if the difference between the first success rate and the second success rate is less than a second difference threshold, further optimize the optimized policy network based on the second task result when the second task result does not meet the preset condition; wherein the second difference threshold is greater than the first difference threshold.
[0017] In some possible implementations, the acquisition module is further configured to acquire the third task results after the mobile platform executes multiple third navigation tasks, wherein the multiple third navigation tasks are tasks in which the mobile platform moves towards a third target location in multiple real environments based on a further optimized policy network; the navigation module is further configured to control the mobile platform to perform navigation operations according to the further optimized policy network when the difference between the first success rate and the third success rate in the third task results is less than the first difference threshold, and the second collision count in the third task results is less than the second number threshold; wherein the third success rate is used to indicate the success rate of the multiple third navigation tasks, and the second collision count is used to indicate the number of collision events that occur in the multiple third navigation tasks.
[0018] In some possible implementations, the simulation environment includes environmental parameters for describing the physical quantities of the interaction between the digital twin platform and the simulation environment; the optimization module is further configured to increase the randomness of the simulation environment by adjusting the environmental parameters if the difference between the first success rate and the second success rate is greater than or equal to the second difference threshold when the result of the second task does not meet the preset conditions.
[0019] In some possible implementations, when the optimization module optimizes the policy network based on the first task result after the mobile platform's digital twin platform performs the first navigation task in a simulation environment, it is configured to: correct the digital twin platform based on the hardware error of the mobile platform; and optimize the policy network based on the first task result after the corrected digital twin platform performs the first navigation task.
[0020] In a third aspect, this application provides an electronic device for controlling a mobile platform, comprising: a memory storing at least one program instruction for controlling the mobile platform; and a processor, wherein when the program instruction is executed by the processor, the mobile platform implements the control method of the first aspect of this application or any possible embodiment of the first aspect.
[0021] In a fourth aspect, this application provides a computer program (product) comprising computer program / instructions, which are executed by a processor to enable a mobile platform to implement the control method of the first aspect or any possible implementation thereof.
[0022] In a fifth aspect, this application provides a computer-readable storage medium having stored thereon program instructions for controlling a mobile platform, which, when executed by one or more processors, cause the mobile platform to implement the control method of the first aspect or any possible implementation of the first aspect of this application.
[0023] The beneficial effects of the technical solution provided in this application include at least the following:
[0024] The technical solution provided in this application, on the one hand, by constructing a digital twin platform for the mobile platform and ensuring that the digital twin platform and the mobile platform have the same or similar responses to the same state vector, allows the test scenario of the policy network to be migrated to a simulation environment. This fully leverages the convenience and variability of setting up the simulation environment to optimize the policy network, which helps reduce the optimization cost and improve optimization efficiency. On the other hand, after optimizing the policy network, the method provided in this application's embodiments further examines the performance of the optimized policy network in the real environment based on the second task result of the mobile platform performing a second navigation task in a real environment using the optimized policy network. This helps improve the policy network's adaptability to the real environment and ensures its reliability and stability. Attached Figure Description
[0025] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0026] Figure 1 This is a schematic diagram of an implementation scenario provided in the embodiments of this application;
[0027] Figure 2 This is a flowchart of the control method for a mobile platform provided in an embodiment of this application;
[0028] Figure 3 This is a schematic diagram of the structure of the control device for the mobile platform provided in the embodiments of this application;
[0029] Figure 4 This is a schematic diagram of the structure of an electronic device for controlling a mobile platform provided in an embodiment of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0032] Figure 1 This is a schematic diagram of an implementation scenario provided in an embodiment of this application. (Reference) Figure 1 The implementation scenarios provided in this application embodiment may include a mobile platform 11 and a control unit 13.
[0033] The mobile platform 11 integrates sensors and a policy network. The sensors can be used, but are not limited to, enabling the mobile platform to perceive its environment. The policy network can be used, but is not limited to, mapping the mobile platform's perception of its environment into its own actions, thereby enabling the mobile platform 11 to have automatic navigation capabilities and complete tasks such as positioning and movement, path planning, and obstacle avoidance in its environment.
[0034] In some embodiments, the mobile platform may be any device with autonomous navigation capabilities, such as a wheeled robot, a tracked robot, a humanoid robot, or an unmanned vehicle; this application makes no limitations in this regard. The sensors integrated into the mobile platform 11 include, but are not limited to, one or more of the following: LiDAR, depth camera, fisheye camera, and ultrasonic radar.
[0035] The control unit 12 can communicate with the mobile platform 11 via wired or wireless means, enabling the control platform 12 to optimize the policy network integrated in the mobile platform 11, acquire data from the mobile platform 11's autonomous navigation in a real environment based on the optimized policy network, and control the mobile platform 11 to perform navigation operations according to the optimized policy network. The navigation operations are used to instruct the mobile platform 11 on relevant operations involved in autonomous navigation, such as sensing the environment through sensors and performing specific actions based on the perceived environmental conditions. This application does not impose any limitations in this regard.
[0036] Those skilled in the art should understand that the mobile platform 11 and control unit 12 described above are merely examples. Other existing or future mobile platforms and control units that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0037] Figure 2 This is a flowchart of a control method for a mobile platform provided in an embodiment of this application. This control method can, for example, be... Figure 1 The control unit shown performs the actions, and this application makes no limitations in this regard. See also Figure 2 The mobile platform control method provided in this application embodiment may include steps S210-S230.
[0038] Step S210: Optimize the policy network based on the first task result after the digital twin platform of the mobile platform performs the first navigation task. When the input state vectors of the digital twin platform and the mobile platform are the same, the output action vector is less than the error threshold. The policy network is used to map the state vector to the action vector. The first navigation task is the task of the digital twin platform moving towards the first target position in the simulation environment based on the policy network.
[0039] For example, a digital twin platform for a mobile platform is used to indicate a high-fidelity, synchronizeable, and evolvable virtual digital copy of the mobile platform built in a computer system. It can be used, but is not limited to, for bidirectional digital binding with the mobile platform and quantifying and predicting the behavior of the mobile platform. In some embodiments, the digital twin platform includes, but is not limited to, a first digital copy corresponding to the geometry and physics of the mobile platform, used to, but not limited to, reproducing characteristics such as the mobile platform's mass inertia, joint friction, motor saturation, and contact stiffness; a second digital copy corresponding to the mobile platform's sensing interface, used to, but not limited to, reproducing characteristics such as the mobile platform's sensor noise, latency, and dropout rate; a third digital copy corresponding to the mobile platform's external coupling, used to, but not limited to, reproducing the mobile platform's interaction characteristics with the environment, such as ground friction, slope, obstacle mass, and dynamic pedestrian distribution; and a fourth digital copy corresponding to the mobile platform's control and execution, used to, but not limited to, reproducing characteristics such as the mobile platform's command latency, current limiting, and fault modes.
[0040] The policy network can be used, but is not limited to, mapping state vectors to action vectors, such as mapping the state vector of a mobile platform to the action vector of the mobile platform. The state vector of the mobile platform indicates the environmental state after the mobile platform's sensors perceive its surroundings. The state vector can be a high-dimensional state vector composed of one or more of the following raw data: laser point cloud data acquired by LiDAR, depth image data acquired by a depth camera, wide-angle image data acquired by a fisheye camera, and time-of-flight-distance data from ultrasonic radar. Alternatively, it can be a low-dimensional state vector synthesized from artificially generated features after feature extraction from the laser point cloud data acquired by LiDAR, depth image data acquired by a depth camera, wide-angle image data acquired by a fisheye camera, and time-of-flight-distance data from ultrasonic radar. This application does not impose any limitations in this regard.
[0041] When the state vector is a high-dimensional state vector composed of one or more of the original data, the policy network can, for example, adopt an end-to-end model that can directly map the high-dimensional state vector to the action vector. This can avoid the accumulation of errors and loss of information in intermediate modules, retain more details in the original data, and ensure the accuracy and generalization ability of navigation operation control of the mobile platform during autonomous navigation.
[0042] When the state vector is a low-dimensional state vector synthesized from manually generated features, the policy network can, for example, include a perception module for dimensionality reduction and clustering of the original data, a mapping module for constructing a two-dimensional or three-dimensional occupancy grid map, a path planning module for searching for collision-free global and local trajectories within the occupancy grid map, and a control module for tracking the trajectory and outputting control commands corresponding to the action vector, thereby reducing the computational requirements on the hardware. The manually generated features are used to indicate manually designed geometric or statistical characteristics.
[0043] The motion vector of the mobile platform is used to indicate the motion control quantities used to control the movement of the mobile platform. These include, for example, a one-dimensional motion vector corresponding to one of linear velocity, angular velocity, joint torque, wheel drive force, etc., or a multi-dimensional vector synthesized from multiple vectors. In some embodiments, the motion vector of the mobile platform can be discrete, to reduce the computing power requirements of the hardware in the mobile platform, or continuous, to improve the smoothness of the mobile platform's movement and reduce motion jitter. It can be adjusted according to actual application requirements, and this application does not impose any limitations in this regard.
[0044] The simulation environment refers to a virtual training space comprised of physical parameters, geometric layout, lighting conditions, dynamic obstacles, and mobile platform-environment coupling characteristics generated by random sampling. It can be used, but is not limited to, to provide policy networks with highly diverse, low-cost training data with controllable errors with the mobile platform. The mobile platform-environment coupling characteristics indicate the physical quantities of the interaction between the mobile platform and the simulation environment, determined by the environmental parameters of both the mobile platform and the simulation environment.
[0045] In some embodiments, the method for optimizing the policy network is, for example, to optimize the policy network using a value network and a composite reward function. The value network can be used, but is not limited to, predicting the current state vector or the expected cumulative reward of the state vector-action vector pair to provide gradient direction for the optimization of the policy network. The composite reward function is, for example, a composite reward function synthesized from at least two of the following: a first reward function for positively rewarding reaching the target position; a second reward function for negatively rewarding collision events (i.e., penalizing collision events); a third reward function for positively rewarding motion stability during motion; a fourth reward function for rewarding time efficiency in completing the navigation task; and a fifth reward function for rewarding dynamic obstacle avoidance behavior during the completion of the navigation task. The specific form of the function can be adjusted according to the actual application, and this application does not impose any limitations in this regard.
[0046] Considering that in real-world scenarios, there may be discrepancies between the calibration parameters and actual parameters of the hardware in a mobile platform, such as the discrepancies between the calibration parameters and actual parameters of sensors and motors, this can lead to different responses from the digital twin platform to the same state vector, meaning the error in the output action vector exceeds an error threshold. Therefore, in some embodiments, the policy network is optimized based on the first task result after the digital twin platform of the mobile platform performs the first navigation task in a simulation environment. This includes: correcting the digital twin platform based on the hardware errors of the mobile platform; and optimizing the policy network based on the first task result after the corrected digital twin platform performs the first navigation task, to ensure high fidelity between the digital twin platform and the mobile platform.
[0047] Step S220: Obtain the second task result after the mobile platform executes the second navigation task. The second navigation task is the task of the mobile platform moving towards the second target location in the real environment based on the optimized policy network.
[0048] For example, the real environment is used to indicate an uncontrollable external space composed of the actual site, real lighting, real dynamic obstacles, and real ground physical parameters. The second task result, which is the second navigation task to reach the second target location without a pre-set guided path or external remote control, but only through the sensors integrated in the mobile platform and the optimized strategy network, can be used, but is not limited to, as the final criterion for determining whether the optimized strategy network is officially online, that is, to solidify the optimized strategy network integrated in the mobile platform.
[0049] A digital twin platform is a high-fidelity, synchronizeable, and evolvable virtual digital copy of a mobile platform, which can be used to quantify and predict the behavior of the mobile platform. In other words, when the input state vectors of the digital twin platform and the mobile platform are the same, the output action vectors are less than an error threshold. Under the same state vector, the output action vectors of the mobile platform and the digital twin platform are similar or identical, and the same action vectors can be transformed into the same motion state. Therefore, the method provided in this application can use the digital twin platform to batch generate navigation scenarios with controllable errors to the mobile platform in a simulation environment, and rapidly iteratively optimize the policy network using the results of the first task. The value of the error threshold can be adjusted according to the actual application situation, and this application does not impose any restrictions in this regard.
[0050] Step S230: If the result of the second task meets the preset conditions, control the mobile platform to perform navigation operations according to the optimized strategy network.
[0051] As described above, the results of the second task can be used, but are not limited to, as the final criterion for determining whether the optimized policy network should be officially launched. For example, the success rate of the second navigation task can be used to determine whether the optimized policy network integrated into the mobile platform should be solidified. The success rate of the second navigation task indicates the proportion of successful tasks in at least one second navigation task out of the total number of tasks. For example, the criteria for determining the success of a second navigation task might be: after the completion of a second navigation task, if the distance between the mobile platform and the second target location is less than a distance threshold, the completion time of the second navigation task is less than a time threshold, and / or the collision rate during the second navigation task is less than a collision rate threshold, then the second navigation task is considered successful; otherwise, the second navigation task is considered a failure. The values of the distance threshold, time threshold, and collision rate threshold can be adjusted according to actual application conditions, and this application does not impose any restrictions in this regard.
[0052] In some embodiments, there may be multiple real-world environments, and the second navigation task may include multiple second navigation tasks corresponding to the multiple real-world environments. The second task result may include a second success rate of the multiple second navigation tasks, and the second task result may also include a first collision count indicating the number of collision events that occurred in the multiple second navigation tasks. The preset conditions may include, for example, that the first collision count is less than a first collision count threshold and the second success rate is greater than a success rate threshold. The values of the first collision count threshold and the success rate threshold can be adjusted according to the actual application, and this application does not impose any restrictions in this regard.
[0053] In this case, when the mobile platform performs the second navigation task based on the optimized policy network, if the success rate of the second navigation task is high and the number of collision events is low, it indicates that the optimized policy network performs relatively stably in the real environment, and the optimized policy network integrated in the mobile platform can be solidified.
[0054] Considering that in practical applications, when the number of real-world environments is small, it may be impossible to fully verify the optimized policy network, leading to an artificially high second success rate. Therefore, in some embodiments, there are multiple simulation environments, and the first navigation task includes multiple first navigation tasks corresponding to each of the simulation environments. The first task result includes, for example, a first success rate for each of the multiple first navigation tasks. Similarly, there are multiple real-world environments, and the second navigation task includes multiple second navigation tasks corresponding to each of the real-world environments. The second task result includes, for example, a second success rate for each of the multiple second navigation tasks. The second task result also includes a first collision count indicating the number of collision events occurring in the multiple second navigation tasks. Preset conditions include, for example, that the first collision count is less than a second collision count threshold, and the difference between the first success rate and the second success rate is less than a first difference threshold. The value of the first difference threshold can be adjusted according to the actual application, and this application does not impose any restrictions in this regard.
[0055] The success rate of the third navigation task is used to indicate the proportion of successful tasks in at least one third navigation task out of the total number of tasks. For example, the criteria for determining whether a third navigation task is successful include: after the completion of a third navigation task, if the distance between the digital twin platform and the third target location is less than a distance threshold, the completion time of the third navigation task is less than a time threshold, and / or the collision rate during the third navigation task is less than a collision rate threshold, then the third navigation task is considered successful; otherwise, the third navigation task is considered a failure. The values of the distance threshold, time threshold, and collision rate threshold can be adjusted according to actual application conditions, and this application does not impose any restrictions in this regard.
[0056] In this scenario, when the mobile platform executes the second navigation task based on the optimized policy network, if the difference between the first success rate and the second success rate is less than the first difference threshold, and the number of collision events is low, it indicates that the simulation environment has fully covered the physical, geometric, and dynamic distribution of the real scene, and the optimized policy network has quantifiable generalization capabilities for unknown real environments. Therefore, the optimized policy network integrated into the mobile platform can be solidified.
[0057] When the result of the second task does not meet the preset conditions, the optimized policy network needs to be optimized again to ensure the stability and reliability of the mobile platform's autonomous navigation capability. Furthermore, considering that the reasons for the second task failing to meet the preset conditions vary under different circumstances, the strategies or methods for further optimizing the optimized policy network will also differ. Therefore, differentiated processing can be applied according to different situations.
[0058] For example, if the result of the second task does not meet the preset conditions, and the difference between the first success rate and the second success rate is less than the second difference threshold, the optimized policy network can be optimized again based on the result of the second task. The second difference threshold is greater than the first difference threshold.
[0059] In this case, the control method for the mobile platform provided in this application embodiment further includes: obtaining the third task results after the mobile platform executes multiple third navigation tasks, wherein the multiple third navigation tasks are tasks in which the mobile platform moves towards a third target location in multiple real environments based on a further optimized policy network; and controlling the mobile platform to execute navigation operations according to the further optimized policy network when the difference between the first success rate and the third success rate in the third task results is less than a first difference threshold, and the second collision count in the third task results is less than a second collision count threshold. Wherein, the third success rate is used to indicate the success rate of the multiple third navigation tasks, and the second collision count is used to indicate the number of collision events that occur in the multiple third navigation tasks.
[0060] In the above method, when the result of the second task does not meet the preset conditions, if the difference between the first success rate and the second success rate is less than the second difference threshold, it means that the performance of the mobile platform in the real environment and the performance of the digital twin platform in the simulation environment have only a small range of distributional deviations, such as slight changes in ground friction and sensor zero drift. The simulation environment is already representative to a certain extent and does not need to be significantly modified. It can be quickly corrected by the incremental data generated by the mobile platform in the real environment, which significantly reduces the deployment time of the policy network and the cost of on-site debugging.
[0061] Alternatively, the simulation environment includes environmental parameters for describing the physical quantities of the interaction between the digital twin platform and the simulation environment; the control method of the mobile platform provided in this application embodiment further includes: if the result of the second task does not meet the preset conditions, and if the difference between the first success rate and the second success rate is greater than or equal to the second difference threshold, the randomness of the simulation environment is increased by adjusting the environmental parameters.
[0062] If the result of the second task does not meet the preset conditions, and the difference between the first success rate and the second success rate is greater than or equal to the second difference threshold, it indicates that the simulation environment does not adequately cover the real environment, resulting in poor adaptability of the optimized policy network to unknown real environments. Therefore, increasing the randomness of the simulation environment can improve its coverage of unknown real environments, thereby improving the reliability and stability of the policy network.
[0063] The technical solution provided in this application, on the one hand, by constructing a digital twin platform for the mobile platform and ensuring that the digital twin platform and the mobile platform have the same or similar responses to the same state vector, allows the test scenario of the policy network to be migrated to a simulation environment. This fully leverages the convenience and variability of setting up the simulation environment to optimize the policy network, which helps reduce the optimization cost and improve optimization efficiency. On the other hand, after optimizing the policy network, the method provided in this application's embodiments further examines the performance of the optimized policy network in the real environment based on the second task result of the mobile platform performing a second navigation task in a real environment using the optimized policy network. This helps improve the policy network's adaptability to the real environment and ensures its reliability and stability.
[0064] In some other possible implementations, this application also provides a control device for a mobile platform. Figure 3 This is a schematic diagram of the structure of the control device for the mobile platform provided in this application embodiment. See also... Figure 3 The control device for the mobile platform provided in this application embodiment includes an optimization module 310, an acquisition module 320, and a navigation module 330.
[0065] The optimization module 310 is configured to optimize the policy network based on the first task result after the digital twin platform of the mobile platform performs the first navigation task. When the digital twin platform and the mobile platform have the same input state vector, the output action vector is less than the error threshold. The policy network is used to map the state vector to the action vector. The first navigation task is the task of the digital twin platform moving towards the first target position in the simulation environment based on the policy network.
[0066] The acquisition module 320 is configured to acquire the second task result after the mobile platform performs the second navigation task, wherein the second navigation task is the task in which the mobile platform moves toward a second target location in a real environment based on an optimized policy network.
[0067] The navigation module 330 is configured to control the mobile platform to perform navigation operations according to the optimized strategy network when the result of the second task meets preset conditions.
[0068] In some possible implementations, there are multiple real environments, the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments, the second task result includes a second success rate of the multiple second navigation tasks, and the second task result also includes a first collision count for indicating the occurrence of collision events in the multiple second navigation tasks; the preset conditions include the first collision count being less than a first collision count threshold, and the second success rate being greater than a success rate threshold.
[0069] In some possible implementations, there are multiple simulation environments, and the first navigation task includes multiple first navigation tasks corresponding to the multiple simulation environments. The first task result includes a first success rate of the multiple first navigation tasks. There are multiple real environments, and the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments. The second task result includes a second success rate of the multiple second navigation tasks. The second task result also includes a first number of collisions to indicate the occurrence of collision events in the multiple second navigation tasks. The preset conditions include that the first number of collisions is less than a second number threshold, and the difference between the first success rate and the second success rate is less than a first difference threshold.
[0070] In some possible implementations, the optimization module 310 is further configured to, if the difference between the first success rate and the second success rate is less than a second difference threshold, further optimize the optimized policy network based on the second task result when the second task result does not meet the preset condition; wherein the second difference threshold is greater than the first difference threshold.
[0071] In some possible implementations, the acquisition module 320 is further configured to acquire the third task results after the mobile platform executes multiple third navigation tasks, wherein the multiple third navigation tasks are tasks in which the mobile platform moves towards a third target location in multiple real environments based on a further optimized policy network; the navigation module 330 is further configured to control the mobile platform to perform navigation operations according to the further optimized policy network when the difference between the first success rate and the third success rate in the third task results is less than the first difference threshold, and the second collision count in the third task results is less than the second collision count threshold; wherein the third success rate is used to indicate the success rate of the multiple third navigation tasks, and the second collision count is used to indicate the number of collision events that occur in the multiple third navigation tasks.
[0072] In some possible implementations, the simulation environment includes environmental parameters for describing the physical quantities of the interaction between the digital twin platform and the simulation environment; the optimization module 310 is further configured to increase the randomness of the simulation environment by adjusting the environmental parameters if the difference between the first success rate and the second success rate is greater than or equal to the second difference threshold when the result of the second task does not meet the preset conditions.
[0073] In some possible implementations, when the optimization module 310 optimizes the policy network based on the first task result after the mobile platform's digital twin platform performs the first navigation task in the simulation environment, it is configured to: correct the digital twin platform based on the hardware error of the mobile platform; and optimize the policy network based on the first task result after the corrected digital twin platform performs the first navigation task.
[0074] It should be understood that the control device for the mobile platform and the control method for the mobile platform provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the control method for the mobile platform embodiments.
[0075] In some other possible implementations, this application also provides an electronic device for controlling a mobile platform. Figure 4 This is a schematic diagram of the structure of an electronic device for controlling a mobile platform provided in an embodiment of this application. See also... Figure 4 The electronic device for controlling a mobile platform provided in this application embodiment includes:
[0076] The memory 410 stores at least one program instruction for the mobile platform.
[0077] When the processor 420 executes the above program instructions, it enables the mobile platform to achieve the above-mentioned combination. Figure 2 The steps of the described control method and its various embodiments are described below. Depending on the implementation, the processor 420 may be one or more types of processors, including but not limited to DSP (digital signal processor), ASIC (application specific integrated circuit), FPGA (field-programmable gate array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and the number of such devices can be determined according to actual needs.
[0078] In some other possible implementations, this application also provides a computer program (product) comprising computer programs / instructions, which are executed by a processor to enable a mobile platform to implement the above-mentioned combination. Figure 2 The steps of the described control method and its various embodiments.
[0079] In some other possible embodiments, this application also provides a computer-readable storage medium storing program instructions for controlling a mobile platform, which, when executed by one or more processors, cause the mobile platform to perform the above-described combination. Figure 2 The steps of the described control method and its various embodiments are described. The computer-readable storage medium can be a readable signal medium or a readable storage medium. A readable storage medium can be, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0080] It should also be noted that the terms "first," "second," etc. (if applicable) in the specification and claims of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0081] The term "and / or" in the embodiments of this application is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0082] The above description is only for the purpose of enabling those skilled in the art to understand the technical solution of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A control method for a mobile platform, characterized in that, The method includes: The policy network is optimized based on the first task result after the digital twin platform of the mobile platform performs the first navigation task. When the input state vectors of the digital twin platform and the mobile platform are the same, the output action vector is less than the error threshold. The policy network is used to map the state vector to the action vector. The first navigation task is the task of the digital twin platform moving towards the first target position in the simulation environment based on the policy network. Obtain the second task result after the mobile platform executes the second navigation task, wherein the second navigation task is the task of the mobile platform moving towards the second target position in a real environment based on the optimized policy network; If the result of the second task meets the preset conditions, the mobile platform is controlled to perform navigation operations according to the optimized strategy network.
2. The method according to claim 1, characterized in that, The real environment is multiple, the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments, the second task result includes the second success rate of the multiple second navigation tasks, and the second task result also includes the first number of collisions used to indicate the number of collision events that occurred in the multiple second navigation tasks. The preset conditions include that the first number of collisions is less than the threshold for the first collision, and the second success rate is greater than the success rate threshold.
3. The method according to claim 1, characterized in that, The simulation environment is multiple, the first navigation task includes multiple first navigation tasks corresponding to the multiple simulation environments, and the first task result includes the first success rate of the multiple first navigation tasks. The real environment is multiple, the second navigation task includes multiple second navigation tasks corresponding to the multiple real environments, the second task result includes the second success rate of the multiple second navigation tasks, and the second task result also includes the first number of collisions used to indicate the number of collision events that occurred in the multiple second navigation tasks. The preset conditions include that the first number of collisions is less than the second number threshold, and the difference between the first success rate and the second success rate is less than the first difference threshold.
4. The method according to claim 3, characterized in that, The method further includes: If the result of the second task does not meet the preset condition, and the difference between the first success rate and the second success rate is less than the second difference threshold, the optimized strategy network is optimized again based on the result of the second task. Wherein, the second difference threshold is greater than the first difference threshold.
5. The method according to claim 4, characterized in that, The method further includes: The third task result is obtained after the mobile platform executes multiple third navigation tasks, wherein the multiple third navigation tasks are tasks in which the mobile platform moves toward a third target location in multiple real environments based on a further optimized policy network. If the difference between the first success rate and the third success rate in the third task result is less than the first difference threshold, and the second number of collisions in the third task result is less than the second number threshold, the mobile platform is controlled to perform navigation operations according to the further optimized strategy network. The third success rate is used to indicate the success rate of the plurality of third navigation tasks, and the second collision count is used to indicate the number of collision events that occur in the plurality of third navigation tasks.
6. The method according to claim 4, characterized in that, The simulation environment includes environmental parameters used to describe the physical quantities of the interaction between the digital twin platform and the simulation environment; The method further includes: If the result of the second task does not meet the preset conditions, and if the difference between the first success rate and the second success rate is greater than or equal to the second difference threshold, the randomness of the simulation environment is increased by adjusting the environmental parameters.
7. The method according to any one of claims 1-6, characterized in that, The optimization of the policy network based on the first task result after the mobile platform's digital twin platform executes the first navigation task in the simulation environment includes: The digital twin platform is calibrated based on the hardware errors of the mobile platform; The policy network is optimized based on the result of the first task after the first navigation task is executed by the corrected digital twin platform.
8. A control device for a mobile platform, characterized in that, The device includes an optimization module, an acquisition module, and a navigation module; The optimization module is configured to optimize the policy network based on the first task result after the digital twin platform of the mobile platform performs the first navigation task. When the digital twin platform and the mobile platform have the same input state vector, the output action vector is less than the error threshold. The policy network is used to map the state vector to the action vector. The first navigation task is the task of the digital twin platform moving towards the first target position in the simulation environment based on the policy network. The acquisition module is configured to acquire the second task result after the mobile platform executes the second navigation task, wherein the second navigation task is the task in which the mobile platform moves toward the second target location in a real environment based on the optimized policy network; The navigation module is configured to control the mobile platform to perform navigation operations according to the optimized strategy network when the result of the second task meets preset conditions.
9. An electronic device, characterized in that, include: The memory stores program instructions for controlling the mobile platform; as well as A processor, when the program instructions are executed by the processor, causes the mobile platform to implement the method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores program instructions for controlling a mobile platform, which, when executed by one or more processors, cause the mobile platform to perform the method described in any one of claims 1-7.