Robot autonomous task planning and execution control method based on reinforcement learning
By combining digital twin models and reinforcement learning strategies, the robot's task planning and execution are optimized, solving the problems of mechanical wear and virtual model deviation, and achieving efficient and accurate task execution and dynamic environment adaptation.
Patent Information
- Application Number
- CN202511221374.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-08-29
AI Technical Summary
In existing technologies, it is difficult to optimize mechanical wear and task efficiency in conjunction with robot task planning and execution. There are deviations between virtual models and physical execution. Reinforcement learning strategies are not adaptable enough in dynamic environments, resulting in shortened equipment lifespan and increased strategy failure rate.
By collecting multi-source data and using a digital twin model for dynamic mapping and loss parameter calibration, a virtual entity state matching dataset is generated. Reinforcement learning strategies are then used to train the strategy in the calibrated simulation environment, generating and selecting the optimal execution strategy to drive the robot to execute. Finally, the weight parameters of the reinforcement learning network are optimized through incremental training.
Reduce mechanical wear and tear, minimize downtime and maintenance, improve production efficiency, enhance the accuracy of entity execution of virtual pre-simulation strategies, strengthen the adaptability of reinforcement learning strategies in dynamic environments, and shorten the strategy adaptation cycle.
Smart Images

Figure CN121004603A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robot control, in particular to a robot autonomous task planning and execution control method based on reinforcement learning. BACKGROUND
[0002] Robots are devices that automatically perform work. In modern industry, robots refer to artificial machine devices that can automatically run tasks, and are used to replace or cooperate with human work. With the increasing number of robots, the application fields are also increasing, and the demand for robots is gradually increasing, requiring robots to not only perform tasks more quickly, but also perform tasks with higher accuracy.
[0003] In the prior art, when robots plan and execute tasks, they only focus on task completion efficiency, without considering the real-time wear and tear of mechanical actuators in the decision model, resulting in shortened equipment life and increased maintenance costs after long-term high-frequency task execution. The physical parameters of the virtual model are mostly static presets and cannot be dynamically adjusted according to the real-time state of the entity robot, resulting in significant deviations between virtual rehearsal strategies and entity execution results, and reduced strategy practicality. Moreover, traditional reinforcement learning strategies are mostly trained based on fixed environment models, and when the environment or state of the entity robot changes dynamically, the pre-trained strategy is difficult to quickly adapt, resulting in an increase in task failure rate. SUMMARY
[0004] The purpose of the present application is to provide a robot autonomous task planning and execution control method based on reinforcement learning to solve the problem that in the prior art, mechanical wear and tear and task efficiency are difficult to optimize simultaneously when robots plan and execute tasks, there is a deviation between virtual model rehearsal strategies and execution results, and reinforcement learning strategies lack adaptability in dynamic environments.
[0005] To achieve the above purpose, the present application provides a robot autonomous task planning and execution control method based on reinforcement learning, which comprises: collecting multi-source data according to a sensor module carried by a robot, and preprocessing the multi-source data to generate an entity state original data set; based on the entity state original data set, using a digital twin model to perform dynamic mapping and wear and tear parameter calibration to generate a virtual entity state matching data set; based on the virtual entity state matching data set, using a preset reinforcement learning strategy rehearsal mechanism, taking the calibrated digital twin model as a simulation environment, performing strategy training to generate a candidate strategy set; using a preset strategy screening mechanism to sort the candidate strategy set, determining an optimal execution strategy, and driving the robot to execute the optimal execution strategy to obtain entity execution data, thereby realizing robot autonomous task planning and execution.
[0006] Optionally, the multi-source data comprises mechanical wear perception data and kinematics and task state data, and the preprocessing of the multi-source data comprises: based on the multi-source data, performing filtering processing and eliminating noise of the multi-source data; and packaging the multi-source data after filtering processing and noise elimination according to a preset data format to obtain an entity state original data set.
[0007] Optionally, the dynamic mapping and wear parameter calibration by using the digital twin model comprises: analyzing the kinematics and task state data by using the digital twin model, and driving a geometric structure model in a virtual space of the digital twin model to perform synchronous motion; and extracting the mechanical wear perception data, and substituting the mechanical wear perception data into a physical property model and a mechanical wear model in the virtual space of the digital twin model to perform wear parameter calibration.
[0008] Optionally, the wear parameter calibration comprises: adjusting a friction coefficient of a virtual joint in the physical property model and correcting a stiffness parameter of a virtual component in the physical property model by using a first preset algorithm; and updating a cumulative wear amount of a virtual actuator in the mechanical wear model by using a second preset algorithm.
[0009] Optionally, the strategy training in the simulation environment by using the calibrated digital twin model by using a preset reinforcement learning strategy rehearsal mechanism comprises: defining a state space based on the virtual entity state matching data set; constructing an initial value of an action space based on a physical constraint of the robot, and iteratively generating a control amount of the action space in the simulation environment by using a reinforcement learning algorithm; calculating a task efficiency score and a mechanical wear penalty value of a reward function based on the kinematics and task state data and the mechanical wear perception data; and performing multi-round strategy training based on the state space, the control amount of the action space, the task efficiency score and the mechanical wear penalty value to generate a candidate strategy set.
[0010] Optionally, the sorting of the candidate strategy set to determine an optimal execution strategy comprises: determining a candidate strategy with the highest task efficiency score as the optimal execution strategy by using a preset strategy screening mechanism; and if the task efficiency scores of multiple candidate strategies are the same, introducing a wear priority coefficient to perform secondary screening to determine a candidate strategy with a low wear amount as the optimal execution strategy.
[0011] Optionally, the driving of the robot to execute the optimal execution strategy to obtain entity execution data comprises: transmitting the determined optimal execution strategy to an execution control module of the robot in the form of a control instruction set; the execution control module of the robot executes the optimal execution strategy, a mechanical wear perception module of the robot records an execution process, and entity execution data is generated; and the entity execution data is transmitted back to the digital twin model and compared with data of virtual rehearsal of the digital twin model.
[0012] Optionally, the control method further comprises: calculating a parameter difference between the candidate strategy set and the entity execution data, triggering a model correction mechanism, correcting the digital twin model, and generating a model correction parameter table; using the reinforcement learning model, calling the model correction parameter table, and combining the entity execution data, incrementally training the reinforcement learning model to generate an optimized candidate strategy set, so as to realize iterative optimization of autonomous task planning and execution of the robot.
[0013] Optionally, the calculation of the parameter difference between the candidate strategy set and the entity execution data, the triggering of the model correction mechanism, and the correction of the digital twin model comprise: if the calculated parameter difference exceeds a preset threshold, triggering the model correction mechanism; using a bias parameter correction formula in the model correction mechanism to adjust the associated parameters of the physical property model in the digital twin model; and based on the deviation between the robot wear and the virtual wear of the digital twin model, adjusting the associated parameters of the mechanical wear model in the digital twin model.
[0014] Optionally, the incremental training of the reinforcement learning model to generate the optimized candidate strategy set comprises: using an incremental training algorithm to increase the training weight corresponding to the working condition with a large deviation value based on the deviation value between the entity execution data and the model correction parameter table; adjusting the calculation coefficient of the loss penalty value in the reward function to make the virtual pre-play wear characteristics of the digital twin model fit the robot wear characteristics, and generating the optimized candidate strategy set.
[0015] Through the above technical solutions, the real-time wear of the mechanical actuator of the robot is included in the decision model by collecting mechanical wear perception data, the mechanical wear of the robot is reduced while the robot is completing a task, the number of downtime maintenance due to excessive wear is reduced, and the production efficiency is improved; the consistency of the virtual and entity states is ensured by dynamically calibrating the physical property parameters of the digital twin model, the entity execution accuracy of the virtual pre-play strategy is improved, the strategy trial and error cost caused by model deviation is reduced, and the credibility of virtual simulation is significantly improved; the weight parameters of the reinforcement learning network are updated by the incremental training algorithm, the adaptation speed of the reinforcement learning strategy to the dynamic environment is improved, the sample demand of strategy retraining is reduced, and the strategy adaptation cycle in a new scenario is shortened.
[0016] Other features and advantages of the present application will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0017] The accompanying drawings are included to provide a further understanding of the embodiments of the application, and constitute a part of the specification, and are used together with the following specific embodiments to explain the embodiments of the application, but do not constitute a limitation on the embodiments of the application. In the drawings: Figure 1is a flowchart of the robot autonomous task planning and execution control method based on reinforcement learning of the present application Figure 1 ; Figure 2 is a flowchart of the dynamic mapping and loss parameter calibration in the present application Figure 3 is a flowchart of the strategy training in the present application Figure 4 is a flowchart of the driving robot to execute the optimal execution strategy in the present application Figure 5 is a flowchart of the digital twin model correction in the present application Figure 6 is a flowchart of the generation of the optimization candidate strategy set in the present application Figure 7 is a flowchart of the robot autonomous task planning and execution control method based on reinforcement learning of the present application Figure 2 . DETAILED DESCRIPTION
[0018] The specific embodiments of the present application will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application.
[0019] It should be noted that the acquisition, transmission, storage, use, processing, etc. of data in the technical solutions of the present application comply with the relevant provisions of national laws and regulations. In the embodiments of the present application, some industry existing solutions, components, models, etc. may be mentioned, which should be considered as exemplary, and the purpose is only to illustrate the feasibility of the implementation of the technical solutions of the present application, but it does not mean that the applicant has or will necessarily use the solutions.
[0020] Please refer to Figure 1 , the embodiments of the present application provide a robot autonomous task planning and execution control method based on reinforcement learning, which comprises: Step S110: According to the sensor module carried by the robot, collect multi-source data, and pre-process the multi-source data to generate entity state original data set.
[0021] In combination with Figure 7 , in the embodiments of the present application, the multi-source data includes mechanical loss perception data and kinematics and task state data, and the pre-processing of the multi-source data can include: Step S111: Based on the multi-source data, filtering processing is performed, and the noise of the multi-source data is eliminated.
[0022] In the preferred embodiments of the present application, the mechanical wear-aware data can include real-time temperature, vibration frequency, torque fluctuation value, lubricant state parameters, etc. of the robot actuators (e.g., joints, motors), and the kinematics and task state data can include position coordinates, motion speed, acceleration, task progress parameters, etc. of the robot end effector (e.g., grasping success rate, path completion degree).
[0023] Step S112: The filtered and noise-removed multi-source data is packaged according to a preset data format (e.g., timestamp, sensor ID, parameter type and value) to obtain an entity state raw data set.
[0024] Step S120: Based on the entity state raw data set, a digital twin model is used to perform dynamic mapping and wear parameter calibration to generate a virtual entity state matching data set.
[0025] Please refer to Figure 2 In the embodiments of the present application, the dynamic mapping and wear parameter calibration using the digital twin model can include: Step S121: The kinematics and task state data (e.g., end effector position, joint rotation angle, etc. geometric and kinematics parameters) are analyzed using the digital twin model, and the geometric structure model in the virtual space of the digital twin model is driven to perform synchronous motion (e.g., the virtual robot arm joints are rotated in real time according to the entity position coordinates). Step S122: The mechanical wear-aware data (e.g., temperature, vibration, etc.) is extracted and substituted into the physical property model and mechanical wear model in the virtual space of the digital twin model to perform wear parameter calibration.
[0026] In the preferred embodiments of the present application, the wear parameter calibration can include: Step S1221: The friction coefficient of the virtual joint in the physical property model is adjusted using a first preset algorithm, and the stiffness parameter of the virtual component in the physical property model is corrected.
[0027] In the preferred embodiments of the present application, for the physical property model, the friction coefficient of the virtual joint can be adjusted by a temperature-material property correlation algorithm (an algorithm for the quantitative relationship between material properties and temperature, which describes the change law of material properties with temperature by model or data driven method, inputs temperature parameter, and outputs material property parameter at corresponding temperature) (e.g., the friction coefficient increases by 0.02 for every 10℃ increase in temperature), and the stiffness parameter of the virtual component is corrected by a vibration frequency-structure stiffness conversion formula.
[0028] Step S1222: The cumulative wear amount of the virtual actuator in the mechanical wear model is updated using a second preset algorithm.
[0029] In a preferred embodiment of the present invention, for the mechanical wear model, the torque fluctuation value (the value of the periodic or non-periodic fluctuation of the output torque during mechanical operation as it changes with time or rotation angle) can be input into the wear accumulation function (which reflects the cumulative effect of wear amount with the action conditions, usually established based on the stage characteristics of the wear curve, and used to predict the wear development trend and remaining life of materials or parts under specific working conditions), and the cumulative wear amount of the virtual actuator is updated in real time. The calibrated digital twin model forms a virtual entity state matching dataset, which includes the correspondence between virtual parameters and entity parameters (for example, a virtual joint temperature of 41.8°C corresponds to an entity temperature of 42°C).
[0030] Step S130: Based on the virtual entity state matching dataset, use the preset reinforcement learning policy pre-playing mechanism, with the calibrated digital twin model as the simulation environment, to train the policy and generate a candidate policy set.
[0031] Please refer to Figure 3 In this embodiment of the invention, policy training is performed using a preset reinforcement learning policy pre-training mechanism and a calibrated digital twin model as the simulation environment. This may include: Step S131: Define the state space based on the virtual entity state matching dataset.
[0032] In a preferred embodiment of the present invention, the definition of the state space directly references the real-time parameters of the virtual model in the virtual entity state matching dataset, which may include virtual joint position, velocity, real-time friction coefficient, cumulative wear, etc.
[0033] Step S132: Based on the robot's physical constraints, construct the initial values of the motion space, and use reinforcement learning algorithms to iteratively generate the control variables of the motion space in the simulation environment.
[0034] In a preferred embodiment of the present invention, the initial value of the motion space can be set based on the physical constraints of the physical robot (e.g., the maximum joint rotation angle ±90°), and specific control quantities (e.g., the joint rotation angle increment) can be generated iteratively in the virtual environment by a reinforcement learning algorithm.
[0035] Step S133: Based on kinematic and task state data and mechanical wear perception data, calculate the task efficiency score of the reward function (a quantitative indicator that measures the efficiency of task completion; a higher score indicates better task execution efficiency, usually derived from the path length and time of the virtual model to complete the task) and the mechanical wear penalty value (energy loss caused by friction, vibration, etc. during equipment operation, usually derived from the virtual wear accumulation of the virtual model).
[0036] Step S134: Based on the control variables of the state space and action space, the task efficiency score, and the mechanical loss penalty value, perform multiple rounds of policy training (generating a set of policy parameters in each round of trial and error) to generate a candidate policy set.
[0037] For example, based on the calibrated virtual robotic arm model, when the reinforcement learning algorithm tests the strategy of "joint rotation angular velocity 1.0 rad / s", the virtual robotic arm model reports a path length of 2m, a completion time of 8s, and a wear accumulation of 0.001mm, with a calculated score of 60 (task efficiency score 80 - mechanical wear penalty value 20). When testing the strategy of "angular velocity 0.8 rad / s", the path length is 2.2m, the time is 10s, the wear accumulation is 0.0005mm, and the score is 60 (task efficiency score 70 - mechanical wear penalty value 10). Finally, a candidate set containing these two strategies is generated.
[0038] Step S140: Using a preset strategy filtering mechanism, sort the candidate strategy set, determine the optimal execution strategy, and drive the robot to execute the optimal execution strategy to obtain entity execution data, so as to realize the robot's autonomous task planning and execution.
[0039] In this embodiment of the invention, sorting the candidate strategy set to determine the optimal execution strategy may include: using a preset strategy screening mechanism (a systematic rule, process, or method for selecting strategies that meet specific conditions or objectives from multiple candidate strategies) to determine the candidate strategy with the highest task efficiency score as the optimal execution strategy; if multiple candidate strategies have the same task efficiency score, a wear priority coefficient is introduced for secondary screening to determine the candidate strategy with low wear as the optimal execution strategy; the selected optimal execution strategy is transmitted to the execution control module of the physical robot in the form of a control instruction set (e.g., "Joint 1 rotation angle +30°, speed 0.8 rad / s; Joint 2 rotation angle -15°, speed 0.6 rad / s") via a secure encryption protocol.
[0040] Please refer to Figure 4 In this embodiment of the invention, driving the robot to execute the optimal execution strategy to obtain entity execution data may include: Step S141: Transmit the determined optimal execution strategy to the robot's execution control module in the form of a control instruction set.
[0041] Step S142: The robot's execution control module executes the optimal execution strategy, the robot's mechanical loss sensing module records the execution process, and generates entity execution data.
[0042] Step S143: Send the entity execution data back to the digital twin model and compare it with the data from the virtual pre-simulation of the digital twin model.
[0043] For example, in the candidate strategy set, the strategy of "angular velocity 0.8 rad / s" was selected as the optimal strategy because of its lower wear (0.0005 mm < 0.001 mm). Its control command "joint 1 rotates to the 60° position at 0.8 rad / s" was sent to the physical robotic arm. During the execution, the physical sensors collected data showing that the actual joint temperature rose from 42°C to 45°C and the vibration peak was 22 Hz. These data were transmitted back to the digital twin model in real time and compared with the data in the virtual pre-simulation, which showed that "the temperature rose to 44°C and the vibration peak was 20 Hz".
[0044] Step S150: Calculate the parameter difference between the candidate strategy set and the entity execution data, trigger the model correction mechanism, correct the digital twin model, and generate a model correction parameter table.
[0045] Please refer to Figure 5 In this embodiment of the invention, calculating the parameter difference between the candidate strategy set and the entity execution data, triggering the model correction mechanism, and correcting the digital twin model may include: Step S151: If the calculated parameter difference (e.g., temperature deviation = physical temperature - virtual temperature, vibration deviation = physical vibration frequency - virtual vibration frequency) exceeds the preset threshold (e.g., temperature deviation > 2℃, vibration deviation > 5Hz), the model correction mechanism is triggered.
[0046] Step S152: Using the deviation parameter correction formula in the model correction mechanism (to adjust or calibrate the deviation parameters existing in the system to improve the accuracy of the model, measurement or calculation results), adjust the correlation parameters of the physical property model in the digital twin model (for example, when the temperature deviation is 3℃, correct the coefficient in the temperature-friction coefficient correlation algorithm from 0.02 / 10℃ to 0.025 / 10℃).
[0047] Step S153: Based on the deviation between the robot wear amount (e.g., which can be obtained through periodic inspection of the physical robot) and the virtual wear amount in the digital twin model, adjust the correlation parameters of the mechanical loss model in the digital twin model (e.g., adjust the weight coefficient of the wear accumulation function; if the robot wear amount is 20% higher than the virtual value, adjust the coefficient of the square term of torque fluctuation from 1.0 to 1.2).
[0048] For example, after comparing the physical robot data with the virtual pre-simulation data, the temperature deviation was 1℃ (derived from 45℃-44℃) and the vibration deviation was 2Hz (derived from 22Hz-20Hz), both of which did not exceed the threshold (temperature 2℃, vibration 5Hz), but the system still recorded the deviation trend; when the fifth execution, the cumulative temperature deviation reached 3℃, triggering the model correction mechanism, which further adjusted the coefficient of the temperature-friction coefficient correlation algorithm in the digital twin model from 0.025 / 10℃ to 0.03 / 10℃, ensuring that subsequent virtual temperature predictions are more accurate.
[0049] Step S160: Using the reinforcement learning model, call the model correction parameter table and combine it with entity execution data to incrementally train the reinforcement learning model and generate an optimized candidate policy set to achieve iterative optimization of the robot's autonomous task planning and execution.
[0050] In a preferred embodiment of the present invention, the reinforcement learning model can call the model correction parameter table, use the corrected digital twin model as a new simulation environment, and add entity execution data (e.g., actual task completion time, actual wear and tear, etc.) to the training dataset to form virtual-entity fusion training samples for subsequent incremental training.
[0051] Please refer to Figure 6 In this embodiment of the invention, incremental training of the reinforcement learning model to generate an optimized candidate policy set may include: Step S161: Using an incremental training algorithm (the model is iteratively updated by gradually adding new data on the basis of existing training, without having to retrain from scratch using all historical data), the training weights corresponding to the conditions with large deviation values (e.g., high vibration conditions correspond to large deviations) are increased based on the deviation values between the entity execution data and the model correction parameter table.
[0052] Step S162: Adjust the calculation coefficient of the loss penalty value in the reward function so that the virtual pre-simulation loss characteristics of the digital twin model match the loss characteristics of the robot (for example, if the physical wear is more sensitive to vibration, increase the proportion of vibration parameters in the loss penalty value) to generate an optimized candidate strategy set.
[0053] For example, the revised digital twin model improves the accuracy of vibration simulation. Combined with the data recorded in the entity execution data that "the entity wears out faster when vibrating at 22Hz", the reinforcement learning model increases the weight of the sample in this scenario during incremental training (e.g., from 1.0 to 1.5) and adjusts the penalty coefficient of the vibration parameter in the reward function from 0.1 to 0.15. In the virtual pre-simulation, the updated digital twin model will prioritize avoiding motion paths that cause vibrations exceeding 20Hz. When the generated strategy is executed by the physical robot, the vibration peak drops to 19Hz, and the deviation from the virtual pre-simulation is reduced to 1Hz.
[0054] Accordingly, this invention provides a robot autonomous task planning and execution control method based on reinforcement learning. The control method includes: collecting multi-source data from the sensor modules mounted on the robot and preprocessing the multi-source data to generate an original dataset of entity states; based on the original dataset of entity states, using a digital twin model, performing dynamic mapping and loss parameter calibration to generate a virtual entity state matching dataset; based on the virtual entity state matching dataset, using a preset reinforcement learning strategy pre-playing mechanism and the calibrated digital twin model as the simulation environment, performing strategy training to generate a candidate strategy set; using a preset strategy screening mechanism, sorting the candidate strategy set, determining the optimal execution strategy, and driving the robot to execute the optimal execution strategy to obtain entity execution data, thereby realizing the robot's autonomous task planning and execution. By collecting mechanical wear perception data, the real-time wear of the robot's mechanical actuators is incorporated into the decision-making model. This reduces mechanical wear while the robot completes its tasks, minimizing downtime for maintenance due to excessive wear and improving production efficiency. Dynamically calibrating the physical property parameters of the digital twin model ensures consistency between the virtual and physical states, improving the accuracy of the physical execution of the virtual pre-simulation strategy, reducing the cost of strategy trial and error due to model bias, and significantly enhancing the credibility of the virtual simulation. Furthermore, updating the weight parameters of the reinforcement learning network through incremental training algorithms improves the speed at which the reinforcement learning strategy adapts to dynamic environments, reduces the sample requirements for strategy retraining, and shortens the strategy adaptation cycle in new scenarios.
[0055] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0056] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0057] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0058] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0059] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0060] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0061] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0062] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0063] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for autonomous task planning and execution control of robots based on reinforcement learning, characterized in that, The control method includes: Based on the sensor modules mounted on the robot, multi-source data is collected, and the multi-source data is preprocessed to generate the original dataset of entity state. Based on the original dataset of the entity state, a virtual entity state matching dataset is generated by using a digital twin model for dynamic mapping and loss parameter calibration. Based on the virtual entity state matching dataset, a preset reinforcement learning policy pre-playing mechanism is used to train the policy in the simulation environment of the calibrated digital twin model, thereby generating a candidate policy set. By using a preset strategy filtering mechanism, the candidate strategy set is sorted to determine the optimal execution strategy, and the robot is driven to execute the optimal execution strategy to obtain entity execution data, thereby realizing the robot's autonomous task planning and execution.
2. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 1, characterized in that, The multi-source data includes mechanical wear sensing data and kinematic and task status data. The preprocessing of the multi-source data includes: Based on the multi-source data, filtering is performed to remove noise from the multi-source data; The filtered and noise-removed multi-source data is packaged according to a preset data format to obtain the original dataset of entity states.
3. The reinforcement learning-based robot autonomous task planning and execution control method according to claim 2, characterized in that, The process of using a digital twin model for dynamic mapping and loss parameter calibration includes: Using a digital twin model, the kinematic and task state data are analyzed, and the geometric structure model in the virtual space of the digital twin model is driven to move synchronously. The mechanical loss sensing data is extracted and substituted into the physical attribute model and mechanical loss model in the virtual space of the digital twin model to perform loss parameter calibration.
4. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 3, characterized in that, The loss parameter calibration includes: Using the first preset algorithm, the friction coefficient of the virtual joint in the physical property model is adjusted, and the stiffness parameters of the virtual component in the physical property model are corrected. The cumulative wear of the virtual actuator in the mechanical wear model is updated using the second preset algorithm.
5. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 1, characterized in that, The method of using a pre-defined reinforcement learning policy pre-training mechanism, with a calibrated digital twin model as the simulation environment, to train the policy includes: Based on the virtual entity state matching dataset, a state space is defined; Based on the robot's physical constraints, the initial values of the motion space are constructed, and the control variables of the motion space are iteratively generated in the simulation environment using reinforcement learning algorithms. Based on kinematic and task state data and mechanical wear perception data, the task efficiency score and mechanical wear penalty value of the reward function are calculated. Based on the control variables of the state space and action space, the task efficiency score, and the mechanical loss penalty value, multiple rounds of policy training are performed to generate a candidate policy set.
6. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 1, characterized in that, The step of sorting the candidate strategy set to determine the optimal execution strategy includes: Using a pre-defined strategy selection mechanism, the candidate strategy with the highest task efficiency score is determined as the optimal execution strategy; If multiple candidate strategies have the same task efficiency score, a loss priority coefficient is introduced for secondary screening to determine the candidate strategy with the lowest wear as the optimal execution strategy.
7. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 6, characterized in that, The drive robot executes the optimal execution strategy to obtain entity execution data, including: The determined optimal execution strategy is transmitted to the robot's execution control module in the form of a set of control instructions; The robot's execution control module executes the optimal execution strategy, and the robot's mechanical wear perception module records the execution process and generates entity execution data; The entity execution data is sent back to the digital twin model and compared with the data from the virtual pre-simulation of the digital twin model.
8. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 1, characterized in that, The control method further includes: Calculate the parameter difference between the candidate strategy set and the entity execution data, trigger the model correction mechanism to correct the digital twin model, and generate a model correction parameter table; By utilizing a reinforcement learning model, the model's parameter table is called, and the entity's execution data is combined to incrementally train the reinforcement learning model, generating an optimized candidate policy set to achieve iterative optimization of the robot's autonomous task planning and execution.
9. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 8, characterized in that, The step of calculating the parameter difference between the candidate strategy set and the entity execution data, triggering the model correction mechanism, and correcting the digital twin model includes: If the calculated parameter difference exceeds a preset threshold, the model correction mechanism is triggered. By using the deviation parameter correction formula in the model correction mechanism, the correlation parameters of the physical property model in the digital twin model are adjusted; Based on the deviation between the robot's wear amount and the virtual wear amount in the digital twin model, the correlation parameters of the mechanical loss model in the digital twin model are adjusted.
10. The robot autonomous task planning and execution control method based on reinforcement learning according to claim 8, characterized in that, The incremental training of the reinforcement learning model to generate an optimized candidate policy set includes: Using an incremental training algorithm, the training weights corresponding to the working conditions with large deviation values are increased based on the deviation values between the entity execution data and the model correction parameter table. Adjust the calculation coefficients of the loss penalty value in the reward function to make the virtual pre-simulation loss characteristics of the digital twin model match the loss characteristics of the robot, and generate an optimized candidate strategy set.
Citation Information
Patent Citations
Milling robot cutter wear state real-time monitoring method fusing digital twinning and deep learning
CN118700161A
Robot multi-mode sensing and motion cooperative control method and device
CN120347724A
Method for optimising the control of a device by means of the combination of a learning method with milp, computer program for carrying out the method and corresponding device
EP4424476A1