A Reinforcement Learning-Based Method for Autonomous Task Planning and Execution Control of Robots
By collecting multi-source data for digital twin model calibration and reinforcement learning strategy training, the problems of mechanical loss and efficiency optimization in robot task planning and execution were solved. This achieved consistency between the virtual model and the physical state and policy adaptability in dynamic environments, thereby improving the reliability and efficiency of robot task execution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2026-03-13
AI Technical Summary
In existing technologies, robot task planning and execution fail to effectively coordinate and optimize mechanical wear and task efficiency, there are deviations between virtual models and physical execution, and reinforcement learning strategies are not adaptable enough to dynamic environments, resulting in shortened equipment lifespan and increased strategy failure rate.
By collecting multi-source data, using a digital twin model for dynamic mapping and loss parameter calibration, and combining reinforcement learning strategies to train and generate a set of candidate strategies in a simulation environment, the strategies are optimized through model correction and incremental training to ensure consistency and adaptability between virtual and physical states.
Reduce mechanical wear and tear, minimize downtime and maintenance, improve task execution accuracy and production efficiency, and enhance the adaptability of reinforcement learning strategies in dynamic environments.
Smart Images

Figure CN121004603B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, and more specifically to a method for autonomous task planning and execution control of robots based on reinforcement learning. Background Technology
[0002] A robot is a device that performs tasks automatically. In modern industry, a robot refers to man-made machines that can automatically run tasks to replace or assist human workers. As the types of robots continue to increase, their application areas are also expanding, and users' demands for robots are gradually increasing, requiring robots to not only perform tasks faster but also with higher precision.
[0003] In existing technologies, when robots plan and execute tasks, they often only focus on task completion efficiency and fail to incorporate the real-time wear and tear of mechanical actuators into the decision-making model. This leads to a shortened equipment lifespan and increased maintenance costs after long-term, high-frequency task execution. The physical parameters of virtual models are mostly statically preset and cannot be dynamically adjusted according to the real-time state of the physical robot. This results in a significant deviation between the virtual pre-simulation strategy and the actual execution effect, reducing the practicality of the strategy. Moreover, traditional reinforcement learning strategies are mostly trained based on fixed environment models. When the environment or state of the physical robot changes dynamically, the pre-trained strategy is difficult to adapt quickly, leading to an increased task failure rate. Summary of the Invention
[0004] The purpose of this invention is to provide a robot autonomous task planning and execution control method based on reinforcement learning, in order to solve the problems in the prior art where it is difficult to optimize mechanical loss and task efficiency in robot task planning and execution, there is a deviation between virtual model pre-simulation strategy and execution effect, and reinforcement learning strategy is not adaptable enough in dynamic environment.
[0005] To achieve the above objectives, this invention provides a robot autonomous task planning and execution control method based on reinforcement learning. The control method includes: collecting multi-source data from sensor modules mounted on the robot and preprocessing the multi-source data to generate a raw entity state dataset; based on the raw entity state dataset, performing dynamic mapping and loss parameter calibration using a digital twin model to generate a virtual entity state matching dataset; based on the virtual entity state matching dataset, using a preset reinforcement learning policy pre-training mechanism and the calibrated digital twin model as a simulation environment, training policies to generate a candidate policy set; and using a preset policy filtering mechanism, sorting the candidate policy set to determine the optimal execution policy and driving the robot to execute the optimal execution policy to obtain entity execution data, thereby achieving autonomous task planning and execution for the robot.
[0006] Optionally, the multi-source data includes mechanical loss sensing data and kinematic and task state data. The preprocessing of the multi-source data includes: filtering the multi-source data and removing noise from the multi-source data; and packaging the filtered and noise-removed multi-source data according to a preset data format to obtain the original dataset of entity state.
[0007] Optionally, the step of using a digital twin model for dynamic mapping and loss parameter calibration includes: using the digital twin model to analyze the kinematic and task state data, and driving the geometric structure model in the virtual space of the digital twin model to move synchronously; extracting the mechanical loss perception data, substituting it into the physical attribute model and mechanical loss model in the virtual space of the digital twin model, and calibrating the loss parameters.
[0008] Optionally, the loss parameter calibration includes: using a first preset algorithm to adjust the friction coefficient of the virtual joint in the physical property model and correct the stiffness parameter of the virtual component in the physical property model; and using a second preset algorithm to update the cumulative wear of the virtual actuator in the mechanical loss model.
[0009] Optionally, the step of using a preset reinforcement learning policy pre-training mechanism, with the calibrated digital twin model as the simulation environment, to train the policy includes: defining a state space based on the virtual entity state matching dataset; constructing initial values for the action space based on the robot's physical constraints, and using a reinforcement learning algorithm to iteratively generate control quantities for the action space in the simulation environment; calculating the task efficiency score and mechanical loss penalty value of the reward function based on kinematic and task state data and mechanical loss perception data; and performing multiple rounds of policy training based on the state space, the control quantities of the action space, the task efficiency score, and the mechanical loss penalty value to generate a candidate policy set.
[0010] Optionally, sorting the candidate strategy set to determine the optimal execution strategy includes: using a preset strategy screening mechanism to determine the candidate strategy with the highest task efficiency score as the optimal execution strategy; if multiple candidate strategies have the same task efficiency score, a loss priority coefficient is introduced for secondary screening to determine the candidate strategy with low wear as the optimal execution strategy.
[0011] Optionally, the process of driving the robot to execute the optimal execution strategy to obtain entity execution data includes: transmitting the determined optimal execution strategy to the robot's execution control module in the form of a control instruction set; the robot's execution control module executing the optimal execution strategy, the robot's mechanical wear perception module recording the execution process and generating entity execution data; and transmitting the entity execution data back to the digital twin model for comparison with the data from the virtual pre-simulation of the digital twin model.
[0012] Optionally, the control method further includes: calculating the parameter difference between the candidate strategy set and the entity execution data, triggering a model correction mechanism to correct the digital twin model and generate a model correction parameter table; using a reinforcement learning model, calling the model correction parameter table and combining it with the entity execution data to incrementally train the reinforcement learning model and generate an optimized candidate strategy set, so as to achieve iterative optimization of the robot's autonomous task planning and execution.
[0013] Optionally, the step of calculating the parameter difference between the candidate strategy set and the entity execution data, and triggering the model correction mechanism to correct the digital twin model, includes: if the calculated parameter difference exceeds a preset threshold, triggering the model correction mechanism; adjusting the correlation parameters of the physical attribute model in the digital twin model using the deviation parameter correction formula in the model correction mechanism; and adjusting the correlation parameters of the mechanical loss model in the digital twin model based on the deviation between the robot wear amount and the virtual wear amount of the digital twin model.
[0014] Optionally, the incremental training of the reinforcement learning model to generate an optimized candidate strategy set includes: using an incremental training algorithm, based on the deviation between the entity execution data and the model correction parameter table, increasing the training weights corresponding to the conditions with large deviation values; adjusting the calculation coefficient of the loss penalty value in the reward function so that the virtual pre-simulation loss characteristics of the digital twin model match the robot loss characteristics, thereby generating an optimized candidate strategy set.
[0015] Through the above technical solutions, this invention incorporates the real-time wear of the robot's mechanical actuators into the decision-making model by collecting mechanical wear perception data. This reduces mechanical wear while the robot completes its tasks, minimizing downtime for maintenance due to excessive wear and improving production efficiency. By dynamically calibrating the physical attribute parameters of the digital twin model, the consistency between the virtual and physical states is ensured, improving the accuracy of the entity execution of the virtual pre-simulation strategy, reducing the cost of strategy trial and error due to model bias, and significantly enhancing the credibility of the virtual simulation. Furthermore, by updating the weight parameters of the reinforcement learning network through incremental training algorithms, the speed at which the reinforcement learning strategy adapts to dynamic environments is improved, reducing the sample requirements for strategy retraining and shortening the strategy adaptation cycle in new scenarios.
[0016] Other features and advantages of the present invention will be described in detail in the following detailed description section. Attached Figure Description
[0017] The accompanying drawings are provided to further illustrate embodiments of the present invention and form part of the specification. They are used together with the following detailed description to explain the embodiments of the present invention, but do not constitute a limitation thereof. In the drawings:
[0018] Figure 1This is a flowchart illustrating the robot autonomous task planning and execution control method based on reinforcement learning, as described in this invention. Figure 1 ;
[0019] Figure 2 This is a schematic diagram of the process of dynamic mapping and loss parameter calibration in this invention;
[0020] Figure 3 This is a schematic diagram of the strategy training process in this invention;
[0021] Figure 4 This is a schematic diagram of the process of driving the robot to execute the optimal execution strategy in this invention;
[0022] Figure 5 This is a schematic diagram of the process for modifying the digital twin model in this invention;
[0023] Figure 6 This is a schematic diagram of the process for generating an optimization candidate strategy set in this invention;
[0024] Figure 7 This is a flowchart illustrating the robot autonomous task planning and execution control method based on reinforcement learning, as described in this invention. Figure 2 . Detailed Implementation
[0025] The specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit the scope of the present invention.
[0026] It should be noted that the acquisition, transmission, storage, use, and processing of data in this application comply with relevant national laws and regulations. In the embodiments of this application, certain existing industry solutions such as software, components, and models may be mentioned. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0027] Please refer to Figure 1 This invention provides a reinforcement learning-based method for autonomous task planning and execution control of robots. The control method includes:
[0028] Step S110: Collect multi-source data based on the sensor modules mounted on the robot, and preprocess the multi-source data to generate the original dataset of entity state.
[0029] Combination Figure 7 In this embodiment of the invention, the multi-source data includes mechanical loss sensing data and kinematic and task state data. Preprocessing of the multi-source data may include:
[0030] Step S111: Based on the multi-source data, perform filtering processing and remove noise from the multi-source data.
[0031] In a preferred embodiment of the present invention, the mechanical wear perception data may include the real-time temperature, vibration frequency, torque fluctuation value, lubricant state parameters, etc. of the robot actuator (e.g., joints, motors), and the kinematic and task status data may include the position coordinates, motion speed, acceleration, task progress parameters, etc. of the robot end effector (e.g., grasping success rate, path completion rate).
[0032] Step S112: Pack the filtered and noise-removed multi-source data into a preset data format (e.g., timestamp, sensor ID, parameter type and value) to obtain the original dataset of entity status.
[0033] Step S120: Based on the original dataset of entity states, use the digital twin model to perform dynamic mapping and loss parameter calibration to generate a virtual entity state matching dataset.
[0034] Please refer to Figure 2 In this embodiment of the invention, using a digital twin model for dynamic mapping and loss parameter calibration may include:
[0035] Step S121: Using the digital twin model, analyze the kinematic and task state data (e.g., geometric and kinematic parameters such as end effector position and joint rotation angle), and drive the geometric structure model in the virtual space of the digital twin model to move synchronously (e.g., the virtual robotic arm joints rotate in real time according to the physical position coordinates).
[0036] Step S122: Extract mechanical loss sensing data (e.g., temperature, vibration, etc.), substitute it into the physical attribute model and mechanical loss model in the virtual space of the digital twin model, and perform loss parameter calibration.
[0037] In a preferred embodiment of the present invention, loss parameter calibration may include:
[0038] Step S1221: Using the first preset algorithm, adjust the friction coefficient of the virtual joint in the physical property model and correct the stiffness parameters of the virtual component in the physical property model.
[0039] In a preferred embodiment of the present invention, for the physical property model, the friction coefficient of the virtual joint can be adjusted by a temperature-material property correlation algorithm (an algorithm that describes the quantitative relationship between material properties and temperature, describes the change law of material properties with temperature through a model or data-driven method, inputs temperature parameters, and outputs material property parameters at the corresponding temperature). For example, the friction coefficient increases by 0.02 for every 10°C increase in temperature. The stiffness parameters of the virtual component can be corrected by a vibration frequency-structural stiffness conversion formula.
[0040] Step S1222: Update the cumulative wear of the virtual actuator in the mechanical wear model using the second preset algorithm.
[0041] In a preferred embodiment of the present invention, for the mechanical wear model, the torque fluctuation value (the value of the periodic or non-periodic fluctuation of the output torque during mechanical operation as it changes with time or rotation angle) can be input into the wear accumulation function (which reflects the cumulative effect of wear amount with the action conditions, usually established based on the stage characteristics of the wear curve, and used to predict the wear development trend and remaining life of materials or parts under specific working conditions), and the cumulative wear amount of the virtual actuator is updated in real time. The calibrated digital twin model forms a virtual entity state matching dataset, which includes the correspondence between virtual parameters and entity parameters (for example, a virtual joint temperature of 41.8°C corresponds to an entity temperature of 42°C).
[0042] Step S130: Based on the virtual entity state matching dataset, use the preset reinforcement learning policy pre-playing mechanism, with the calibrated digital twin model as the simulation environment, to train the policy and generate a candidate policy set.
[0043] Please refer to Figure 3 In this embodiment of the invention, policy training is performed using a preset reinforcement learning policy pre-training mechanism and a calibrated digital twin model as the simulation environment. This may include:
[0044] Step S131: Define the state space based on the virtual entity state matching dataset.
[0045] In a preferred embodiment of the present invention, the definition of the state space directly references the real-time parameters of the virtual model in the virtual entity state matching dataset, which may include virtual joint position, velocity, real-time friction coefficient, cumulative wear, etc.
[0046] Step S132: Based on the robot's physical constraints, construct the initial values of the motion space, and use reinforcement learning algorithms to iteratively generate the control variables of the motion space in the simulation environment.
[0047] In a preferred embodiment of the present invention, the initial value of the motion space can be set based on the physical constraints of the physical robot (e.g., the maximum joint rotation angle ±90°), and specific control quantities (e.g., the joint rotation angle increment) can be generated iteratively in the virtual environment by a reinforcement learning algorithm.
[0048] Step S133: Based on kinematic and task state data and mechanical wear perception data, calculate the task efficiency score of the reward function (a quantitative indicator that measures the efficiency of task completion; a higher score indicates better task execution efficiency, usually derived from the path length and time of the virtual model to complete the task) and the mechanical wear penalty value (energy loss caused by friction, vibration, etc. during equipment operation, usually derived from the virtual wear accumulation of the virtual model).
[0049] Step S134: Based on the control variables of the state space and action space, the task efficiency score, and the mechanical loss penalty value, perform multiple rounds of policy training (generating a set of policy parameters in each round of trial and error) to generate a candidate policy set.
[0050] For example, based on the calibrated virtual robotic arm model, when the reinforcement learning algorithm tests the strategy of "joint rotation angular velocity 1.0 rad / s", the virtual robotic arm model reports a path length of 2m, a completion time of 8s, and a wear accumulation of 0.001mm, with a calculated score of 60 (task efficiency score 80 - mechanical wear penalty value 20). When testing the strategy of "angular velocity 0.8 rad / s", the path length is 2.2m, the time is 10s, the wear accumulation is 0.0005mm, and the score is 60 (task efficiency score 70 - mechanical wear penalty value 10). Finally, a candidate set containing these two strategies is generated.
[0051] Step S140: Using a preset strategy filtering mechanism, sort the candidate strategy set, determine the optimal execution strategy, and drive the robot to execute the optimal execution strategy to obtain entity execution data, so as to realize the robot's autonomous task planning and execution.
[0052] In this embodiment of the invention, sorting the candidate strategy set to determine the optimal execution strategy may include: using a preset strategy screening mechanism (a systematic rule, process, or method for selecting strategies that meet specific conditions or objectives from multiple candidate strategies) to determine the candidate strategy with the highest task efficiency score as the optimal execution strategy; if multiple candidate strategies have the same task efficiency score, a wear priority coefficient is introduced for secondary screening to determine the candidate strategy with low wear as the optimal execution strategy; the selected optimal execution strategy is transmitted to the execution control module of the physical robot in the form of a control instruction set (e.g., "Joint 1 rotation angle +30°, speed 0.8 rad / s; Joint 2 rotation angle -15°, speed 0.6 rad / s") via a secure encryption protocol.
[0053] Please refer to Figure 4 In this embodiment of the invention, driving the robot to execute the optimal execution strategy to obtain entity execution data may include:
[0054] Step S141: Transmit the determined optimal execution strategy to the robot's execution control module in the form of a control instruction set.
[0055] Step S142: The robot's execution control module executes the optimal execution strategy, the robot's mechanical loss sensing module records the execution process, and generates entity execution data.
[0056] Step S143: Send the entity execution data back to the digital twin model and compare it with the data from the virtual pre-simulation of the digital twin model.
[0057] For example, in the candidate strategy set, the strategy of "angular velocity 0.8 rad / s" was selected as the optimal strategy because of its lower wear (0.0005 mm < 0.001 mm). Its control command "joint 1 rotates to the 60° position at 0.8 rad / s" was sent to the physical robotic arm. During the execution, the physical sensors collected data showing that the actual joint temperature rose from 42°C to 45°C and the vibration peak was 22 Hz. These data were transmitted back to the digital twin model in real time and compared with the data in the virtual pre-simulation, which showed that "the temperature rose to 44°C and the vibration peak was 20 Hz".
[0058] Step S150: Calculate the parameter difference between the candidate strategy set and the entity execution data, trigger the model correction mechanism, correct the digital twin model, and generate a model correction parameter table.
[0059] Please refer to Figure 5 In this embodiment of the invention, calculating the parameter difference between the candidate strategy set and the entity execution data, triggering the model correction mechanism, and correcting the digital twin model may include:
[0060] Step S151: If the calculated parameter difference (e.g., temperature deviation = physical temperature - virtual temperature, vibration deviation = physical vibration frequency - virtual vibration frequency) exceeds the preset threshold (e.g., temperature deviation > 2℃, vibration deviation > 5Hz), the model correction mechanism is triggered.
[0061] Step S152: Using the deviation parameter correction formula in the model correction mechanism (to adjust or calibrate the deviation parameters existing in the system to improve the accuracy of the model, measurement or calculation results), adjust the correlation parameters of the physical property model in the digital twin model (for example, when the temperature deviation is 3℃, correct the coefficient in the temperature-friction coefficient correlation algorithm from 0.02 / 10℃ to 0.025 / 10℃).
[0062] Step S153: Based on the deviation between the robot wear amount (e.g., which can be obtained through periodic inspection of the physical robot) and the virtual wear amount in the digital twin model, adjust the correlation parameters of the mechanical loss model in the digital twin model (e.g., adjust the weight coefficient of the wear accumulation function; if the robot wear amount is 20% higher than the virtual value, adjust the coefficient of the square term of torque fluctuation from 1.0 to 1.2).
[0063] For example, after comparing the physical robot data with the virtual pre-simulation data, the temperature deviation was 1℃ (derived from 45℃-44℃) and the vibration deviation was 2Hz (derived from 22Hz-20Hz), both of which did not exceed the threshold (temperature 2℃, vibration 5Hz), but the system still recorded the deviation trend; when the fifth execution, the cumulative temperature deviation reached 3℃, triggering the model correction mechanism, which further adjusted the coefficient of the temperature-friction coefficient correlation algorithm in the digital twin model from 0.025 / 10℃ to 0.03 / 10℃, ensuring that subsequent virtual temperature predictions are more accurate.
[0064] Step S160: Using the reinforcement learning model, call the model correction parameter table and combine it with entity execution data to incrementally train the reinforcement learning model and generate an optimized candidate policy set to achieve iterative optimization of the robot's autonomous task planning and execution.
[0065] In a preferred embodiment of the present invention, the reinforcement learning model can call the model correction parameter table, use the corrected digital twin model as a new simulation environment, and add entity execution data (e.g., actual task completion time, actual wear and tear, etc.) to the training dataset to form virtual-entity fusion training samples for subsequent incremental training.
[0066] Please refer to Figure 6 In this embodiment of the invention, incremental training of the reinforcement learning model to generate an optimized candidate policy set may include:
[0067] Step S161: Using an incremental training algorithm (the model is iteratively updated by gradually adding new data on the basis of existing training, without having to retrain from scratch using all historical data), the training weights corresponding to the conditions with large deviation values (e.g., high vibration conditions correspond to large deviations) are increased based on the deviation values between the entity execution data and the model correction parameter table.
[0068] Step S162: Adjust the calculation coefficient of the loss penalty value in the reward function so that the virtual pre-simulation loss characteristics of the digital twin model match the loss characteristics of the robot (for example, if the physical wear is more sensitive to vibration, increase the proportion of vibration parameters in the loss penalty value) to generate an optimized candidate strategy set.
[0069] For example, the revised digital twin model improves the accuracy of vibration simulation. Combined with the data recorded in the entity execution data that "the entity wears out faster when vibrating at 22Hz", the reinforcement learning model increases the weight of the sample in this scenario during incremental training (e.g., from 1.0 to 1.5) and adjusts the penalty coefficient of the vibration parameter in the reward function from 0.1 to 0.15. In the virtual pre-simulation, the updated digital twin model will prioritize avoiding motion paths that cause vibrations exceeding 20Hz. When the generated strategy is executed by the physical robot, the vibration peak drops to 19Hz, and the deviation from the virtual pre-simulation is reduced to 1Hz.
[0070] Accordingly, this invention provides a robot autonomous task planning and execution control method based on reinforcement learning. The control method includes: collecting multi-source data from the sensor modules mounted on the robot and preprocessing the multi-source data to generate an original dataset of entity states; based on the original dataset of entity states, using a digital twin model, performing dynamic mapping and loss parameter calibration to generate a virtual entity state matching dataset; based on the virtual entity state matching dataset, using a preset reinforcement learning strategy pre-playing mechanism and the calibrated digital twin model as the simulation environment, performing strategy training to generate a candidate strategy set; using a preset strategy screening mechanism, sorting the candidate strategy set, determining the optimal execution strategy, and driving the robot to execute the optimal execution strategy to obtain entity execution data, thereby realizing the robot's autonomous task planning and execution. By collecting mechanical wear perception data, the real-time wear of the robot's mechanical actuators is incorporated into the decision-making model. This reduces mechanical wear while the robot completes its tasks, minimizing downtime for maintenance due to excessive wear and improving production efficiency. Dynamically calibrating the physical property parameters of the digital twin model ensures consistency between the virtual and physical states, improving the accuracy of the physical execution of the virtual pre-simulation strategy, reducing the cost of strategy trial and error due to model bias, and significantly enhancing the credibility of the virtual simulation. Furthermore, updating the weight parameters of the reinforcement learning network through incremental training algorithms improves the speed at which the reinforcement learning strategy adapts to dynamic environments, reduces the sample requirements for strategy retraining, and shortens the strategy adaptation cycle in new scenarios.
[0071] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0072] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0073] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0074] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0075] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0076] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0077] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0078] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0079] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for autonomous task planning and execution control of a robot based on reinforcement learning, characterized in that, The control method comprises: According to the sensor module carried by the robot, multi-source data is collected, and the multi-source data is preprocessed to generate an entity state original data set; Based on the entity state original data set, a digital twin model is used for dynamic mapping and loss parameter calibration to generate a virtual entity state matching data set; Based on the virtual entity state matching data set, a preset reinforcement learning strategy rehearsal mechanism is used, and the calibrated digital twin model is used as a simulation environment for strategy training to generate a candidate strategy set; The preset strategy screening mechanism is used to sort the candidate strategy set, determine the optimal execution strategy, and drive the robot to execute the optimal execution strategy to obtain entity execution data, so as to realize autonomous task planning and execution of the robot; The preset reinforcement learning strategy rehearsal mechanism is used, and the calibrated digital twin model is used as a simulation environment for strategy training, which comprises: Based on the virtual entity state matching data set, a state space is defined; Based on the physical constraints of the robot, the initial value of the action space is constructed, and the reinforcement learning algorithm is used to iteratively generate the control amount of the action space in the simulation environment; Based on the kinematics and task state data and the mechanical loss perception data, the task efficiency score and the mechanical loss penalty value of the reward function are calculated; According to the state space, the control amount of the action space, the task efficiency score and the mechanical loss penalty value, multi-round strategy training is performed to generate a candidate strategy set; The preset strategy screening mechanism is used to determine the candidate strategy with the highest task efficiency score as the optimal execution strategy; If the task efficiency scores of multiple candidate strategies are the same, a loss priority coefficient is introduced for secondary screening to determine the candidate strategy with low wear as the optimal execution strategy. The multi-source data comprises mechanical loss perception data and kinematics and task state data, and the preprocessing of the multi-source data comprises:
2. The reinforcement learning based robot autonomous task planning and execution control method according to claim 1, characterized in that, Based on the multi-source data, filtering processing is performed, and noise of the multi-source data is removed; The multi-source data after filtering processing and noise removal is packaged according to a preset data format to obtain the entity state original data set. The digital twin model is used for dynamic mapping and loss parameter calibration, which comprises:
3. The reinforcement learning based robot autonomous task planning and execution control method according to claim 2, characterized in that, The digital twin model is used to analyze the kinematics and task state data, and drive the geometric structure model in the digital twin model virtual space to move synchronously; The mechanical loss perception data is extracted and substituted into the physical property model and the mechanical loss model in the digital twin model virtual space for loss parameter calibration. The loss parameter calibration comprises:
4. The reinforcement learning based robot autonomous task planning and execution control method according to claim 3, characterized in that, A first preset algorithm is used to adjust the friction coefficient of the virtual joint in the physical property model and correct the stiffness parameter of the virtual component in the physical property model; A second preset algorithm is used to update the cumulative wear of the virtual actuator in the mechanical loss model. The robot is driven to execute the optimal execution strategy to obtain entity execution data, which comprises:
5. The reinforcement learning based robot autonomous task planning and execution control method according to claim 1, wherein, The determined optimal execution strategy is transmitted to the execution control module of the robot in the form of a control instruction set; The execution control module of the robot executes the optimal execution strategy, and the mechanical wear perception module of the robot records the execution process and generates entity execution data; The entity execution data is fed back to the digital twin model, and compared with the data of the virtual pre-performance of the digital twin model.
6. The reinforcement learning based robot autonomous task planning and execution control method according to claim 1, wherein, The control method further comprises: calculating the parameter difference between the candidate strategy set and the entity execution data, triggering the model correction mechanism, correcting the digital twin model, and generating a model correction parameter table; using the reinforcement learning model, calling the model correction parameter table, and combining the entity execution data, incrementally training the reinforcement learning model to generate an optimized candidate strategy set, to realize iterative optimization of robot autonomous task planning and execution.
7. The reinforcement learning based robot autonomous task planning and execution control method according to claim 6, characterized in that, The calculation of the parameter difference between the candidate strategy set and the entity execution data, the triggering of the model correction mechanism, and the correction of the digital twin model comprise: if the calculated parameter difference exceeds a preset threshold, the model correction mechanism is triggered; using a bias parameter correction formula in the model correction mechanism, adjusting the associated parameters of the physical property model in the digital twin model; based on the deviation between the robot wear and the virtual wear of the digital twin model, adjusting the associated parameters of the mechanical wear model in the digital twin model. 8.The reinforcement learning based robot autonomous task planning and execution control method according to claim 6, wherein, The incrementally training of the reinforcement learning model to generate an optimized candidate strategy set comprises: using an incremental training algorithm, based on the deviation between the entity execution data and the model correction parameter table, increasing the training weight corresponding to the large deviation value working condition; adjusting the calculation coefficient of the loss penalty value in the reward function, so that the virtual pre-performance loss characteristics of the digital twin model fit the robot loss characteristics, and generating an optimized candidate strategy set.
Citation Information
Patent Citations
Milling robot cutter wear state real-time monitoring method fusing digital twinning and deep learning
CN118700161A
Robot multi-mode sensing and motion cooperative control method and device
CN120347724A