A robot polishing speed control method, system, device and storage medium based on SAC reinforcement learning

CN122807856APending Publication Date: 2026-09-25HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610834527.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,现有技术中,针对机器人打磨任务的强化学习研究仍相对有限,尤其缺乏一种能够面向不同表面区域特征、兼顾轨迹跟踪与速度调节要求的机器人打磨自适应速度控制技术方案

Benefits of technology

本发明通过将机器人打磨任务建模为强化学习决策过程,建立了面向不同表面区域特征的速度自适应控制机制,使机器人能够根据当前轨迹状态、区域类型和目标速度动态调整打磨速度,克服了现有固定速度控制方式和经验规则控制方式灵活性不足、自适应能力较弱的问题;同时,通过构建仿真环境并结合状态空间、动作空间及奖励函数设计,实现了对机器人打磨过程的智能化训练与验证,使机器人在保证轨迹跟踪稳定性的同时兼顾加工效率和加工质量,具有较好的智能性、稳定性、通用性和扩展性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122807856A_ABST
    Figure CN122807856A_ABST
Patent Text Reader

Abstract

The application provides a robot polishing speed control method based on SAC reinforcement learning, which comprises the following steps: analyzing the characteristics of the robot polishing task, determining the polishing control elements, establishing the polishing task state attributes, constructing the robot polishing simulation environment and polishing trajectory data model, establishing the robot polishing speed control model based on SAC reinforcement learning, training and executing the strategy by using the reinforcement learning control model, outputting the robot polishing speed control instruction, and realizing the adaptive speed adjustment under different polishing areas; by introducing the reinforcement learning algorithm, the speed adjustment strategy under different area conditions in the robot polishing process is autonomously learned and optimized, and the adaptive control ability of the system under the complex workpiece surface environment is improved. The method can train and verify the control strategy under different polishing areas and task states, provide a reference basis for the design and optimization of the robot polishing speed control system, and has good intelligence, universality and expansibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control of industrial robots, specifically to a robot grinding speed control method based on SAC reinforcement learning. Background Technology

[0002] With the development of modern manufacturing towards automation and intelligence, industrial robots are increasingly widely used in processing scenarios such as deburring castings, polishing molds, and surface treatment of aerospace parts. Robotic grinding technology has become an important technical means to improve processing efficiency and consistency. In the robotic grinding process, grinding speed is one of the key parameters affecting processing quality and efficiency. Since workpiece surfaces typically have different processing areas such as normal areas, burr areas, and gate areas, the required grinding speed varies significantly between these areas. For example, normal surface areas are usually suitable for higher speeds to improve processing efficiency, while burr or gate areas require lower grinding speeds to avoid over-processing or damaging the workpiece surface. Therefore, when performing grinding tasks, robots not only need to run stably along a predetermined trajectory but also need to dynamically adjust the grinding speed based on the characteristics of the current processing area.

[0003] Existing robotic grinding speed control methods mostly employ fixed speed control, speed switching based on empirical rules, or a combination of position control, force control, and impedance control to adjust the processing. While these methods can accomplish grinding tasks to a certain extent, they typically suffer from insufficient speed adjustment flexibility and weak adaptability when dealing with complex workpiece surfaces, frequent area switching, or uncertain contact states. They struggle to simultaneously achieve processing efficiency, processing quality, and system stability. Particularly in complex trajectory and multi-area continuous grinding scenarios, traditional methods find it difficult to achieve precise, continuous, and adaptive speed control based on real-time task status.

[0004] In recent years, reinforcement learning, as a data-driven intelligent decision-making method, has been widely applied in robot control, path planning, and complex operational tasks. Reinforcement learning learns control strategies through continuous interaction with the environment, enabling adaptive decision-making in complex and uncertain environments without the need for precise system models. However, current research on reinforcement learning for robotic polishing tasks remains relatively limited, particularly lacking an adaptive speed control technology solution for robotic polishing that can address different surface region characteristics and balance trajectory tracking and speed adjustment requirements.

[0005] Therefore, how to construct an intelligent control method that can adaptively adjust the speed of the robot during the grinding process by combining the characteristics of the grinding area, the trajectory execution status, and the speed control requirements has become a technical problem that urgently needs to be solved in the field of industrial robot grinding. Summary of the Invention

[0006] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: In a first aspect, this application provides a robot grinding speed control method based on SAC reinforcement learning, the grinding speed control method comprising the following steps: Step S1: Analyze the characteristics of the robot polishing task, determine the components of the polishing control, and establish the polishing task state attributes; the polishing task state attributes include position-related states. Speed-related states Task objective related status Region semantic related state Historical dynamic related status ; Step S2: Construct a robot grinding simulation environment and a grinding trajectory data model; the trajectory data model is used to describe the sequence of trajectory points that the robot's end effector needs to reach sequentially during the grinding process, represented as: ; in, ( ): represents the three-dimensional spatial coordinates of i trajectory points. This indicates the surface region type to which the trajectory point belongs. The region type is classified according to the processing characteristics of the workpiece surface, including normal regions, burr regions, and gate regions. : This indicates the recommended polishing speed corresponding to the trajectory point. The recommended polishing speed is used to provide a speed adjustment target reference for the reinforcement learning control model. Different region types correspond to different recommended polishing speed ranges.

[0007] Step S3: Establish a robot grinding speed control model based on SAC reinforcement learning; Step S4: Use the reinforcement learning control model to train and execute the strategy, and output the robot grinding speed control command to achieve adaptive speed adjustment in different grinding areas.

[0008] Furthermore, the polishing speed control method is executed by the corresponding polishing speed control system. Based on the actual operation of the robotic automated polishing system, the overall control system can be divided into a trajectory input submodule, an environment perception submodule, a reinforcement learning decision-making submodule, and a robot execution submodule. The trajectory input submodule provides trajectory point data for the robot's polishing task, including the spatial location of the trajectory points, surface region type, and recommended polishing speed. The environment perception submodule acquires real-time status information such as the robot's end-effector position, end-effector speed, current polishing speed, target polishing speed, trajectory error, trajectory execution progress, and region category. The reinforcement learning decision-making submodule outputs polishing speed adjustment actions based on the current status information. The robot execution submodule receives the adjustment actions and converts them into speed control commands for the robot to move along a predetermined trajectory. Therefore, this invention establishes corresponding state attributes for each of the above functional submodules, and the set of state attributes of each submodule together constitutes the overall state space of the robot polishing speed control model, thus providing a complete state description framework for subsequent strategy learning and control execution.

[0009] To enable the reinforcement learning control model to uniformly represent the robot's current position, velocity state, task progress, and area changes during the polishing process, preferably, the state attributes corresponding to each functional sub-module are vectorized and combined to form the overall state representation at the current moment. In this way, the originally scattered motion information, task information, and area information can be mapped into unified input variables, facilitating subsequent learning and inference of the control strategy. Accordingly, in step S1, the polishing task state attributes are represented as follows: ; in: Indicates location-related status. Indicates speed-related states. Indicates the status related to the task objective. Indicates the semantically related state of a region.

[0010] Furthermore, in step S2, to construct a simulation environment for the robot grinding task, the processing characteristics of the workpiece surface are first analyzed, and the workpiece surface is divided into multiple different processing areas according to the grinding task requirements. In a preferred embodiment, the processing areas include at least a normal area, a burr area, and a gate area, with different areas corresponding to different target grinding speed ranges. Secondly, a robot grinding trajectory data model is established, wherein the trajectory data includes at least trajectory point numbers, three-dimensional spatial coordinates, surface area types, and recommended grinding speeds. To uniformly describe each discrete trajectory point on the robot grinding path and enable the trajectory information to be jointly accessed by the simulation environment and the reinforcement learning control model, preferably, a single trajectory point is abstractly represented as a data structure containing spatial location, area labels, and speed reference information. This forms a trajectory point data model suitable for continuous robot grinding tasks.

[0011] In constructing a robot grinding simulation environment based on a robot simulation platform, the simulation environment includes a robot model, a workpiece model, a worktable model, a trajectory tracking and control module, a state observation module, and a reward feedback module. This simulates the process of a robot's end effector continuously grinding along a workpiece surface trajectory. The simulation environment can dynamically switch the target position and target grinding speed based on trajectory point information, and update state and reward information after the robot performs grinding actions, providing a training foundation for reinforcement learning algorithms. The simulation environment is preferably constructed using the RoboSuite simulation framework combined with the MuJoCo physics engine to simulate the robot's kinematics and dynamics, as well as the interaction between the robot and the workpiece surface, thereby improving the approximation of the actual grinding task by the training environment.

[0012] Furthermore, in step S3, since the robot needs to continuously adjust its speed according to the current environmental state during the polishing process and obtain new state feedback and reward feedback after the action is executed, the control process can be abstracted as a sequential decision problem. Preferably, a Markov decision process is used to formally describe the robot's polishing speed control task in order to establish the correspondence between state, action, reward and state transition.

[0013] ; in: Representing the state space, Represents the action space. Represents the state transition function. Represents the reward function, This represents the discount factor.

[0014] The problem of controlling the speed of robot grinding is modeled as a reinforcement learning decision problem. A state space, action space and reward function are established, and a reinforcement learning algorithm is used to learn the control strategy.

[0015] Specifically, the state space is used to describe the robot's current motion state, task target state, and region semantic state, including at least one or more of the following: robot end-effector position, end-effector velocity component, current grinding speed, target grinding speed, trajectory error, trajectory execution progress, current region type, and subsequent region type. To improve the model's ability to perceive velocity smoothness and region switching trends, the state space may further include the velocity change at the previous moment, local trajectory curvature features, and the distance from the current position to the next target trajectory point.

[0016] To enable the reinforcement learning control model to simultaneously perceive the robot's current motion, task progress, and surface region change trends, the state space is preferably designed as an enhanced state representation composed of multiple types of state variables. Further, in a preferred embodiment, the position-related state, velocity-related state, task-objective-related state, and region semantic-related state can be represented as follows: ; ; ; ; in, : Represents the spatial position coordinates of the robot's end effector at time t; : Represents the terminal velocity component; : Indicates the current polishing speed; Indicates the target polishing speed; : Indicates the change in velocity at the previous moment; : Indicates the progress of trajectory execution; : Represents the curvature characteristics of a local trajectory; : Indicates the distance from the current position to the next target trajectory point; : These represent the current region type and the region type corresponding to subsequent trajectory points, respectively.

[0017] In terms of motion modeling, the motion space is defined as the adjustment amount of the robot's current grinding speed, and the motion space is a continuous motion space. Preferably, the motion can be represented as a control intention to increase, maintain, or decrease the speed, and is converted into an actual speed change through constraint mapping to achieve continuous adjustment of the robot's current grinding speed.

[0018] To enable the robot to continuously, smoothly, and dynamically adjust the grinding speed under different surface area conditions, preferably, the output action of the reinforcement learning control model is defined as the adjustment command for the current grinding speed. The adjustment command is first output in a normalized form, and then mapped to the actual speed change based on the speed adjustment constraints of the current area.

[0019] 1. Range of motion ; 2. Motion is mapped to velocity change. Furthermore, to enable the same control strategy to have different adjustment sensitivities in different regions, a corresponding maximum speed adjustment range can be set according to the current region type, and the action can be mapped to the actual speed change, expressed as: ; in, Set the maximum speed adjustment range for the current region type; 3. Speed ​​Update and Regional Upper and Lower Limit Constraints After obtaining the speed change, the robot's current grinding speed can be updated. Considering that different regions correspond to different allowable speed ranges, preferably, a regional constraint is applied to the updated grinding speed, expressed as: ; ; in, For the updated polishing speed, These are the lower and upper limits of the polishing speed corresponding to the current region type; 4. Rate of change of velocity limit Furthermore, to avoid excessive speed changes between adjacent control cycles that could lead to robot motion oscillations or unstable workpiece surface processing, it is preferable to further constrain the rate of speed change, as follows: ; This is the constraint threshold.

[0020] In terms of reward mechanism design, the reward function is used to characterize the degree to which the current control action satisfies the grinding task objective, and includes at least one or more of the following: speed error term, speed smoothing term, trajectory error term, and region constraint term; wherein, the speed error term is used to encourage the robot's current grinding speed to approach the target grinding speed, the speed smoothing term is used to limit speed abrupt changes between adjacent moments, the trajectory error term is used to ensure that the robot runs stably along the predetermined grinding trajectory, and the region constraint term is used to suppress excessively high speed control behavior in low-speed regions or complex structure regions.

[0021] To ensure that the reinforcement learning control model, during training, not only focuses on whether the grinding speed approaches the target speed but also simultaneously considers trajectory tracking accuracy, speed adjustment smoothness, and safety in the low-speed region, the reward function is preferably constructed as a weighted combination of multiple evaluation metrics. This approach allows the control model to progressively learn a speed control strategy that better meets the actual needs of industrial grinding under multi-objective constraints. The reward function can be expressed as: ; Each reward item describes the degree to which different control objectives are met. In a preferred embodiment, they can be represented as follows: ; ; ; ; By weighting and combining the above reward items, the reinforcement learning control model can be guided to gradually form an adaptive speed regulation strategy that takes into account processing efficiency, trajectory stability and regional safety during the training process.

[0022] Regarding algorithm selection, the reinforcement learning algorithm preferably employs the Soft Actor-Critic algorithm to achieve robot grinding speed control strategy learning in a continuous action space. The reinforcement learning control model includes a policy network and a value evaluation network. The policy network adjusts the action based on the output speed according to the current state, while the value evaluation network evaluates the value of the current state and action combination, thereby guiding the policy network update.

[0023] Furthermore, in step S4, after completing the construction of the robot grinding simulation environment and the establishment of the reinforcement learning control model, policy training is performed through continuous interaction between the agent and the simulation environment. At the start of training, the robot model, workpiece model, grinding trajectory data, and robot initial state are initialized. In each control cycle, the environment inputs the current state information to the reinforcement learning control model, and the reinforcement learning control model outputs the speed adjustment action at the current moment. The simulation environment updates the robot's current grinding speed according to the speed adjustment action and drives the robot to move along a predetermined trajectory. Subsequently, the instantaneous reward is calculated based on the speed error, trajectory error, and speed change, and the state transition data obtained from the interaction is stored in the experience playback buffer. To improve the utilization efficiency of historical interaction data and reduce the temporal correlation between continuously sampled data, preferably, the state transition samples generated in each control cycle are stored in the experience playback buffer, represented as: ; in: : Indicates time Environmental status information; : Indicates time Reinforcement learning controls the speed adjustment of the model's output. : Represents the immediate reward value fed back by the simulation environment after the action is performed; : Represents the environmental state information obtained at the next moment after the action is performed; : Represents the experience replay buffer; : Indicates the control cycle number.

[0024] Furthermore, the reinforcement learning algorithm periodically samples historical interaction data from the experience replay buffer to update the parameters of the policy network and the value evaluation network until the control policy reaches the preset convergence condition. After training, the trained control policy is used to execute the robot's polishing task, enabling the robot to automatically adjust the polishing speed under different surface area conditions and maintain stable trajectory tracking performance.

[0025] During strategy execution, corresponding upper and lower speed limits and speed change rate limits can be set for different grinding areas to ensure that the output control speed meets the safety and stability requirements of industrial grinding scenarios.

[0026] Compared with existing technologies, the robot grinding speed control method based on SAC reinforcement learning provided by this invention has the following advantages: This invention models the robot grinding task as a reinforcement learning decision-making process and establishes a speed adaptive control mechanism for different surface region characteristics. This enables the robot to dynamically adjust the grinding speed according to the current trajectory state, region type, and target speed, overcoming the problems of insufficient flexibility and weak adaptability of existing fixed speed control and empirical rule control methods. At the same time, by constructing a simulation environment and combining the design of state space, action space, and reward function, intelligent training and verification of the robot grinding process are realized. This allows the robot to ensure trajectory tracking stability while taking into account processing efficiency and quality, exhibiting good intelligence, stability, versatility, and scalability.

[0027] Secondly, this application provides a robot grinding speed control system based on SAC reinforcement learning, wherein the grinding speed control system is used to execute the grinding speed control method provided in the first aspect.

[0028] Based on the actual operation of the robotic automated polishing system, the overall control process can be divided into a trajectory input submodule, an environment perception submodule, a reinforcement learning decision-making submodule, and a robot execution submodule. The trajectory input submodule provides trajectory point data for the robot polishing task, including the spatial location of the trajectory points, surface region type, and recommended polishing speed. The environment perception submodule acquires real-time status information such as the robot's end-effector position, end-effector speed, current polishing speed, target polishing speed, trajectory error, trajectory execution progress, and region category. The reinforcement learning decision-making submodule outputs polishing speed adjustment actions based on the current status information. The robot execution submodule receives the adjustment actions and converts them into speed control commands for the robot to move along a predetermined trajectory. Therefore, this invention establishes corresponding state attributes for each of the above functional submodules, and the set of state attributes of each submodule together constitutes the overall state space of the robot polishing speed control model, thus providing a complete state description framework for subsequent strategy learning and control execution.

[0029] Thirdly, this application also provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, it implements the SAC reinforcement learning-based robot polishing speed control method of the first aspect described above.

[0030] Fourthly, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the SAC reinforcement learning-based robot polishing speed control method described in the first aspect.

[0031] Through the above technical solution, the present invention has the following beneficial effects: This invention models the robot grinding task as a reinforcement learning decision-making process and establishes a speed adaptive control mechanism for different surface region characteristics. This enables the robot to dynamically adjust the grinding speed according to the current trajectory state, region type, and target speed, overcoming the problems of insufficient flexibility and weak adaptability of existing fixed speed control and empirical rule control methods. At the same time, by constructing a simulation environment and combining the design of state space, action space, and reward function, intelligent training and verification of the robot grinding process are realized. This allows the robot to ensure trajectory tracking stability while taking into account processing efficiency and quality, exhibiting good intelligence, stability, versatility, and scalability.

[0032] It is understandable that reconstruction devices, terminal devices, and readable storage media that can implement the above methods have the same beneficial effects. Attached Figure Description

[0033] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0034] Figure 1 The flowchart shows a robot grinding speed control method based on SAC reinforcement learning. Figure 2 A schematic diagram of the grinding speed control process executed by each module in the robot grinding speed control system; Figure 3 This is a schematic diagram of the state change process of the grinding speed control model; Figure 4 To refine the overall architecture hierarchy of the reinforcement learning model for the speed control system; Figure 5 Design the state space diagram; Figure 6 Action space and hierarchical control diagram; Figure 7 This is a flowchart illustrating the grinding speed control process of the grinding speed control system in this application. Figure 8 A comparison chart showing the actual speed and target speed during the grinding speed control process; Figure 9 This is a schematic diagram of the workpiece to be polished. Figure 10 This is a structural block diagram of the terminal device of this application. Detailed Implementation

[0035] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0036] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples.

[0037] Example 1 like Figure 1 The flowchart of the robot grinding speed control method based on SAC reinforcement learning is shown. In this embodiment of the invention, a robot grinding speed control method based on SAC reinforcement learning is provided, applied to the scenario of automated grinding of workpiece surfaces by industrial robots. The method can output robot grinding speed control commands in real time for different processing areas of the workpiece surface, combined with the robot's current motion state, trajectory execution state, and target grinding requirements. This allows the robot to continuously, adaptively, and stably adjust the grinding speed while ensuring trajectory tracking accuracy. The method specifically includes the following steps: Step S1: Analyze the characteristics of robot polishing tasks, determine the components of polishing speed control, and establish model state attributes; Furthermore, based on the functional division principle of robotic automated polishing tasks, the entire robotic polishing speed control system can be divided into four collaborative sub-modules: trajectory input sub-module, environmental perception sub-module, reinforcement learning decision-making sub-module, and robot execution sub-module. The polishing speed control processing procedures executed by each module in the robotic polishing system are as follows: Figure 2 As shown.

[0038] The trajectory input submodule provides trajectory data for the robot's grinding task. This trajectory data, generated by an offline planning method, includes multiple consecutive trajectory points. Each trajectory point includes at least three-dimensional spatial coordinates, surface region type information, and a corresponding recommended grinding speed. The three-dimensional spatial coordinates indicate the target grinding position of the robot's end effector on the workpiece surface; the surface region type characterizes the processing area attributes corresponding to the current trajectory point; and the recommended grinding speed provides a speed adjustment target reference for the reinforcement learning control model.

[0039] The environmental perception submodule is used to acquire the robot's task and motion states in real time during the polishing process. This state information includes, but is not limited to, the robot's current end-effector coordinates, end-effector linear velocity components, current polishing speed, target polishing speed, distance from the current position to the target trajectory point, trajectory error, trajectory execution progress, current polishing area type, area type of subsequent trajectory points, and velocity change at the previous moment. By collecting and processing this state information, a complete state vector describing the robot's current operating state and task progress can be obtained.

[0040] The reinforcement learning decision-making submodule generates speed control actions based on the state information provided by the environment perception submodule. These speed control actions can be represented as incremental adjustments to the current grinding speed, or as a combination of high-level speed adjustment intentions and low-level local corrections. Through continuous interaction and learning with the simulation environment, the reinforcement learning decision-making submodule enables the robot to gradually acquire the ability to adaptively adjust its speed under different surface area conditions.

[0041] The robot execution submodule converts the actions output by the reinforcement learning decision submodule into actual motion control commands that the robot can execute. These motion control commands are used not only to adjust the robot's current grinding speed but also to drive the robot's end effector to move continuously along a predetermined grinding trajectory. When the robot reaches the vicinity of the current trajectory point, the system automatically updates the target trajectory point and continues subsequent trajectory tracking and speed adjustment, thereby achieving a complete continuous grinding process.

[0042] Furthermore, in this invention, based on the characteristics of the robot polishing task, the system state attributes are uniformly modeled to form the overall state space of the reinforcement learning control model. For example... Figure 3 As shown, the overall state space can be composed of the following types of state attributes: 1. Kinematic state attributes, used to describe the current spatial position and velocity information of the robot's end effector; 2. Speed ​​state attribute, used to describe the robot's current grinding speed, target grinding speed, and speed change trend in adjacent time intervals; 3. Task target status attributes, used to describe trajectory execution progress, local trajectory geometric changes, and distance to the next target point, etc.; 4. Regional semantic state attributes, used to describe the current regional type and the regional change trends corresponding to several future trajectory points; 5. Historical dynamic state attributes, used to describe the continuous information of speed and trajectory changes over a period of time.

[0043] By establishing the aforementioned state attributes, the system can accurately describe the current operating status of the robot's polishing task, providing a complete, continuous, and learnable state input foundation for the reinforcement learning control model. To uniformly represent the various state attributes, preferably, the overall state of the robot's polishing task at time t is constructed as a state vector, represented as: ; Step S2: Construct a robot grinding simulation environment and grinding trajectory data model; In order to enable the reinforcement learning control method to be trained under safe, low-cost and reproducible conditions, this embodiment constructs a robot grinding simulation environment and establishes a grinding trajectory data model and regionalized task scenarios in the simulation environment.

[0044] Furthermore, the simulation environment is built on a robot operation simulation platform, preferably using the RoboSuite simulation framework combined with the MuJoCo physics engine. RoboSuite provides a unified environment interface for robot operation tasks, robot model management, and controller interface encapsulation, while the MuJoCo physics engine accurately simulates the robot's dynamic behavior, the contact relationship between the end effector and the workpiece surface, and the state evolution during continuous motion.

[0045] In this embodiment, the robot model is preferably the Franka Panda seven-DOF industrial robot. This robot has high control precision and good motion flexibility, making it suitable for robot operation research and automated grinding scenarios. The robot is mounted near the worktable, with its end effector facing the workpiece surface, and completes the grinding task by moving along a preset trajectory on the workpiece. The workpiece model is arranged on the worktable surface, and the workpiece surface contains multiple different grinding areas to simulate complex surface structures that may exist on real industrial workpieces, such as smooth areas, burr areas, and gate areas.

[0046] Furthermore, based on the processing characteristics of the workpiece surface, the grinding area is divided into at least three different types of areas, specifically including: 1. Normal area, used to indicate areas with relatively flat surfaces suitable for high-speed processing; 2. Burr area: This indicates areas with burrs, protrusions, or uneven structures, where the sanding speed needs to be reduced appropriately. 3. The gate area is used to indicate areas with complex structures, high grinding requirements, and where grinding speed needs to be significantly reduced.

[0047] Different zones correspond to different recommended grinding speed ranges. For example, in the normal zone, a higher speed can be used to improve processing efficiency; in the burr zone, a medium speed can be used for stable removal; and in the gate zone, a lower speed should be used to ensure processing quality and reduce the risk of damage to the workpiece.

[0048] Furthermore, a robot grinding trajectory data model is established. This model describes the sequence of trajectory points that the robot's end effector needs to reach sequentially during the grinding process. Each trajectory point in the sequence includes at least the following fields: (1) Track point numbering; (2) Three-dimensional spatial coordinate position; (3) The type of surface region to which it belongs; (4) Recommended polishing speed.

[0049] During simulation, the environment module sequentially reads trajectory data and uses the current trajectory point's position as the target position for the robot's end effector. Simultaneously, based on the region type to which the trajectory point belongs, it sets the target grinding speed for the reinforcement learning control model at the current moment. Within each control cycle, the robot needs to dynamically adjust its current speed while maintaining trajectory tracking, gradually approaching the target grinding speed.

[0050] Furthermore, a state observation module, an action execution module, and a reward calculation module are constructed in the simulation environment. The state observation module acquires the current robot motion state and task state; the action execution module receives the speed adjustment actions output by the reinforcement learning control model and converts these actions into actual speed changes; the reward calculation module calculates immediate rewards based on the current control effect and feeds them back to the reinforcement learning control model.

[0051] To facilitate reinforcement learning algorithm training, the simulation environment is further encapsulated as an environment class conforming to the Gymnasium interface specification, implementing the reset and step interfaces. The reset interface is used to initialize the robot model, workpiece model, trajectory data, and environment state, while the step interface is used to execute actions and return a new state, reward value, task completion flag, and auxiliary information. Standardizing the environment interface improves the compatibility and scalability of the system, facilitating subsequent training and validation of the environment using different reinforcement learning algorithms.

[0052] Step S3: Establish a robot grinding speed control model based on SAC reinforcement learning; In robotic polishing tasks, the control objective is not merely to enable the robot to complete a trajectory, but more importantly, to enable the robot to adaptively adjust its polishing speed based on the current area and task status. Therefore, this invention models the robotic polishing speed control problem as a reinforcement learning decision problem.

[0053] Furthermore, the robot polishing process is abstracted as a Markov decision process. At each moment, the reinforcement learning agent outputs a control action based on the current environmental state. The environment updates the robot's current speed and trajectory execution state based on this control action and returns the new state along with the corresponding reward value. Through continuous interaction and parameter updates, a control strategy that maximizes cumulative rewards is ultimately learned.

[0054] (a) State-space design: like Figure 3 As shown in this embodiment, the state space of the reinforcement learning control model is composed of multiple types of information, including kinematic state, velocity state, task target state, region semantic state, and historical dynamic state.

[0055] 1. Kinematic state Kinematic state is used to describe the current spatial position and trajectory deviation of the robot's end effector. Preferably, it includes the current position coordinates of the end effector. , , The trajectory error from the robot's end effector to the current target trajectory point is also considered. This trajectory error explicitly reflects whether the robot is currently deviating from the predetermined grinding trajectory.

[0056] 2. Speed ​​Status The velocity state is used to characterize the real-time motion velocity of the robot's end effector and the grinding speed. Preferably, it includes an end effector velocity component. , , Current polishing speed Target polishing speed and the change in velocity at the previous moment. By incorporating the velocity change from the previous moment, the model can sense control inertia, thus avoiding violent oscillations between adjacent control cycles.

[0057] 3. Task objective status The task objective state describes the current progress of the polishing task and the local trajectory geometry. Preferably, it includes the trajectory execution progress. Local trajectory curvature characteristics And the distance from the robot's current position to the next target trajectory point. Trajectory curvature features can help reinforcement learning models distinguish between "smooth straight grinding segments" and "turning segments with large curvature changes," thereby automatically tending to reduce speed at turning points.

[0058] 4. Region semantic state The region semantic state is used to characterize the changing trends of the current surface region and the future trajectory region. Preferably, it includes the current region type. And the prediction region types corresponding to several subsequent trajectory points. , Region types can be encoded using one-hot encoding. By incorporating future region information, reinforcement learning models not only know "what region they are currently in," but also can anticipate "what region they are about to enter," thereby enabling them to decelerate or accelerate in advance and reduce abrupt changes in region switching.

[0059] 5. Historical Dynamic Status Historical dynamic states are used to represent short-term continuous trends. Preferably, they include sequences of velocity changes, trajectory error sequences, or region switching sequences from past time points. These states help enhance the reinforcement learning model's understanding of temporal continuity, thereby further improving the smoothness of velocity control.

[0060] Since the dimensions of different state variables differ—for example, position error is usually measured in meters (m), while grinding speed is usually measured in millimeters per second (mm / s), and trajectory progress is a proportional quantity—it is preferable to normalize all state variables to reduce the impact of dimension differences on the stability of network training.

[0061] (II) Action Space Design like Figure 4 As shown, in this embodiment, the action space of the reinforcement learning control model is defined as the adjustment amount of the robot's current grinding speed, and the action space is a continuous action space. Preferably, the original action value range is [-1, 1]. Accordingly, the original action can be expressed as: ; Furthermore, to improve the adaptability of the control strategy to complex regions and complex trajectory conditions, this invention preferably employs a hierarchical motion control mechanism. The hierarchical motion control mechanism includes: 1. High-level actions High-level actions are used to generate macro-level speed control intentions, indicating the overall speed control trend that the robot should execute at the current moment: "accelerate," "decelerate," or "maintain."

[0062] 2. Low-level actions Low-level actions are used to make fine-grained corrections to the speed based on the local conditions, in addition to the speed adjustment intentions of higher levels.

[0063] Preferably, the final initial motion is formed by the combined action of higher-level and lower-level motions. This initial motion is then mapped to the actual change in velocity. And adjust it according to the maximum speed adjustment range corresponding to the current area.

[0064] Furthermore, to meet the requirements of different grinding areas regarding the upper speed limit and speed variation range, the following constraints need to be applied after the action is mapped: 1. Regional velocity constraints Based on the current region type The updated polishing speed will be limited to the speed range allowed in the corresponding area.

[0065] 2. Rate of change of velocity constraint Limit the rate of change of velocity between adjacent time points to prevent sudden velocity changes from causing system oscillations or workpiece damage.

[0066] Furthermore, after obtaining the speed change, the robot's current grinding speed can be updated, and constraints can be imposed based on the allowable speed range of the current area, as follows: ; ; in, For the updated polishing speed, These are the lower and upper limits of the polishing speed corresponding to the current region type; The current polishing speed; ; in, The normalized motion range of the robot grinding speed control command: , Set the maximum speed adjustment range for the current region type; Set constraints on the rate of change of grinding speed , This is the constraint threshold.

[0067] Through the above design, the same strategy network can automatically exhibit different speed regulation sensitivity and control strength in different regions, thereby improving the engineering feasibility of the control strategy.

[0068] (III) Reinforcement Learning Algorithm and Network Structure Design like Figure 5 and Figure 6As shown in this embodiment, the grinding speed control model based on SAC reinforcement learning is preferably constructed using the Soft Actor-Critic (SAC) algorithm. The Soft Actor-Critic algorithm is a maximum entropy reinforcement learning algorithm suitable for continuous motion spaces, characterized by high training stability, high sample utilization, and strong exploration capabilities. It is well-suited for continuous speed adjustment control scenarios during robot grinding. Unlike traditional methods that set speed based on fixed process parameters, switch preset speeds in segments, or perform closed-loop speed adjustment based solely on a single error, this embodiment incorporates robot motion state, trajectory tracking state, regional constraint information, and multimodal environmental characteristics into the reinforcement learning decision-making process. This allows the control model to adaptively output speed adjustment actions according to different grinding areas and different trajectory states, thereby achieving grinding speed control that balances trajectory accuracy, speed stability, and local area processing safety.

[0069] like Figure 5 As shown, the reinforcement learning control model for the grinding task is trained through continuous interaction between the agent and the robot grinding simulation environment. At the start of training, the robot model, workpiece model, grinding trajectory data, and the robot's initial state are initialized. Within each control cycle, the simulation environment inputs the current state information to the reinforcement learning control model. The reinforcement learning control model outputs the speed adjustment action at the current moment. The simulation environment updates the robot's current grinding speed based on the actions and drives the robot to move along a predetermined trajectory, thereby utilizing state transition probabilities. Obtain the state in the next moment Subsequently, the simulation environment calculates the instantaneous reward based on speed error (speed accuracy), trajectory error (trajectory accuracy), speed variation (speed smoothness), and regional speed limit constraints (such as parameters for regional switching adaptability and safety in dangerous areas). and interactive samples Stored in the experience replay buffer In this process, reinforcement learning algorithms periodically sample historical data from an experience replay buffer to update the parameters of the policy network and the value evaluation network until the control policy reaches a preset convergence condition.

[0070] like Figure 6As shown, the reinforcement learning control model is preferably constructed as a semantically enhanced deep decision network, which includes at least a multi-branch state encoding network, an attention fusion network, an Actor policy network, a dual-Critic value network, and a target Critic network. The multi-branch state encoding network is used to extract visual information, point cloud information, pose / joint information, and semantic description information respectively; the attention fusion network is used to adaptively weight and fuse the features of each branch to obtain integrated state features; the Actor policy network outputs the speed adjustment action at the current moment based on the integrated state features; the dual-Critic value network and the target Critic network are used to evaluate the value of the action and improve training stability. Through the above structure, this application can not only use motion information such as current position and speed error for control, but also combine trajectory semantics, regional constraints, and other information to achieve adaptive adjustment of grinding speed, thereby improving trajectory tracking accuracy, speed control stability, and processing safety in low-speed sensitive areas.

[0071] like Figure 7 As shown, the reinforcement learning model for the polishing task includes a state perception layer, a decision generation layer, and a safe execution layer. The first layer, the state perception layer, normalizes and structures the original polishing task state in the input environment data to obtain structured state data usable by the reinforcement learning model. The second layer, the decision generation layer, performs multi-level semantic decomposition on the structured state, extracting features such as task stage, action intensity, spatial location, and posture. These features are then fused through an attention fusion network and input into the Soft Actor-Critic model. The Actor network generates original action data, and a dual Critic network performs value evaluation and parameter updates. The third layer, the safe execution layer, maps and corrects the original action data for safety. Combining the regional speed limit, speed change rate limit, and safety rule constraints, it generates executable control commands, which are then sent to the polishing robot to execute the polishing task.

[0072] Furthermore, the reinforcement learning control model includes a policy network and a value evaluation network: 1. Policy Network (Actor Network) The policy network is used to generate speed adjustment actions based on the current state. The policy network is preferably implemented using a multi-layer fully connected neural network, with its output layer employing the tanh activation function to restrict the action output to the [-1, 1] interval. The action output can be converted into the actual speed adjustment amount after scaling.

[0073] 2. Double Critic Network The value network is used to evaluate the value of the current state and action combination. It is preferable to use two independent Q-networks to reduce Q-value overestimation and improve training stability.

[0074] Furthermore, to enhance the model's ability to extract different types of state information, this invention preferably employs a multi-branch state coding structure. The multi-branch state coding structure includes: The kinematics branch is used to encode position, velocity, and trajectory error information; Task branches are used to encode trajectory progress, curvature, and target point distance information; Region semantic branch, used to encode the current region and future region types; Historical dynamic branches are used to encode short-term velocity change trends and continuous state change characteristics.

[0075] Furthermore, the features output by each branch are dynamically weighted using an attention fusion mechanism, enabling the reinforcement learning model to automatically select more important state information at different task stages. For example, at region switching locations, region semantic features have higher weights; at trajectory turning points, trajectory curvature and velocity information have higher weights. The fused global state representation is then used as input to the policy network and the value evaluation network.

[0076] (iv) Reward Function Design In robotic polishing tasks, reinforcement learning control models need to simultaneously consider speed control accuracy, trajectory tracking accuracy, speed smoothness, and area safety. Therefore, this invention constructs a multi-factor comprehensive reward function. To comprehensively evaluate the control effect of the current control action in terms of speed tracking, trajectory tracking, speed smoothness, and area safety, preferably, the reward function is constructed as a weighted combination of multiple reward items, expressed as: ; , , , The weights for the velocity error term, velocity smoothing term, trajectory error term, and region constraint term are respectively defined, and the sum of these weights is 1. The reward function includes at least one or more of the following reward terms: 1. Speed ​​error term

[0077] Used to encourage the robot's current polishing speed to approach the target polishing speed; ; in, : for time The robot's current actual grinding speed; To achieve the target of speed.

[0078] 2. Speed ​​smoothing term

[0079] Used to limit excessive velocity changes between adjacent moments and suppress velocity oscillations; ; in, : for time 1. The robot's actual grinding speed at the previous moment.

[0080] 3. Trajectory Error Term

[0081] Used to encourage robots to operate stably along a predetermined trajectory; ; in, : for time The current position of the robot's end effector; : for time The corresponding target trajectory point location.

[0082] 4. Region Constraints

[0083] Used to impose additional penalties on excessively high-speed behavior in areas requiring low-speed fine processing, such as burr areas and gate areas; ; in, : The maximum speed allowed in the current processing area; : This is the overspeed penalty coefficient, used to adjust the intensity of the penalty when the speed exceeds the limit.

[0084] By weighting and combining the above reward items, the reinforcement learning model can gradually learn the optimal speed control strategy that simultaneously satisfies the requirements of processing efficiency, polishing quality, trajectory stability, and safety during the training process.

[0085] Step S4: Use a reinforcement learning control model to train and execute the policy, and output the robot's grinding speed control command; After completing the construction of the robot polishing simulation environment and the establishment of the reinforcement learning control model, the training and execution phase begins.

[0086] (I) Training Process At the start of training, the simulation environment is first initialized. The initialization process includes loading the robot model, workpiece model, grinding trajectory data, and setting the robot's initial state. Preferably, the robot's initial grinding speed is set to a medium value to ensure sufficient exploration space in the early stages of training. The environment also sets the current trajectory point as the starting trajectory point and initializes the trajectory execution progress.

[0087] In each control cycle, the simulation environment first provides the reinforcement learning control model with the current state information. This state information includes the robot's end-effector position, end-effector velocity, current grinding speed, target grinding speed, trajectory error, current trajectory point number, trajectory execution progress, and region type, etc.

[0088] The reinforcement learning control model generates an action output based on the current state input, whereby the action represents the adjustment amount of the robot's grinding speed at the current moment. Upon receiving this action, the environment maps it to the actual speed change and updates the robot's current grinding speed accordingly, while simultaneously driving the robot's end effector to continue moving along a predetermined trajectory.

[0089] After the robot performs an action, the simulation environment updates the robot's motion state and calculates the new environmental state. Subsequently, the environment calculates an immediate reward based on the current speed error, trajectory error, speed change, and whether the area constraints are met. This immediate reward reflects the quality of the current control strategy; a higher reward is given when the robot's speed is closer to the target speed, the speed change is smoother, and the trajectory error is smaller; conversely, a penalty is incurred.

[0090] The environment feeds back the new state, reward value, and task completion flag to the reinforcement learning control model, and stores the state transition data generated by this interaction in the experience replay buffer. The reinforcement learning algorithm periodically samples data randomly from the experience replay buffer to update the parameters of the policy network and the dual-value network. By continuously repeating the above interaction and update process, the reinforcement learning control policy gradually converges, ultimately obtaining an adaptive speed control policy suitable for different polishing areas.

[0091] Furthermore, to improve training stability, a target network is set in the SAC algorithm, and the target network parameters are updated using a soft update method, making the change of the target Q value smoother, thereby reducing the oscillation problem during training.

[0092] (II) Strategy Execution Process After the reinforcement learning control model is trained, the trained policy model is saved and used for subsequent testing or actual deployment. During the policy execution phase, the robot sequentially passes through multiple trajectory points according to the pre-input grinding trajectory, collects current state information in each control cycle, and inputs it into the trained reinforcement learning control model.

[0093] The reinforcement learning control model outputs a speed adjustment action based on the current state. After the action is mapped by region constraints and limited by the rate of change of speed, it generates a control command for the robot's actual grinding speed. The robot's execution module adjusts the current grinding speed according to the control command and continues to move along the predetermined trajectory.

[0094] When the robot moves from the normal area to the burr area or the gate area, the reinforcement learning control model can automatically reduce the grinding speed according to the change in area; when the robot re-enters the normal area, the reinforcement learning control model can appropriately increase the grinding speed according to the target speed requirement. Thus, the robot can smoothly switch speeds between different surface areas, achieving continuous, adaptive, and stable control throughout the entire grinding process.

[0095] In a preferred embodiment, the robot maintains a small trajectory error throughout the polishing task and makes the actual polishing speed close to the target polishing speed at most trajectory points, thereby balancing processing efficiency and processing quality.

[0096] Example The method of the present invention will be further described below with reference to specific embodiments.

[0097] In this embodiment, the robot model uses the Franka Panda seven-DOF robotic arm, and the simulation platform uses RoboSuite and MuJoCo. Multiple grinding zones are set on the workpiece surface, including at least a normal zone, a burr zone, and a gate zone. Different recommended grinding speeds are set for different zones, with higher speeds for the normal zone, medium speeds for the burr zone, and lower speeds for the gate zone.

[0098] At the start of the simulation, pre-planned grinding trajectory data is imported into the simulation environment. This trajectory data contains 70 trajectory points, each with its corresponding 3D position, region type, and recommended grinding speed. The robot begins its grinding task from the first trajectory point.

[0099] During the training phase, the reinforcement learning control model continuously interacts with the simulation environment. As training progresses, the robot gradually learns the following control laws: it tends to use a higher grinding speed in normal areas to improve processing efficiency; it automatically reduces and maintains a stable speed in burr areas; it significantly reduces the speed in gate areas to improve grinding precision; and it can adjust the speed in advance when areas switch to reduce sudden changes in speed.

[0100] During the testing phase, for example Figure 8 The workpiece to be polished is then polished, and the robot executes the complete polishing trajectory by invoking a pre-trained control strategy. Experimental results show that, combined with the attached... Figure 9 The comparison diagram of the actual speed and target speed in the grinding speed control process is shown to illustrate that at most trajectory points, the error between the actual grinding speed and the target grinding speed of the robot is small, and the trajectory error remains at a low level. This shows that the control method described in this invention can achieve adaptive speed control for different grinding areas, while ensuring stable movement of the robot along the trajectory.

[0101] Furthermore, by combining the training reward curve and the target speed comparison curve, it can be seen that in the early stage of training, due to the immaturity of the control strategy, there is a large error between the robot's polishing speed and the target speed, and the average reward value is low. As the number of training steps increases, the control strategy gradually converges, the reward value gradually increases and tends to stabilize, indicating that the reinforcement learning model has learned to adjust the speed reasonably according to different regions and trajectory states.

[0102] Therefore, the robot grinding speed control method based on SAC reinforcement learning proposed in this invention has stronger adaptability, continuity and intelligence compared with traditional fixed speed control methods and empirical rule control methods, and can more effectively adapt to complex workpiece surfaces, complex trajectories and multi-area continuous grinding tasks.

[0103] This invention is primarily used for adaptive speed control in robotic automated grinding systems. By performing state modeling, region modeling, simulation modeling, and reinforcement learning control modeling on the robotic grinding task, this invention proposes a robotic grinding speed adaptive control technology scheme for processing different surface regions. This scheme can autonomously learn speed adjustment rules based on the current trajectory state, region type, and target speed, and continuously output control commands that meet region constraints and speed smoothing requirements during execution.

[0104] Compared with the prior art, the present invention has at least the following advantages: 1. It can overcome the limitations of fixed speed control and manual experience-based control methods, and improve the flexibility of speed adjustment under complex working conditions; 2. It can automatically adjust the speed according to the characteristics of different grinding areas, balancing processing efficiency and processing quality; 3. It can maintain good trajectory tracking stability while controlling speed; 4. It can achieve training and verification through a simulation environment, reducing the risks and costs associated with direct training of real robots; 5. It has good modularity, scalability and engineering application potential, and can be extended to other robot surface processing tasks.

[0105] Example 2 This application also provides a robot grinding speed control system based on SAC reinforcement learning, which is used to execute the grinding speed control method provided in Embodiment 1 when running.

[0106] Based on the actual operation of the robotic automated polishing system, the overall control process can be divided into a trajectory input submodule, an environment perception submodule, a reinforcement learning decision-making submodule, and a robot execution submodule. The trajectory input submodule provides trajectory point data for the robot polishing task, including the spatial location of the trajectory points, surface region type, and recommended polishing speed. The environment perception submodule acquires real-time status information such as the robot's end-effector position, end-effector speed, current polishing speed, target polishing speed, trajectory error, trajectory execution progress, and region category. The reinforcement learning decision-making submodule outputs polishing speed adjustment actions based on the current status information. The robot execution submodule receives the adjustment actions and converts them into speed control commands for the robot to move along a predetermined trajectory. Therefore, this invention establishes corresponding state attributes for each of the above functional submodules, and the set of state attributes of each submodule together constitutes the overall state space of the robot polishing speed control model, thus providing a complete state description framework for subsequent strategy learning and control execution.

[0107] Example 3 like Figure 10 As shown, this application also provides a terminal device 30, including a memory 31, a processor 32, and a computer program 33 stored in the memory and executable on the processor. For example, when the processor 32 executes the computer program 33, it implements the steps in the above embodiments of the robot grinding speed control method based on SAC reinforcement learning. When the processor 32 executes the computer program 33, it implements the functions of each module in the above device embodiments, such as the functions of each module and unit in Embodiment 2.

[0108] For example, the computer program 33 may be divided into one or more modules, which are stored in the memory 31 and executed by the processor 32 to complete Embodiment 1 of this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 33 in the terminal device 30.

[0109] The terminal device 30 may be a receiving terminal device. The terminal device may include, but is not limited to, a memory 31 and a processor 32. Those skilled in the art will understand that... Figure 10 This is merely an example of terminal device 30 and does not constitute a limitation on terminal device 30. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal device may also include input / output devices, network access devices, buses, etc.

[0110] The memory 31 can be an internal storage unit of the terminal device 30, such as a hard disk or RAM of the terminal device 30. The memory 31 can also be an external storage device of the terminal device 30, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal device 30. Furthermore, the memory 31 can include both internal and external storage units of the terminal device 30. The memory 31 is used to store the computer program and other programs and data required by the terminal device. The memory 31 can also be used to temporarily store data that has been output or will be output.

[0111] The processor 32 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0112] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0113] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0114] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0115] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / terminal devices and methods can be implemented in other ways. For example, the apparatus / terminal device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0116] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0117] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0118] If the integrated module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0119] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A robot grinding speed control method based on SAC reinforcement learning, characterized in that, The grinding speed control method includes the following steps: Step S1: Analyze the characteristics of the robot polishing task, determine the components of the polishing control, and establish the polishing task state attributes; the polishing task state attributes include position-related states. Speed-related states Task objective related status Region semantic related state Historical dynamic related status ; Step S2: Construct a robot grinding simulation environment and a grinding trajectory data model; the trajectory data model is used to describe the sequence of trajectory points that the robot's end effector needs to reach sequentially during the grinding process, represented as: ; in, ( ): represents the three-dimensional spatial coordinates of i trajectory points. This indicates the surface region type to which the trajectory point belongs. The region type is classified according to the processing characteristics of the workpiece surface, including normal regions, burr regions, and gate regions. : This indicates the recommended polishing speed corresponding to the trajectory point. The recommended polishing speed is used to provide a speed adjustment target reference for the reinforcement learning control model. Different region types correspond to different ranges of recommended polishing speeds. Step S3: Establish a robot grinding speed control model based on SAC reinforcement learning; Step S4: Use the SAC reinforcement learning-based robot grinding speed control model to train and execute the strategy, and output robot grinding speed control commands to achieve adaptive speed adjustment in different grinding areas. In step S3, a Markov decision process is used to formally describe the robot's grinding speed control task, establishing the correspondence between states, actions, rewards, and state transitions: ; in: Representing the state space, Represents the action space. Represents the state transition function. Represents the reward function, Indicates the discount factor; The robot grinding speed control problem is modeled as a reinforcement learning decision problem. A state space, action space and reward function are established, and a reinforcement learning algorithm is used to learn the control policy, thus establishing a robot grinding speed control model based on SAC reinforcement learning.

2. The method according to claim 1, characterized in that, The position-related state is used to describe the current spatial position and velocity information of the robot's end effector; the velocity state attribute is used to describe the robot's current grinding speed, target grinding speed, and velocity change trend in adjacent moments; the task target-related state is used to describe the trajectory execution progress, local trajectory geometric changes, and distance to the next target point; the region semantic-related state is used to describe the current region type and the region type corresponding to subsequent trajectory points; the historical dynamic state is used to describe the continuity information of velocity and trajectory changes over a period of time; The polishing task status attributes are respectively represented as follows: ; ; ; ; ; ; in, : Represents the spatial position coordinates of the robot's end effector at time t; : Represents the terminal velocity component; : Indicates the current polishing speed; Indicates the target polishing speed; : Indicates the change in velocity at the previous moment; : Indicates the progress of trajectory execution; : Represents the curvature characteristics of a local trajectory; : Indicates the distance from the current position to the next target trajectory point; : These represent the current region type and the region type corresponding to subsequent trajectory points, respectively; These represent the polishing task status at the previous moment and the moment before that, respectively.

3. The method according to claim 2, characterized in that, The robot grinding speed control model based on SAC reinforcement learning is implemented as a Markov decision process. The reinforcement learning agent outputs control actions at each moment according to the current environmental state. The environment updates the robot's current speed and trajectory execution state according to the control actions and returns the new state and corresponding reward value.

4. The method according to claim 3, characterized in that, During the operation of the grinding simulation environment, the current robot motion state and task state are obtained through the state observation module; the action execution module receives the speed adjustment action output by the robot grinding speed control model and converts the adjustment action into actual speed change; the reward calculation module calculates the real-time reward based on the current control effect and feeds it back to the robot grinding speed control model.

5. The method according to claim 4, characterized in that, The robot grinding speed control model adopts the SoftActor-Critic model; The reward function of the robot grinding speed control model is expressed as: ; in, This is a speed error term, used to encourage the robot's current polishing speed to approach the target polishing speed; This is a velocity smoothing term, used to limit excessive velocity changes between adjacent time points and suppress velocity oscillations; This is the trajectory error term, used to encourage the robot to run stably along a predetermined trajectory; This is a region constraint term used to impose additional penalties on excessively high-speed behavior in areas requiring low-speed fine machining, such as burr areas and gate areas. The weights are respectively the weights of the velocity error term, velocity smoothing term, trajectory error term, and region constraint term, and the sum of the weights is 1.

6. ; ; ; ; in, : for time The robot's current actual grinding speed; Sharpening speed to achieve the target; : for time 1. The robot's actual grinding speed at the previous moment; : for time The current position of the robot's end effector; : for time The corresponding target trajectory point location; : The maximum speed allowed in the current processing area; : This is the overspeed penalty coefficient, used to adjust the intensity of the penalty when the speed exceeds the limit.

7. The method according to claim 1, characterized in that, In step S4, the robot grinding speed control command is first output in normalized form, and then mapped to the actual speed change by combining the speed adjustment constraint of the current region. Specifically, it includes the following steps: S401. Set the corresponding maximum speed adjustment range according to the current area type, and map the action to the actual speed change, expressed as: ; in, The normalized motion range of the robot grinding speed control command: , Set the maximum speed adjustment range for the current region type; S402. Update the current grinding speed based on the speed change, and apply a region constraint to the updated grinding speed, as follows: ; ; in, For the updated polishing speed, This indicates the current polishing speed. These are the lower and upper limits of the polishing speed corresponding to the current region type; Set constraints on the rate of change of grinding speed , This is the constraint threshold.

8. The method according to claim 6, characterized in that, In step S4, after the construction of the robot grinding simulation environment and the establishment of the reinforcement learning control model are completed, strategy training is carried out through continuous interaction between the agent and the simulation environment. At the start of training, the robot model, workpiece model, grinding trajectory data and robot initial state are initialized. In each control cycle, the environment inputs the current state information to the reinforcement learning control model, and the reinforcement learning control model outputs the speed adjustment action at the current moment. The simulation environment updates the robot's current grinding speed according to the speed adjustment action and drives the robot to move along the predetermined trajectory. Subsequently, the instantaneous reward is calculated based on the speed error, trajectory error, and speed change, and the interactively obtained state transition data is stored in the experience playback buffer; specifically, the state transition samples generated within each control cycle are stored in the experience playback buffer, represented as: ; in: : Indicates time Environmental status information; : Indicates time Reinforcement learning controls the speed adjustment of the model's output. : Represents the immediate reward value fed back by the simulation environment after the action is performed; : Represents the environmental state information obtained at the next moment after the action is performed; : Represents the experience replay buffer; : Indicates the control cycle number.

9. The reinforcement learning algorithm periodically samples historical interaction data from the experience replay buffer to update the parameters of the policy network and value evaluation network in the Soft Actor-Critic model until the control policy reaches the preset convergence condition. After training, the robot uses the trained control policy to perform the robot polishing task, enabling the robot to automatically adjust the polishing speed under different surface area conditions and maintain stable trajectory tracking performance.

10. A robot grinding speed control system based on SAC reinforcement learning, used to execute the grinding speed control method according to any one of claims 1-7, characterized in that, The grinding speed control system includes: a trajectory input submodule, an environmental perception submodule, a reinforcement learning decision-making submodule, and a robot execution submodule; The trajectory input submodule is used to provide trajectory point data for the robot polishing task, including the spatial location of the trajectory points, surface area type, and recommended polishing speed; The environmental perception submodule is used to acquire the robot's current status information in real time. The status information includes the robot's end position, end speed, current polishing speed, target polishing speed, trajectory error, trajectory execution progress, and area category status information. The reinforcement learning decision-making submodule is used to output grinding speed adjustment actions based on the current state information; The robot execution submodule is used to receive the adjustment action and convert it into control instructions for the robot to move along a predetermined trajectory and grind at a predetermined speed. A terminal device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor; when the processor executes the computer program, it implements the steps of the robot polishing speed control method based on SAC reinforcement learning as described in any one of claims 1-7; A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the robot grinding speed control method based on SAC reinforcement learning as described in any one of claims 1-7.