Robot training method and system, electronic equipment and storage medium

High-precision simulation scenarios are constructed through implicit rendering equations and contact impact models, and combined with near-end strategy optimization algorithms and dynamic reward updates, the problems of inefficient scene construction, insufficient physical simulation and weak generalization of strategies in the existing technology are solved, and efficient and highly adaptable robot training is achieved.

CN120095828APending Publication Date: 2025-06-06SHANGHAI ELECTRICGROUP CORP

Patent Information

Application Number
CN202510502360.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing robot training methods have problems such as inefficient scenario construction, insufficient physical simulation, and weak generalization of strategies, resulting in high deployment costs and weak scenario adaptability.

Method used

High-precision simulation scenarios are constructed based on implicit rendering equations and contact impact models, and on this basis, the robot is controlled to perform training tasks using a near-end strategy optimization algorithm, and dynamically update the reward function and policy network model.

Benefits of technology

It significantly improves the efficiency of scenario construction, reduces manual intervention, improves the adaptability of the strategy model to complex physical interactions, improves the task success rate and reduces the cost of development hardware.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120095828A_ABST
    Figure CN120095828A_ABST
Patent Text Reader

Abstract

The invention provides a robot training method and system, electronic equipment and a storage medium. The training method comprises the following steps: constructing a training scene according to simulation environment configuration parameters and based on an implicit rendering equation and a contact impact model; the simulation environment configuration parameters comprise robot model parameters, scene asset parameters and physical parameters; and controlling the robot to execute a training task based on a near-end strategy optimization algorithm in the training scene. According to the method, the high-precision simulation scene is constructed through the implicit rendering equation and the contact impact model, the scene construction efficiency is remarkably improved, and manual intervention is reduced; the robot is driven to execute the training task based on the near-end strategy optimization algorithm, rapid adaptation of the strategy model to complex physical interaction is achieved, and compared with a traditional fixed trajectory planning scheme, the task success rate is increased, and the hardware development cost is reduced; the core pain points of low scene construction efficiency, insufficient physical simulation and weak strategy generalization ability in traditional robot training are comprehensively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of robots, and in particular to a robot training method, system, electronic device and storage medium. Background Art

[0002] In the context of the intelligent transformation of the manufacturing industry, robotics technology has gradually evolved from traditional preset trajectory control to data-driven autonomous learning. Traditional solutions rely on computer-aided design (CAD) and motion planning algorithms to generate the motion path of the robot arm, and convert the trajectory into joint instructions through inverse kinematics, but there are problems such as poor flexibility and reliance on manual experience. In recent years, embodied intelligence technology has given robots the ability to acquire autonomous skills through imitation learning and reinforcement learning: imitation learning is based on teleoperation or teaching to collect sample data (such as joint angles, end forces, visual images, etc.) to train input-output mapping models; reinforcement learning interacts with the environment through trial and error to optimize the strategy model with a reward function. However, existing technologies still face significant bottlenecks - imitation learning requires massive labeled data and has limited generalization, and reinforcement learning relies on real-time environmental interaction and precise reward design, resulting in high deployment costs and weak scene adaptability. Therefore, there is an urgent need for a robot autonomous training method that can integrate data-driven and physical constraints, and take into account efficiency and generalization capabilities, to break through the dependence of traditional technologies on human intervention and the limitations of applicability in complex scenarios. Summary of the invention

[0003] The technical problem to be solved by the present disclosure is to overcome the defects of the prior art in robot training, such as inefficient scene construction, insufficient physical simulation, and weak strategy generalization, and to provide a robot training method, system, electronic device, and storage medium.

[0004] The present invention solves the above technical problems through the following technical solutions:

[0005] A robot training method is disclosed, the training method comprising:

[0006] Constructing a training scene according to simulation environment configuration parameters and based on implicit rendering equations and contact impact models; the simulation environment configuration parameters include: robot model parameters, scene asset parameters, and physical parameters;

[0007] In the training scenario, the robot is controlled to perform a training task based on a proximal strategy optimization algorithm.

[0008] Optionally, configuring parameters according to the simulation environment and constructing a training scene based on an implicit rendering equation and a contact impact model includes:

[0009] Obtaining the environment configuration parameters;

[0010] Determine the scene object corresponding to the training scene according to the environment configuration parameters;

[0011] Generate a geometric model of the scene object through an implicit rendering equation;

[0012] Calculating a physical interaction model of the scene object based on a contact impact model;

[0013] The geometric model and the physical interaction model are integrated and embedded into an initial simulation environment, and the rendered initial simulation environment is used as a training scene.

[0014] Optionally, controlling the robot to perform a training task based on a proximal strategy optimization algorithm in the training scenario includes:

[0015] Initialize the robot's state vector and action vector according to the training task, and load the reward function corresponding to the training task;

[0016] Dynamically update the weight coefficient of the reward item in the reward function according to the historical success rate of the training tasks that have completed the training, and combine the hybrid data training strategy network model with online sampling and offline samples;

[0017] In response to the number of training times reaching a preset number, the current policy network model is evaluated and the policy network model is adjusted according to the evaluation result.

[0018] Optionally, dynamically updating the weight coefficient of the reward item in the reward function according to the historical success rate of the training task that has completed the training, and combining the hybrid data training strategy network model of online sampling and offline samples, includes:

[0019] Extracting a number of target reward items from the training tasks of the completed training;

[0020] Dynamically modifying the weight coefficient of the target reward item according to the achievement rate corresponding to the target reward item to update the reward function;

[0021] Adjust the focus of data collection according to the updated reward function, and collect the robot's current operating data in real time;

[0022] Filter the runs that match the updated reward function from the successful subset of training tasks that have completed training;

[0023] The current operation data and the historical operation data are mixed according to a preset ratio to train the strategy network model.

[0024] Optionally, evaluating the current policy network model and adjusting the policy network model according to the evaluation result includes:

[0025] The comprehensive evaluation score is calculated based on the success rate of completed training tasks, the cumulative value of the reward function, and the stability of the policy network model;

[0026] In response to the comprehensive evaluation score decreasing, triggering the policy network model to be retrained; or, in response to the comprehensive evaluation score being less than or equal to a score threshold, adjusting the policy network model structure;

[0027] The weights of the policy network model are updated by gradient descent.

[0028] Optionally, configuring parameters according to the simulation environment and constructing a training scene based on an implicit rendering equation and a contact impact model includes:

[0029] Retrieving the most similar initial scene template from the template library according to the simulation environment configuration parameters;

[0030] A correction instruction based on the initial scene template is obtained, and the initial scene template is corrected according to the correction instruction; the corrected initial scene template is used to construct the training scene.

[0031] A robot training system is disclosed, the training system comprising:

[0032] A scene construction module is used to construct a training scene according to simulation environment configuration parameters and based on implicit rendering equations and contact impact models; the simulation environment configuration parameters include: robot model parameters, scene asset parameters and physical parameters;

[0033] The training module is used to control the robot to perform training tasks based on the proximal strategy optimization algorithm in the training scenario.

[0034] Optionally, the scene construction module is specifically used to:

[0035] Obtaining the environment configuration parameters;

[0036] Determine the scene object corresponding to the training scene according to the environment configuration parameters;

[0037] Generate a geometric model of the scene object through an implicit rendering equation;

[0038] Calculating a physical interaction model of the scene object based on a contact impact model;

[0039] The geometric model and the physical interaction model are integrated and embedded into an initial simulation environment, and the rendered initial simulation environment is used as a training scene.

[0040] Optionally, the training module is specifically used to:

[0041] Initialize the robot's state vector and action vector according to the training task, and load the reward function corresponding to the training task;

[0042] Dynamically update the weight coefficient of the reward item in the reward function according to the historical success rate of the training tasks that have completed the training, and combine the hybrid data training strategy network model with online sampling and offline samples;

[0043] In response to the number of training times reaching a preset number, the current policy network model is evaluated and the policy network model is adjusted according to the evaluation result.

[0044] Optionally, the training module is specifically used to:

[0045] Extracting a number of target reward items from the training tasks of the completed training;

[0046] Dynamically modifying the weight coefficient of the target reward item according to the achievement rate corresponding to the target reward item to update the reward function;

[0047] Adjust the focus of data collection according to the updated reward function, and collect the robot's current operating data in real time;

[0048] Filter the runs that match the updated reward function from the successful subset of training tasks that have completed training;

[0049] The current operation data and the historical operation data are mixed according to a preset ratio to train the strategy network model.

[0050] Optionally, the training module is specifically used to:

[0051] The comprehensive evaluation score is calculated based on the success rate of completed training tasks, the cumulative value of the reward function, and the stability of the policy network model;

[0052] In response to the comprehensive evaluation score decreasing, triggering the policy network model to be retrained; or, in response to the comprehensive evaluation score being less than or equal to a score threshold, adjusting the policy network model structure;

[0053] The weights of the policy network model are updated by gradient descent.

[0054] Optionally, the scene construction module is specifically used to:

[0055] Retrieving the most similar initial scene template from the template library according to the simulation environment configuration parameters;

[0056] A correction instruction based on the initial scene template is obtained, and the initial scene template is corrected according to the correction instruction; the corrected initial scene template is used to construct the training scene.

[0057] An electronic device is disclosed, comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, wherein the processor implements any one of the above-mentioned robot training methods when executing the computer program.

[0058] A computer-readable storage medium is disclosed, on which a computer program is stored. When the computer program is executed by a processor, the training method of the robot described in any one of the above is implemented.

[0059] A computer program product is disclosed, comprising a computer program, wherein when the computer program is executed by a processor, the training method of the robot described in any one of the above is implemented.

[0060] On the basis of being in accordance with the common sense in the art, the above-mentioned preferred conditions can be arbitrarily combined to obtain the preferred embodiments of the present disclosure.

[0061] The positive progress of the present disclosure is that high-precision simulation scenes are constructed through implicit rendering equations and contact impact models, which significantly improves the efficiency of scene construction and reduces manual intervention; based on the proximal policy optimization algorithm, the robot is driven to perform training tasks, which realizes the rapid adaptation of the policy model to complex physical interactions. Compared with the traditional fixed trajectory planning scheme, the task success rate is improved and the development hardware cost is reduced, which comprehensively solves the core pain points of inefficient scene construction, insufficient physical simulation and weak policy generalization ability in traditional robot training. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 A flowchart of a robot training method provided by an exemplary embodiment of the present disclosure;

[0063] Figure 2 A flowchart of step 101 is provided for an exemplary embodiment of the present disclosure;

[0064] Figure 3 A flowchart of another step 101 provided for an exemplary embodiment of the present disclosure;

[0065] Figure 4 A flowchart of step 102 provided for an exemplary embodiment of the present disclosure;

[0066] Figure 5 A schematic diagram of the architecture of a robot training method provided by an exemplary embodiment of the present disclosure;

[0067] Figure 6 A schematic diagram of a multi-person collaborative workflow of a robot training method provided by an exemplary embodiment of the present disclosure;

[0068] Figure 7A schematic diagram of a jacking action of a robot training method provided by an exemplary embodiment of the present disclosure;

[0069] Figure 8 A schematic diagram of a bowl-shaped workpiece handling action of a robot training method provided by an exemplary embodiment of the present disclosure;

[0070] Fig. 9 A schematic diagram of a module of a robot training system provided by an exemplary embodiment of the present disclosure;

[0071] Fig.10 The present invention provides a schematic structural diagram of an electronic device according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0072] The present disclosure is further described below by way of examples, but the present disclosure is not limited to the scope of the examples.

[0073] Prefixes such as "first" and "second" are used in the embodiments of the present disclosure only to distinguish different description objects, and have no limiting effect on the position, order, priority, quantity or content of the described objects. The use of prefixes such as ordinal numbers to distinguish description objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. For the statement of the described objects, please refer to the description in the context of the embodiments, and no unnecessary limitation should be constituted due to the use of such prefixes. In addition, in the description of the present embodiment, unless otherwise specified, the meaning of "plurality" is two or more.

[0074] In the embodiments of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0075] Example 1

[0076] Figure 1 A flowchart of a robot training method provided by an exemplary embodiment of the present disclosure is provided.

[0077] A robot training method is disclosed, the training method comprising:

[0078] Step 101: construct a training scene according to simulation environment configuration parameters and based on implicit rendering equations and contact impact models. The simulation environment configuration parameters include: robot model parameters, scene asset parameters, and physical parameters.

[0079] The core goal of step 101 is to build a high-precision and reusable robot training scene by combining simulation environment configuration parameters, implicit rendering equations and contact impact models. The core is to solve the problems of insufficient physical simulation accuracy and low scene reuse efficiency in traditional methods, and ensure that the robot can simulate real physical interaction behaviors in the simulation environment, thereby improving training effects and mission success rates.

[0080] Alternatively, see Figure 2 It can be seen that step 101, configuring parameters according to the simulation environment and building a training scene based on the implicit rendering equation and the contact impact model, specifically includes:

[0081] Step 1011: Obtain environmental configuration parameters.

[0082] The core purpose of step 1011 is to build the basic link of the robot training scene. Its core goal is to extract the parameter data required for the robot body, scene assets and physical interaction from the user input or preset template library, and provide accurate input basis for subsequent geometric model generation and physical simulation. This step aims to solve the simulation distortion problem caused by missing or incorrect parameters in traditional methods, and ensure the high consistency between the robot training environment and the real physical scene. Among them, the core actions and functions of the step execution are as follows:

[0083] 1. Parameter sources and classification

[0084] The acquisition of environmental configuration parameters must cover the physical characteristics of the robot body, the geometric and material properties of scene objects, and the key coefficients of physical interaction. Specifically, it includes:

[0085] Robot model parameters: such as joint stiffness (determines the motion stiffness of the end effector), damping coefficient (affects the motion damping effect), control mode (position or torque control), etc. For example, in the jacking task, the robot joint stiffness needs to be set to 0.5N·m / rad to ensure that the motion trajectory of the end effector is close to the response of the real physical system.

[0086] Scene asset parameters: including the geometric dimensions of the workpiece (such as hole diameter Φ5mm±0.1mm), material density (such as aluminum alloy 2700kg / m³) and surface roughness, etc. For example, in the bowl-shaped workpiece handling scene, the bowl mouth diameter (Φ12cm±0.5cm) and center of gravity offset (0.5cm) need to be defined to ensure the stability of the gripper when grasping.

[0087] Physical parameters such as the coefficient of friction (0.3), the coefficient of restitution (0.8), and the acceleration due to gravity (9.8 m / s²) are used to calculate the dynamic interaction behavior of the robot during motion. For example, the coefficient of friction directly affects the sliding resistance when the robot end contacts the workpiece.

[0088] 2. Parameter input and verification

[0089] Parameters can be obtained in two ways:

[0090] Manual user input: Key parameters are manually set for specific task requirements. For example, in the jack task, the user enters the hole size of Φ5mm and specifies the material as aluminum alloy, and the system automatically associates the density parameter of 2700kg / m³.

[0091] Template library call: retrieve similar scene templates from the preset template library and extract verified parameter combinations. For example, in a handling task, call the "standard bowl-shaped workpiece" template and directly load its geometric dimensions and material parameters to reduce repeated configuration time.

[0092] The parameters obtained need to be verified for consistency, such as checking whether the aperture size is within the allowable error range (such as ±0.1mm), or whether the friction coefficient conforms to physical common sense (such as 0.3±0.05), to avoid simulation distortion due to parameter errors.

[0093] 3. Parameter passing and storage

[0094] The validated parameters will be passed to subsequent steps (such as the geometric model generation in step 1013 ) and stored in a simulation environment configuration file in a structured data format (such as JSON or YAML).

[0095] Here is a specific example:

[0096] Example 1: Jack task parameter configuration

[0097] In the jacking task, the user needs to build a training scene of a Φ5mm deep hole. Obtain the following parameters by manually inputting or calling the template library:

[0098] Robot parameters: joint stiffness 0.5N·m / rad, damping coefficient 0.1N·s / m, control mode is position control;

[0099] Scene parameters: aperture Φ5mm±0.1mm, material is aluminum alloy (density 2700kg / m³);

[0100] Physical parameters: friction coefficient 0.3, restitution coefficient 0.8.

[0101] These parameters are passed to the implicit rendering equation to generate the SDF function of the hole wall, and the collision force between the shaft and the hole wall is calculated through the contact impact model. For example, when the robot end is inserted into the hole at a speed of -0.1m / s, the system calculates the normal contact force to be 12N and the impact recovery coefficient to be 0.85 based on the parameters, ensuring that the impact force attenuation in the simulation conforms to the real physical laws.

[0102] Example 2: Reuse of bowl-shaped workpiece handling scenarios

[0103] For the task of handling bowl-shaped workpieces, call the "Φ12cm standard bowl" template from the template library and automatically load its parameters:

[0104] Geometric parameters: bowl diameter Φ12cm, center of gravity height 5cm;

[0105] Material parameters: ceramic material (density 2500kg / m³), surface roughness Ra 0.8μm;

[0106] Physical parameters: friction coefficient 0.25, restitution coefficient 0.75.

[0107] If the size needs to be adjusted, the user enters the correction command (such as enlarging the bowl mouth diameter to Φ15cm), and the system dynamically updates the parameters and regenerates the geometric model to ensure that the contact force distribution when the gripper grasps meets the new size requirements.

[0108] By accurately obtaining the environment configuration parameters, step 1011 achieves the following technical effects:

[0109] 1. Improved simulation accuracy: Parameter-driven geometric model generation keeps the hole wall surface error within 5%. Compared with the traditional polyhedron convex decomposition method (error ≥ 10%), the robot grasping success rate is increased from 80% to 95%.

[0110] 2. Efficiency optimization: Template library calls reduce scene construction time by 60%. For example, the jack task scene is shortened from traditional modeling of several hours to 30 minutes.

[0111] 3. Enhanced generalization capability: The dynamic parameter correction mechanism supports reuse across task scenarios. For example, the bowl-shaped workpiece template can be adapted to different sizes (Φ10cm to Φ20cm) and materials (ceramic, plastic), reducing the cost of repeated modeling.

[0112] "Obtaining environmental configuration parameters" provides high-precision, reusable basic data support for the construction of robot training scenarios through structured parameter input, template library call and dynamic verification. Whether it is the precise parameter control of the jack task or the rapid reuse of the bowl-shaped workpiece template, this step plays a core role in improving simulation authenticity and training efficiency.

[0113] Step 1012: Determine the corresponding scene object in the training scene according to the environment configuration parameters.

[0114] The core purpose of step 1012 is to build a key link in the robot training scene. Its core goal is to convert abstract environmental configuration parameters into specific, operational scene objects (such as holes, bowls, obstacles, etc.), and provide clear object definitions for subsequent geometric model generation and physical simulation. This step ensures that the simulation environment can truly reflect the actual task requirements through accurate mapping of parameters and scene objects, avoiding training deviations caused by ambiguous object definitions. The core actions and functions of the step execution are as follows:

[0115] 1. Parameter parsing and object matching

[0116] The environment configuration parameters contain detailed data about the robot model, scene assets, and physical properties. These parameters need to be parsed to determine the specific objects that need to be built in the scene. For example, in the jack task, the parameter "aperture Φ5mm" corresponds to the scene object "deep hole", while "material aluminum alloy" defines the physical properties of the hole. The system needs to automatically match the preset object template or generate a new object based on the parameter type (such as size, material, shape). For example, if the parameter contains "bowl mouth diameter Φ12cm", the system will call the "bowl-shaped workpiece" template and adjust its size and material based on the parameters.

[0117] 2. Attribute inheritance and dynamic adjustment

[0118] The properties of scene objects must inherit the physical properties (such as density and friction coefficient) in the environment configuration parameters, and support dynamic adjustment. For example, in a handling task, if the user enters "center of gravity offset 0.5cm", the system will adjust the center of gravity position of the bowl-shaped workpiece according to this parameter, and update the torque calculation parameters in its physical interaction model to ensure motion stability in the simulation. In addition, if certain properties (such as surface roughness) are not explicitly specified in the parameters, the system will call the default value or intelligently fill in based on historical data to avoid incomplete object definition due to missing parameters.

[0119] 3. Scene object verification and optimization

[0120] After determining the scene objects, it is necessary to verify whether their parameters meet the physical rationality and task requirements. For example, the aperture size must meet the actual processing tolerance (such as Φ5mm±0.1mm), and the material density must be consistent with the actual material (such as aluminum alloy density 2700kg / m³). If a parameter conflict is detected (such as the aperture is too small so that the robot end cannot be inserted), the system will trigger an alarm and suggest corrections. In addition, adjusting the object layout through optimization algorithms (such as genetic algorithms) can improve the efficiency of scene reuse. For example, in multi-person collaborative training, the system automatically adjusts the distribution density of holes or bowls according to the number of robots and the complexity of the task to avoid spatial conflicts.

[0121] Here are some specific examples:

[0122] Example 1: Jack Task Scene Object Determination

[0123] In the jack task, the environment configuration parameters clearly include "aperture diameter Φ5mm", "material aluminum alloy" and "friction coefficient 0.3". The system performs the following operations based on these parameters:

[0124] Object matching: retrieve the "deep hole" object template from the template library, whose default attributes are cylindrical structure and smooth opening end;

[0125] Property settings: Set the aperture to Φ5mm, the material density inherits the aluminum alloy parameters (2700kg / m³), and the friction coefficient is set to 0.3;

[0126] Dynamic adjustment: If the user further inputs "hole depth 20mm", the system will update the depth parameters of the hole object and recalculate its contact distance with the robot's end effector to ensure that the end can complete the complete insertion action.

[0127] Example 2: Reuse of bowl-shaped workpiece handling scenarios

[0128] For the task of handling bowl-shaped workpieces, the environmental configuration parameters include "bowl mouth diameter Φ12cm", "material ceramic" and "center of gravity height 5cm". The system executes the following process:

[0129] Template call: load the "standard bowl" object from the template library, whose default diameter is Φ12cm and material density is 2500kg / m³;

[0130] Parameter override: If the user enters "Material Ceramic", the system replaces the default material properties and adjusts the center of gravity calculation parameters according to the density of ceramic (2500kg / m³);

[0131] Layout optimization: In a multi-person collaboration scenario, the system automatically adjusts the placement of the bowl according to the number of robots to ensure that the operating spaces of each robot do not interfere with each other. At the same time, it verifies through physical simulation whether the center of gravity offset of the bowl will cause the risk of tipping over.

[0132] Through precise “parameter-object” mapping, step 1012 achieves the following technical effects:

[0133] 1. Improved simulation authenticity: Object attributes are highly consistent with actual physical properties. For example, the center of gravity offset error of a bowl-shaped workpiece is controlled within ±0.2cm, ensuring the stability of the robot’s grasping action.

[0134] 2. Efficiency optimization: Template library call reduces scene object generation time by 70%. For example, the configuration time of hole objects in the jack task is reduced from 1 hour of traditional manual modeling to 15 minutes.

[0135] 3. Enhanced generalization capability: The dynamic adjustment mechanism supports reuse across task scenarios. For example, the "deep hole" template can adapt to different diameters (Φ3mm to Φ10mm) and materials (metal, plastic), reducing the cost of repeated modeling.

[0136] "Determine scene objects based on environment configuration parameters" converts abstract parameters into high-fidelity scene objects through parameter parsing, attribute inheritance and dynamic optimization, laying the foundation for subsequent geometric modeling and physical simulation. Whether it is the precise parameter control of the jack task or the rapid template reuse of the bowl-shaped workpiece, this step plays a core role in improving simulation accuracy and training efficiency.

[0137] Step 1013: Generate a geometric model of the scene object through implicit rendering equations.

[0138] The core purpose of step 1013 is to use mathematical modeling and physical simulation technology to convert the environment configuration parameters into a high-precision scene object geometry model, solve the shape distortion problem caused by the convex decomposition of polyhedrons in traditional methods, and ensure the geometric consistency between the robot training environment and the real physical world. This step achieves accurate expression of complex surfaces and structures through implicit rendering equations (such as modeling methods based on signed distance functions SDF), providing a reliable basic model for subsequent physical interaction simulation. The core actions and functions of the step execution are as follows:

[0139] 1. Definition and calculation of implicit rendering equation

[0140] The implicit rendering equation describes the relationship between the surface of the scene object and the spatial point through a mathematical function. The core is to define the signed distance function (SDF). Taking the deep hole scene as an example, the SDF function can be expressed as:

[0141] ;

[0142] Among them, r hole is the hole radius, z center is the coordinate of the hole center. This function generates an accurate surface model of the hole wall by calculating the shortest distance from the spatial point (x, y, z) to the hole surface. Compared with the traditional polyhedron convex decomposition method (such as simplifying the hole into a cylinder), SDF can capture the curvature change of the hole wall and avoid the robot end grasping deviation caused by geometric approximation.

[0143] 2. Application of the radiation rendering equation

[0144] After the geometric model is generated, the radiation rendering equation is used to simulate the interaction between light and scene objects and calculate the pixel color distribution. The radiation rendering equation integrates the camera light to calculate the transmittance and volume density. The formula is:

[0145] ;

[0146] Where T(t) is the transmittance, is the volume density, and c(r(t),d) is the color of the light direction d at position r(t). Taking a bowl-shaped workpiece as an example, radiosity rendering can accurately simulate the light and shadow distribution on its surface, ensuring that the visual feedback in the simulation is consistent with the real material, such as the highlight and diffuse reflection characteristics of the ceramic bowl surface.

[0147] 3. Dynamic optimization of geometric models

[0148] The generated geometric model needs to be dynamically optimized in combination with physical parameters. For example, in the task of inserting the shaft into the hole, if the collision force distribution between the robot end and the hole wall is found to be abnormal in the simulation, the system will optimize the geometric model by adjusting the parameters of the SDF function (such as the aperture size or curvature smoothness) to ensure that the simulation results match the real physical interaction. In addition, for complex shapes (such as workpieces with grooves), the implicit rendering equation supports hierarchical modeling, and achieves accurate expression of multi-level geometric features by superimposing multiple SDF functions.

[0149] Here are some specific examples:

[0150] Example 1: Jack Task Geometry Generation

[0151] In the jack task, the environment configuration parameters specify that the hole diameter is Φ5mm±0.1mm and the material is aluminum alloy. The system generates the geometric model of the hole through the implicit rendering equation:

[0152] SDF function definition: Taking the hole center as the origin, define the radial distance function to ensure that the hole wall curvature is continuous and meets the tolerance requirements;

[0153] Radiation rendering calculation: simulates the scattering and absorption of light when it passes through the hole wall to generate realistic metal reflection effects;

[0154] Dynamic verification: Through the robot end trajectory test, verify whether the hole model supports accurate grasping. If it is detected that the collision force between the end and the hole wall exceeds the threshold, the system automatically fine-tunes the aperture parameters to Φ5.05mm to ensure the physical consistency between the simulation and the real task.

[0155] Example 2: Modeling of a bowl-shaped workpiece handling scenario

[0156] For the task of handling bowl-shaped workpieces, the environmental configuration parameters include bowl diameter Φ12cm, material ceramic and center of gravity height 5cm. The system performs the following process:

[0157] Surface generation: Build a surface model of the bowl wall based on the SDF function to accurately capture the elliptical curvature of the bowl mouth and the concave structure of the bowl bottom;

[0158] Light and shadow simulation: The diffuse reflection characteristics of the ceramic surface are simulated through the radiation rendering equation to generate light and shadow effects that conform to the real material;

[0159] Layout optimization: In multi-person collaboration scenarios, the system dynamically adjusts the placement of the bowl according to the robot's operating space to avoid collisions, and verifies through physical simulation whether the center of gravity shift will lead to the risk of tipping.

[0160] By generating the geometric model through the implicit rendering equation, step 1013 achieves the following technical breakthroughs:

[0161] 1. Improved geometric accuracy: The SDF function supports millimeter-level precision expression of complex surfaces. For example, the curvature error of a bowl-shaped workpiece is controlled within ±0.05mm. Compared with traditional polyhedron modeling (error ≥1mm), the robot grasping success rate is increased from 75% to 95%.

[0162] 2. Physical consistency guarantee: The light and shadow and material properties (such as metal reflection and ceramic diffuse reflection) simulated by the radiation rendering equation are highly consistent with the real world, reducing training deviations caused by visual errors.

[0163] 3. Dynamic adaptability: By parameterizing the SDF function, the system can quickly adapt to different task requirements. For example, the same hole model can adapt to a variety of jack tasks from Φ3mm to Φ10mm by modifying the radius parameter, increasing reuse efficiency by 60%.

[0164] "Generating the geometric model of scene objects through implicit rendering equations" solves the core pain points of geometric distortion and physical inconsistency in traditional methods through the deep integration of mathematical modeling and physical simulation technology. Whether it is the precise aperture modeling of the jack task or the complex surface generation of the bowl-shaped workpiece, it reflects the key role of this step in improving simulation accuracy and training reliability, laying a solid foundation for subsequent physical interaction simulation and robot skill learning.

[0165] Step 1014: Calculate a physical interaction model of the scene objects based on the contact impact model.

[0166] The core purpose of step 1014 is to accurately simulate the dynamic interaction behavior (such as collision force, friction force, impact force, etc.) between the robot and the scene objects during movement through physical simulation technology, and solve the interaction distortion problem caused by the simplification of physical models in traditional methods. This step realizes the refined modeling of physical interaction through contact impact models (such as viscoelastic contact force model, impact recovery coefficient model), ensures that the robot training environment is consistent with the dynamic behavior of the real physical world, and provides high-fidelity data support for subsequent policy network training. The core actions and functions of the step execution are as follows:

[0167] 1. Mathematical modeling of contact impact model

[0168] The core of the contact impact model is to describe the interaction force between the robot end and the scene object through physical formulas. Taking the shaft insertion task as an example, the viscoelastic contact force model calculates the normal contact force through the following formula:

[0169] ;

[0170] Among them, v n is the robot terminal velocity, x is the normal relative deformation, a rbt is the robot end acceleration, m, k, d are mass, elastic coefficient and damping coefficient respectively. Impact recovery coefficient (e r ) model is through the formula:

[0171] ;

[0172] Dynamically adjust the impact force attenuation to ensure that the impact force distribution in the simulation is consistent with the real physical law. For example, during the shaft insertion process, if the robot end speed is -0.1m / s, the above formula can be used to calculate a reasonable contact force distribution to avoid simulation errors leading to task failure.

[0173] 2. Dynamic binding of physical parameters

[0174] The physical properties of scene objects (such as material density and friction coefficient) need to be dynamically bound to the contact impact model. For example, in a bowl-shaped workpiece handling task, the density of the ceramic material (2500kg / m³) and the friction coefficient (0.25) will directly affect the calculation of the collision force. The system automatically updates the relevant coefficients in the physical interaction model by reading the environment configuration parameters in real time to ensure that the simulation behavior matches the actual task requirements.

[0175] 3. Real-time verification and optimization of interactive behaviors

[0176] The generated physical interaction model needs to be verified through simulation to verify its consistency with real physics. For example, in the jacking task, if the simulation finds that the peak collision force between the robot end and the hole wall exceeds the actual threshold (such as more than 15N), the system will automatically adjust the damping coefficient in the contact impact model (such as from 0.1 to 0.15) until the error between the simulation result and the real physical data is less than 5%. In addition, for complex interaction scenarios (such as multi-object collision), the system supports hierarchical modeling, which achieves accurate expression of multi-level interaction behaviors by superimposing multiple contact impact models.

[0177] Here are some specific examples:

[0178] Example 1: Collision Force Simulation in a Jack-in-Jack Task

[0179] In the jack task, the environment configuration parameters specify that the hole diameter is Φ5mm and the material is aluminum alloy (density 2700kg / m³). The system performs the following operations based on the contact impact model:

[0180] Collision force calculation: When the end of the robot is inserted into the hole at a speed of -0.1m / s, the normal contact force is calculated to be 12N and the impact recovery coefficient is 0.85 by the formula;

[0181] Dynamic adjustment: If the contact force between the shaft and the hole wall is found to be unevenly distributed during simulation (such as local pressure exceeding the material strength limit), the system automatically optimizes the elastic coefficient of the contact impact model (such as adjusting from 500N / m to 550N / m) to ensure that the simulated force distribution is consistent with the real physics;

[0182] Verification results: Through multiple iterative simulations, the collision force error was ultimately controlled within ±3%, and the robot’s grasping success rate was increased from 80% to 95%.

[0183] Example 2: Friction simulation for handling bowl-shaped workpieces

[0184] For the task of handling bowl-shaped workpieces, the environmental configuration parameters include the bowl diameter of Φ12cm and the material of ceramic (friction coefficient 0.25). The system performs the following process:

[0185] Friction modeling: Define the sliding resistance between the robot end and the bowl wall based on the friction coefficient;

[0186] Dynamic optimization: If the simulation finds that the sliding angle of the bowl during handling exceeds the allowable range (e.g., more than 5°), the system automatically adjusts the friction coefficient to 0.28 and recalculates the restitution coefficient in the contact impact model to ensure the stability of the bowl;

[0187] Scenario reuse: In multi-person collaborative training, the system dynamically adjusts the center of gravity position and friction parameters of the bowl according to the size of the end effectors of different robots (such as the opening and closing range of the mechanical claws) to adapt to multi-task requirements.

[0188] Through the refined modeling of the contact impact model, step 1014 achieved the following technical breakthroughs:

[0189] 1. Improved physical interaction accuracy: The combination of the viscoelastic model and the impact recovery coefficient controls the collision force error within ±5%. Compared with the traditional constant friction model (error ≥ 15%), the robot mission success rate is increased by 20%;

[0190] 2. Enhanced dynamic adaptability: Through the automatic parameter optimization mechanism, the system can quickly adapt to different materials and task requirements. For example, the same contact impact model can be adapted to metal (hard contact) and silicone (soft contact) scenarios by adjusting the elastic coefficient, and the reuse efficiency is increased by 70%;

[0191] 3. Ensure the authenticity of simulation: The interaction force calculation based on physical laws ensures that the robot training results can be directly transferred to the real environment. For example, in the task of handling bowl-shaped workpieces, the center of gravity offset error in the simulation is less than 0.1cm, and the actual grasping stability of the robot is improved by 30%.

[0192] "Calculating the physical interaction model of scene objects based on the contact impact model" solves the core pain point of physical interaction distortion in traditional methods through refined physical modeling and dynamic parameter optimization. Whether it is the precise collision force simulation of the jack task or the complex friction simulation of the bowl-shaped workpiece, it reflects the key role of this step in improving simulation accuracy and training reliability, providing high-fidelity and adaptive physical environment support for robot skill learning.

[0193] Step 1015: Integrate and embed the geometric model and the physical interaction model into the initial simulation environment, and use the rendered initial simulation environment as a training scene.

[0194] The core purpose of step 1015 is to deeply integrate the geometric models (such as holes, bowls and other scene objects) generated in the early stage with the physical interaction models (such as dynamic behaviors such as collision force and friction) to build a high-fidelity, interactive robot training environment. This step integrates the physics engine, rendering engine and scene management module to ensure the authenticity and real-time nature of physical interactions during robot training, and provide a reliable environment support for subsequent policy network training. The core actions and functions of the step execution are as follows:

[0195] 1. Model integration and physics engine configuration

[0196] Import the geometric model generated by the implicit rendering equation (such as the hole wall surface described by the SDF function) and the physical interaction parameters (such as elastic coefficient and restitution coefficient) calculated by the contact impact model into the physics engine (such as Isaac Gym). For example, in the jacking task, the geometric model of the hole needs to be bound to the normal force calculation module in the contact impact model to ensure that the distribution of the collision force when the robot end is inserted conforms to the real physical laws. The physics engine initializes the physical rules of the scene by configuring parameters (such as gravity acceleration and friction coefficient). For example, the friction coefficient of the bowl-shaped workpiece is set to 0.25 to simulate the sliding resistance of the real ceramic surface.

[0197] 2. Real-time data synchronization and dynamic interaction

[0198] In the simulation environment, the position and speed of the robot end must be synchronized with the physical state of the scene objects (such as hole wall deformation and bowl center of gravity offset) in real time. For example, when the robot end is inserted into the hole at a speed of -0.1m / s, the physical engine needs to dynamically calculate the deformation and collision force of the hole wall, and feed the data back to the strategy network model to drive the robot to adjust the action strategy. Dynamic interaction also supports multi-object collision scenarios, such as the calculation of the contact force between a bowl-shaped workpiece and other workpieces during transportation, to ensure that the dynamic behavior of the simulation environment is consistent with the real task.

[0199] 3. Scenario verification and parameter optimization

[0200] The accuracy of the model is verified through multiple simulation iterations. For example, in a bowl-shaped workpiece handling task, if the simulation finds that the center of gravity of the bowl is offset, resulting in a risk of tipping, the coefficient of restitution in the physical interaction model needs to be adjusted (such as from 0.8 to 0.75) and the support point distribution of the geometric model needs to be recalculated until the error between the simulation result and the real physical data is less than 5%. The optimized scene needs to be checked for matching between the geometric structure and the physical behavior through visualization tools (such as the debugging interface of Isaac Sim) to ensure the reliability of the training environment.

[0201] 4. Scenario deployment and multi-task adaptation

[0202] Deploy the integrated simulation environment to the robot training platform (such as a distributed system based on Docker containers) to support multi-task parallel training. For example, in a multi-person collaborative scenario, the asset team calls the "Φ5mm hole handling" template, the body team loads the robot model parameters, and the algorithm team injects the policy network. The three synchronize the scene status in real time through the shared data volume. The physics engine needs to adapt to the size and control mode (such as position control or torque control) of different robot end effectors to ensure seamless switching of multi-task scenarios.

[0203] Here are some specific examples:

[0204] Example 1: Socket Task Scenario Integration

[0205] In the jacking task, the geometric model (SDF function of Φ5mm deep hole) and the physical interaction model (contact force formula between shaft and hole wall) are integrated into the Isaac Gym physics engine:

[0206] Physical rule configuration: set the hole wall material density to 2700kg / m³, the friction coefficient to 0.3, and the restitution coefficient to 0.85;

[0207] Dynamic interaction verification: When the robot end is inserted into the hole at a speed of -0.1m / s, the physical engine calculates the normal contact force in real time as 12N and the impact recovery coefficient as 0.85, ensuring that the simulated force distribution is consistent with the real physics;

[0208] Multi-task adaptation: Call the "Φ5mm hole" template through the template library to quickly generate hole scenes of different materials (such as stainless steel, plastic). You only need to adjust the density and friction coefficient parameters without re-modeling.

[0209] Example 2: Bowl-shaped workpiece handling scenario deployment

[0210] For the bowl-shaped workpiece handling task, the integrated scenario needs to support multi-person collaborative training:

[0211] Binding geometry and physics: The geometric model of the bowl (surface described by SDF function) is bound to the physical interaction model (center of gravity offset 0.5cm, friction coefficient 0.25) to ensure that the sliding and tipping behaviors during transportation conform to the real physical laws;

[0212] Real-time data synchronization: The trajectory data of the robot grabbing the bowl is transmitted to the physics engine through ROS 2, and the rotation angle and contact force distribution of the bowl are dynamically updated;

[0213] Distributed training support: By deploying a shared data volume in a Docker container, the asset team can obtain the latest scene data without retraining after updating the material parameters of the bowl, improving training efficiency by 40%.

[0214] Through model integration and scenario deployment, step 1015 achieved the following technical breakthroughs:

[0215] 1. High-fidelity interactive environment: The deep binding between the physical engine and the geometric model controls the collision force error within ±3%. Compared with the traditional separate modeling method (error ≥ 15%), the robot mission success rate is increased by 25%;

[0216] 2. Dynamic adaptation capability: Through parameterized configuration and template library call, the same scene can adapt to multiple task requirements (such as Φ5mm to Φ10mm aperture, workpieces of different materials), and the reuse efficiency is increased by 60%;

[0217] 3. Low-cost multi-task training: Docker-based containerized deployment supports simultaneous training by a team of 10, reducing hardware costs by 50%, and GPU acceleration of the physics engine (such as RTX 4090) increases training speed by 3 times.

[0218] "Integrating geometric models with physical interaction models and embedding them into the simulation environment" solves the core pain points of physical behavior distortion and low scene reuse efficiency in traditional methods through physical engine drive, dynamic data synchronization and multi-task adaptation. Whether it is the precision force control simulation of the jack task or the dynamic handling simulation of the bowl-shaped workpiece, this step plays a key role in building a high-precision, adaptive robot training environment, providing a reliable foundation for the efficient learning of the policy network.

[0219] Alternatively, see Figure 3 It can be seen that step 101, configuring parameters according to the simulation environment and building a training scene based on the implicit rendering equation and the contact impact model, includes:

[0220] Step 1016: retrieve the most similar initial scene template from the template library according to the simulation environment configuration parameters.

[0221] The core purpose of step 1016 is to quickly match the basic scene model closest to the current task requirements through the preset scene template library to reduce the cost of repeated modeling. This step is based on the intelligent matching of environmental configuration parameters (such as robot size, workpiece geometric features, and physical properties) with the historical scene templates stored in the template library to ensure that the generated initial scene can inherit the physical consistency of the existing templates and adapt to the new task requirements through local corrections, providing an efficient starting point for subsequent dynamic adjustments. The core actions and functions of the step execution are as follows:

[0222] 1. Template library construction and parameter indexing

[0223] The template library stores preset templates for a variety of typical scenarios (such as deep holes, bowl-shaped workpieces, handling tracks, etc.). Each template contains a complete geometric model, physical interaction parameters and metadata (such as material density, friction coefficient). The system achieves efficient retrieval through the following methods:

[0224] Keyword matching: filter candidate templates based on object types (such as "hole" and "bowl") in environment configuration parameters;

[0225] Parameter threshold comparison: Calculate the similarity between the current parameters and the template parameters (such as aperture deviation ≤ ±0.5mm, material density error ≤5%), and select the template with the highest matching degree. For example, when the user enters "Φ5mm hole handling", the system will first search for templates with apertures of Φ4.5mm to Φ5.5mm.

[0226] 2. Similarity calculation and sorting

[0227] A weighted scoring mechanism is used to comprehensively sort candidate templates, with weights including geometric similarity (40%), physical parameter matching (30%), and historical mission success rate (30%). For example, a template with an aperture of Φ5mm (complete match), a material density of 2700kg / m³ (error +2%), and a historical success rate of 95% will have a higher comprehensive score than a template with an aperture of Φ6mm but matching other parameters.

[0228] 3 Dynamic Correction Instruction Generation

[0229] If no completely matching template is found, the system automatically generates correction instructions. For example, when the user's requirement is "Φ6mm hole handling" but the template library only contains Φ5mm templates, the system prompts: "A hole diameter difference is detected. Do you want to scale up the template size by 1.2 times?" Correction instructions can be generated based on user confirmation or algorithm recommendation to ensure the reusability of the initial template.

[0230] Here are some specific examples:

[0231] Example 1: JACK task template call

[0232] In the jack task, the environment configuration parameters specify that the hole diameter is Φ5mm±0.1mm and the material is aluminum alloy (density 2700kg / m³). The system performs the following operations:

[0233] Template retrieval: The "Φ5mm deep hole" template was selected from the template library, with a historical success rate of 95% and a physical parameter matching degree of 98%;

[0234] Direct reuse: No modification is required, and the template can be directly called as the initial scene. Only the corresponding SDF geometry model and contact impact parameters need to be loaded;

[0235] Efficiency comparison: Compared with traditional manual modeling, template reuse reduces scene construction time from 2 hours to 10 minutes.

[0236] Example 2: Bowl-shaped workpiece template correction

[0237] For the task of handling bowl-shaped workpieces, the user inputs the requirements as "Φ15cm bowl diameter, ceramic material". The system operates as follows:

[0238] Preliminary search: There is a "Φ12cm standard bowl" template in the template library, the material is ceramic, but the size does not match;

[0239] Similarity calculation: The aperture deviation is +25% (Φ12→Φ15cm), which exceeds the default threshold of ±10%, triggering the correction process;

[0240] Correction instruction generation: The system recommends enlarging the template size by 1.25 times and adjusting the center of gravity position parameters;

[0241] Manual confirmation: After the user accepts the correction, the system generates a Φ15cm bowl template, which inherits the physical interaction parameters of the original template (such as friction coefficient of 0.25). It only needs to verify the collision force distribution under the new size.

[0242] Through template library retrieval and dynamic correction, step 1016 achieves the following technical breakthroughs:

[0243] 1. Significantly improved efficiency: The construction time of repeated task scenarios is reduced by 80%. For example, the template reuse of the transport task reduces the modeling time from 3 hours to 0.5 hours.

[0244] 2. Physical consistency guarantee: The physical parameters of the inherited template (such as material density and friction coefficient) reduce the simulation error to within ±3%, which increases the mission success rate by 20% compared with the new modeling (error ≥ 10%);

[0245] 3. Flexible adaptation capability: supports cross-task template migration. For example, the "Φ5mm hole" template can be enlarged to adapt to the Φ10mm requirement by adjusting the geometric parameters without redesigning the physical interaction model.

[0246] "Retrieving the most similar initial scene template from the template library based on the simulation environment configuration parameters" solves the core pain points of low scene modeling efficiency and easy distortion of physical parameters in traditional methods through intelligent matching and dynamic correction mechanisms. Whether it is the precise template reuse of the jack task or the dynamic adjustment of the size of the bowl-shaped workpiece, this step plays a key role in improving scene construction efficiency and training reliability, providing efficient and scalable basic support for robot skill learning.

[0247] Step 1017: Obtain a correction instruction based on the initial scene template, and correct the initial scene template according to the correction instruction. The corrected initial scene template is used to construct a training scene.

[0248] The core purpose of step 1017 is to dynamically adjust the parameters and structure of the preset template to adapt it to specific task requirements, solving the problem of insufficient flexibility in scene reuse caused by fixed templates in traditional methods. This step drives the local optimization of the template by correcting the instructions, ensuring that the initial scene can inherit the physical consistency of the template and meet the differentiated requirements of the new task through parameter adjustment, providing a highly adaptable simulation environment for subsequent training. The core actions and functions of the step execution are as follows:

[0249] 1. Correction instruction generation and parsing

[0250] Correction instructions are generated based on the difference between the matching degree of the initial scene template and the task requirements. For example, in the jack task, if the retrieved template aperture is Φ5mm, but the actual requirement is Φ6mm, the system will automatically generate a correction instruction: "Enlarge the aperture by 1.2 times, and keep the material density of aluminum alloy at the default value of 2700kg / m³". Correction instructions can be generated in the following ways:

[0251] Automatic matching algorithm: Calculate the scaling ratio or displacement based on parameter deviation (such as aperture error ±0.5mm);

[0252] Manual input by the user: For complex adjustment requirements (such as asymmetric structural deformation), the user directly enters the correction parameters (such as "the diameter of the bowl mouth increases by 2cm").

[0253] 2. Dynamic modification and verification of templates

[0254] The correction instructions act on the geometric model and physical parameters of the initial template to generate scene objects adapted to the new task. For example, for a bowl-shaped workpiece handling task, if the initial template is a Φ12cm ceramic bowl, but the required size is Φ15cm and the center of gravity needs to be offset by 0.5cm, the system will:

[0255] Geometric adjustment: Proportionally enlarge the bowl diameter to Φ15cm, and recalculate the bowl wall curvature and center of gravity position;

[0256] Physics parameters updated: Inherit the friction coefficient of the original template (0.25), but adjust the restitution coefficient according to the new size (e.g. from 0.8 to 0.75 to match the impact force attenuation of the larger size).

[0257] After correction, it needs to be verified through physical simulation, such as simulating the risk of the bowl tipping over during transportation, to ensure that the error of the corrected center of gravity offset is less than 0.1cm.

[0258] 3. Scenario adaptability and physical consistency assurance

[0259] The modified template needs to ensure that the physical interaction behavior is consistent with the real task. For example, in the task of inserting the shaft into the hole, if the template aperture is enlarged to Φ6mm, the system needs to recalculate the elastic coefficient in the contact impact model (such as adjusting from 500N / m to 550N / m) to match the change in contact force distribution caused by the larger aperture. In addition, the modified geometric model needs to regenerate the SDF function through the implicit rendering equation to avoid surface distortion caused by size changes.

[0260] Here are some specific examples:

[0261] Example 1: Socket Task Template Modification

[0262] In the jack task, the template initially retrieved is a Φ5mm deep hole, but the actual requirement is Φ6mm. The system performs the following operations:

[0263] Correction command generation: Based on the aperture deviation (+1mm), the zoom command "aperture enlargement 1.2 times" is automatically generated;

[0264] Geometry adjustment: Increase the radius of the original hole model from 2.5mm to 3mm, and regenerate the SDF function to describe the hole wall surface;

[0265] Physical parameter adaptation: The restitution coefficient in the contact impact model is adjusted to 0.85 according to the new aperture to ensure that the impact force attenuation during shaft insertion conforms to the real physical laws;

[0266] Verification results: Simulation tests show that the modified hole model increases the robot's grasping success rate from 85% to 95%.

[0267] Example 2: Dynamic adjustment of the template for a bowl-shaped workpiece

[0268] For the task of handling bowl-shaped workpieces, the initial template is a Φ12cm ceramic bowl, but the requirement is Φ15cm and the center of gravity height needs to be lowered:

[0269] Geometric correction: Proportionally enlarge the diameter of the bowl mouth to Φ15cm, and adjust the thickness of the bowl bottom to maintain structural strength;

[0270] Physical parameter update: Inherit the friction coefficient of the original template (0.25), but adjust the restitution coefficient to 0.78 according to the new center of gravity position (offset by 0.5cm);

[0271] Interactive verification: By simulating the sliding angle of the bowl during the handling process, it was found that the center of gravity offset error after correction was 0.08cm, which met the task requirements.

[0272] By modifying the dynamic adjustment of the template driven by the instruction, step 1017 achieves the following technical breakthroughs:

[0273] 1. Improved scene reuse efficiency: Template correction reduces scene construction time by 70%. For example, a Φ15cm bowl scene can be adapted in just 5 minutes, while traditional modeling takes 2 hours.

[0274] 2. Physical consistency guarantee: inherit the physical parameters of the template (such as material density and friction coefficient) and dynamically optimize key coefficients (such as the coefficient of restitution) to control the simulation error within ±3%. Compared with the new modeling (error ≥ 10%), the mission success rate is increased by 20%;

[0275] 3. Flexible adaptation capability: supports cross-task template migration. For example, the "Φ5mm hole" template can be adapted to the Φ10mm requirement by enlarging the size and adjusting the parameters, increasing the reuse efficiency by 60%.

[0276] "Obtaining correction instructions based on the initial scene template and correcting the template" solves the core pain points of poor scene reuse flexibility and difficulty in ensuring physical consistency in traditional methods through dynamic parameter adjustment and physical simulation verification. Whether it is the aperture adaptation of the jack task or the size and center of gravity adjustment of the bowl-shaped workpiece, this step plays a key role in improving scene construction efficiency and training reliability, providing an efficient and accurate simulation foundation for robot skill learning.

[0277] Regarding step 1016 and step 1017, the present disclosure proposes a modular scene self-generation method, which realizes rapid reuse and dynamic adjustment of similar scenes by precipitating existing scenes as standardized templates, and significantly improves the efficiency of simulation scene construction. The specific process is as follows:

[0278] Step 1-1, modular scene serial number. All modular scenes are numbered with positive integers from small to large. The modular scene serial number is The total number of scenes is Modular scene overview The modular scene number is composition.

[0279] Step 1-2: Matrix of scene numbers to be built Extract the keywords of the scene to be built, and use the text fuzzy search method to find the existing modular scenes, and record the scene sequence number set of the search results , is the modular scene number, and the number of elements in the set is .Will Each element in Diagonal matrix, which is the sequence matrix of the scenes to be built .

[0280] Steps 1-3: Set the scene labels Select all the serial numbers in the , select all the related modular scenes and splice them. Since the shape of each scene is a cube, the splicing method is: scale the size of all scenes to the largest scene size, and then splice them in a linear shape to finally form a linear area.

[0281] Step 1-4: Verify the matching degree between the current generated scene and the demand. Therefore, calculate the scene to be built and the overall modular scene All existing scenes match . is the sequence number matrix of the scenes to be built, calculated by steps 1-2. For modular scene overall The serial number of an existing modular scene in composition The diagonal matrix of Calculate scene matching .

[0282] ;

[0283] The smaller the value, the less match there is; the larger the value, the more match there is.

[0284] If the matching degree meets the requirements, jump to step 1-7 to automatically generate the scene, otherwise go to step 1-5 to step 1-6 and modify the generated scene.

[0285] Steps 1-5: Manually modify the scene content, such as the position, quantity, size, etc. of each model in the scene, so that the entire scene is more in line with the requirements. The modified content is determined by the correction coefficient. Record,

[0286] ;

[0287] in, Indicates the total number of models in the scene, is the model number, They are the correction coefficients for the position, quantity, and size of each model:

[0288] ;

[0289] ;

[0290] ;

[0291] in, They are manual correction coefficients, which are adjusted according to actual needs. The correction coefficient is determined by the matching degree calculated in steps 1-4 and the manual correction coefficient. It can automatically approach the direction that needs to be adjusted while retaining space for manual correction, and has good adaptability.

[0292] Correct the scene number matrix that needs to be built , ;

[0293] This is the revised sequence number matrix of the scenes that need to be built.

[0294] Step 1-6: Display the corrected scene, and repeat step 1-4 to recalculate the matching degree until the matching degree meets the requirements, and then proceed to step 1-7 to generate the scene.

[0295] Step 1-7, confirm the automatic scene generation. Call the existing modular scene or the revised scene model, splice and display it according to the method of steps 1-3, and complete the automatic generation of the new scene.

[0296] Step 1-8: Manually correct the scene. Manually adjust the scene arrangement, such as the position of the spliced ​​scene, to make the entire scene more in line with the requirements.

[0297] Step 102: Control the robot to perform the training task based on the proximal strategy optimization algorithm in the training scenario.

[0298] The core purpose of step 102 is to drive the autonomous learning and dynamic optimization of the robot policy network through the reinforcement learning framework (proximal policy optimization, PPO), to solve the problem of insufficient generalization ability caused by the fixed reward function in the traditional method. This step ensures that the robot can quickly adapt to complex task requirements (such as high-precision jacks, smooth handling) by dynamically adjusting the reward weights, mixed data training and model evaluation mechanisms, and improves the robustness and task success rate of the policy network in the real environment. The core actions and functions of the step execution are as follows:

[0299] 1. Strategy network initialization and reward function loading

[0300] At the beginning of training, the system initializes the robot's state vector (such as joint angle, end position) and action vector (such as joint torque, moving speed) according to the task requirements, and loads the preset reward function. For example, in the jack task, the reward function may include a safety impact force term (weight 0.4), an end position accuracy term (weight 0.3), and an energy efficiency term (weight 0.3). These weight coefficients are set based on historical task data or manual experience to provide a benchmark for subsequent dynamic adjustments.

[0301] 2. Dynamic Optimization of Reward Function Weights

[0302] Based on the historical success rate of completed training tasks (e.g., the success rate of the past 100 jacking tasks was 85%), the system dynamically corrects the weight coefficients of each reward item in the reward function. For example, if it is detected that the robot's success rate has dropped due to excessive impact force, the weight of the safety impact force item is increased to 0.5, while the weight of the energy efficiency item is reduced to 0.2. This process analyzes the achievement rate of the target reward item (e.g., "impact force ≤ 0.5N") and optimizes the weight distribution using the gradient ascent method to ensure that the reward function always focuses on key performance indicators.

[0303] 3. Hybrid data training and policy network update

[0304] To balance exploration and exploitation, the system combines online real-time data (such as joint torque data in the current training cycle) with offline historical success samples (such as past successful socket trajectory data), and mixes them in a preset ratio (such as 7:3) to train the policy network. For example, in a handling task, if the online data shows that the robot's center of gravity offset increases, the weight of the trajectory with a stable center of gravity in the offline success sample is increased, and the network parameters are updated through the PPO loss function (Formula 1):

[0305] ;

[0306] in, is the probability ratio before and after the strategy update, A t is the advantage function, is the clipping threshold, is the entropy regularization coefficient. This process avoids model oscillation caused by single data deviation by limiting the amplitude of policy update.

[0307] 4. Model evaluation and adaptive adjustment

[0308] After each preset number of training times (such as 1000 iterations), the system calculates a comprehensive evaluation score based on the task success rate (such as the current success rate of 92%), the cumulative value of the reward function (such as the cumulative reward value of 5000), and the stability of the policy network (such as KL divergence ≤ 0.01). If the score is lower than the threshold (such as 80 points), the model is retrained; if the score continues to decrease, the network structure is adjusted (such as increasing the number of hidden layer nodes). For example, in the task of handling bowl-shaped workpieces, if the robot's success rate decreases due to friction prediction deviation, the system will automatically expand the state vector dimension, add surface roughness parameters, and optimize the new parameter weights through the gradient descent method.

[0309] Example 1: Dynamic reward adjustment for the jack task

[0310] In the Φ5mm deep hole jacking task, the initial reward function set the safety impact force weight to 0.4. After 500 trainings, the system detected that the failure rate caused by excessive impact force reached 15%, so it dynamically increased the weight to 0.6 and reduced the energy consumption weight to 0.1. At the same time, trajectory data with impact force ≤0.5N were selected from historical successful samples and mixed with online data for training at a ratio of 7:3. Two weeks later, the task success rate increased from 80% to 95%, and the impact force distribution error in the simulation was controlled within ±3%.

[0311] Example 2: Multimodal data fusion for a bowl handling task

[0312] For the task of moving a Φ15cm ceramic bowl, the system is initialized with a reward function that includes the center of gravity offset (±0.5cm). During training, it was found that when the diameter of the bowl mouth was enlarged to Φ18cm, the original center of gravity model failed. At this time, the system calls the "large-size bowl" offline samples (such as 100 sets of stable trajectories) from the template library, mixes them with the online data in a 6:4 ratio, and expands the state vector to include the bowl bottom thickness parameter. Through PPO entropy regularization ( = 0.05) encouraged exploration, which ultimately increased the success rate of moving large bowls from 70% to 88%.

[0313] Through dynamic weight adjustment and mixed data training, step 102 achieves the following beneficial effects:

[0314] 1. Enhanced task adaptability: Dynamic optimization of reward function weights increases the model’s response speed to key indicators (such as impact force and accuracy) by 40% and increases the task success rate by 25%;

[0315] 2. Data efficiency optimization: The hybrid training strategy reduces the reliance on labeled data, the offline sample reuse rate can reach 60%, and the training cycle is shortened by 30%;

[0316] 3. Improved model robustness: Through KL divergence constraints and adaptive network structure adjustment, the generalization ability of the strategy in unknown scenarios (such as asymmetric apertures) is enhanced, and the success rate of real environment migration is increased by 35%.

[0317] "Controlling robots to perform training tasks based on proximal policy optimization algorithms" solves the pain points of reward solidification and weak generalization in traditional reinforcement learning through dynamic reward optimization, hybrid data-driven and adaptive model adjustment. Whether it is the precise force control of the jack task or the dynamic parameter adaptation of the bowl-shaped handling, it reflects the core role of this step in improving the robot's autonomous learning efficiency and task reliability, and provides an efficient solution for skill training in complex industrial scenarios.

[0318] Alternatively, see Figure 4 It can be seen that step 102, controlling the robot to perform the training task based on the proximal strategy optimization algorithm in the training scenario includes:

[0319] Step 1021: Initialize the robot's state vector and action vector according to the training task, and load the reward function corresponding to the training task.

[0320] The core purpose of step 1021 is to provide the initial physical state and goal-oriented incentive mechanism for the robot policy network training. By accurately defining the initial motion parameters of the robot (such as joint angles, end positions) and task-related reward functions (such as impact force limits, accuracy requirements), it is ensured that the policy network can start learning around the core goals of the task at the beginning of training, avoid invalid exploration, and improve training efficiency and convergence speed. The core actions and functions of the step execution are as follows:

[0321] 1. Initialization of state vector and action vector

[0322] The state vector contains the robot's current motion state parameters, such as:

[0323] Joint angles (such as the rotation angles of each joint of a six-axis robot);

[0324] End effector position (e.g. 3D coordinates of the gripper);

[0325] Velocity and acceleration (such as end-movement velocity and joint angular velocity).

[0326] Action vectors define the actions that the robot can perform, for example:

[0327] Joint torque output (such as the driving torque of each joint);

[0328] End effector motion instructions (such as linear movement speed or rotation angle).

[0329] In the plug-in task, the initial state vector must include the position parameters of the robot end aligning with the center of the hole, and the action vector must support control instructions for high-precision plug-in and pull-out actions.

[0330] 2. Loading and binding of reward function

[0331] The reward function presets the weights of key performance indicators according to task requirements, for example:

[0332] Safety impact force item (weight 0.4): limits the collision force between the robot end and the scene object to not exceed the threshold (e.g. ≤10N);

[0333] End position accuracy item (weight 0.3): The degree of proximity between the end point and the target position is rewarded;

[0334] Energy efficiency term (weight 0.3): Penalizes the energy consumption of joint torque output.

[0335] These weight coefficients are set based on historical task data or expert experience, providing a benchmark for subsequent dynamic adjustments. For example, in a handling task, if the fragility of a ceramic bowl is required to be high, the system will increase the initial weight of the safety impact force item to 0.5 and reduce the energy consumption weight to 0.2.

[0336] 3. Mapping of physical parameters and mission objectives

[0337] The initialization of the state vector and action vector needs to be tightly coupled with the reward function. For example, in the bowl-shaped workpiece handling task:

[0338] The state vector must contain the relative distance between the center of gravity of the bowl and the end grip point;

[0339] The motion vector needs to support dynamic adjustment of the grasping force;

[0340] The reward function needs to include the center of gravity offset (such as ≤0.5cm) and grasping stability (such as no slip) into the scoring system.

[0341] This mapping ensures that the robot can perceive the key constraints of the task from the early stages of training and reduce the exploration of invalid actions.

[0342] Here are some specific examples:

[0343] Example 1: Initialization configuration of the jack task

[0344] In the Φ5mm deep hole jack task, step 1021 performs the following operations:

[0345] State vector initialization: Set the initial position of the robot end to 10 cm above the center of the hole, and the angles of each joint to the preset starting posture of the jack;

[0346] Action vector definition: supports precise plug-in and pull-out motion of the end along the Z axis, and the torque output is limited to 0.5N·m to prevent overload;

[0347] Reward function loading: preset safety impact force weight 0.4, position accuracy weight 0.3, energy consumption weight 0.3. When the robot inserts into the hole at a speed of -0.1m / s, if the impact force exceeds 8N, the system immediately triggers the penalty mechanism to guide the policy network to learn gentle insertion action.

[0348] Example 2: Parameter binding for a bowl-shaped transport task

[0349] For the Φ15cm ceramic bowl handling task, the initialization phase:

[0350] The state vector includes the bowl diameter, center of gravity height, and the coordinates of the end grabbing point;

[0351] The motion vector supports multi-level adjustment of the grasping force (e.g. 0.1N to 10N);

[0352] The reward function uses the center of gravity offset (weight 0.5) and the bowl wall slip (weight 0.3) as core indicators, and inherits the historical parameters of "fragile handling" in the template library (such as the friction coefficient of 0.25). If the simulation finds that the center of gravity offset of the bowl exceeds 0.8cm, the system will automatically increase the weight of the center of gravity stability item to 0.6, forcing the policy network to optimize the grasping posture.

[0353] Through precise initialization configuration, step 1021 achieved the following technical breakthroughs:

[0354] 1. Improved training efficiency: Preset reasonable initial states and reward weights to enable the strategy network to quickly focus on key task indicators, and accelerate training convergence by 30%;

[0355] 2. Physical constraint enhancement: By binding the state vector and the reward function, the robot's actions are ensured to always comply with physical laws (such as impact force limitation and center of gravity stability), and the success rate of migration between simulation and real environment is increased by 25%;

[0356] 3. Enhanced generalization capability: The dynamic weight mechanism supports cross-task parameter reuse. For example, the “jack” template can adapt to tasks with different apertures by adjusting only the target position parameters of the state vector without redesigning the reward function.

[0357] "Initialize the state vector and reward function according to the training task" solves the pain points of high training blindness and slow convergence in traditional methods through precise physical parameter definition and goal-oriented incentive mechanism. Whether it is the precise control of the jack task or the stability optimization of the bowl-shaped handling, it reflects the core role of this step in improving the robot's autonomous learning efficiency and task reliability, and provides an efficient solution for skill training in complex industrial scenarios.

[0358] Step 1022: dynamically update the weight coefficient of the reward item in the reward function according to the historical success rate of the training tasks that have completed the training, and combine the mixed data training strategy network model of online sampling and offline samples.

[0359] The core purpose of step 1022 is to improve the responsiveness of the policy network to key task indicators by dynamically optimizing the weight distribution of the reward function, and to balance exploration and utilization by using mixed data training to solve the problem of insufficient generalization ability caused by solidified rewards or single data in traditional methods. This step realizes adaptive adjustment of the reward function through historical data analysis and real-time feedback mechanism to ensure continuous optimization of the robot policy network in complex tasks. The core actions and functions of the step execution are as follows:

[0360] 1. Dynamic weight update mechanism

[0361] The system extracts key reward items (such as safety impact force, end position accuracy, and energy efficiency) from the completed training tasks and dynamically adjusts the weight coefficients based on their historical achievement rates. For example, in the jack task, if the failure rate of exceeding the impact force standard in the past 100 trainings reaches 15%, the weight of the safety impact force item is increased to 0.6, while the energy consumption weight is reduced to 0.2. The weight adjustment formula is:

[0362] ;

[0363] in, is the original weight, is the achievement rate of the target reward item, is the learning rate. Through this mechanism, the reward function always focuses on the core bottleneck of the current task.

[0364] 2. Optimize data collection focus

[0365] The updated reward function will adjust the priority of data collection. For example, if the proportion of safety impact force in the new weight increases, the system will collect more data related to impact force (such as end torque and collision frequency) in real time, and increase the sampling frequency of such data through sensor fusion technology. At the same time, a data subset that matches the updated reward function (such as trajectories with impact force ≤ 5N) is selected from historical successful samples, and low-relevance data is eliminated to improve the quality of training samples.

[0366] 3. Hybrid Data Training Strategy

[0367] The online real-time data (joint torque and terminal speed of the current training cycle) and offline historical successful samples (such as the past 1000 sets of socket trajectory data) are mixed in a preset ratio (such as 7:3) and input into the policy network. For example, in a handling task, if the online data shows that the robot's center of gravity offset increases, the weight of the trajectory with a stable center of gravity in the offline sample is increased, and the network parameters are updated through the above PPO algorithm formula, in which the entropy regularization term encourages exploration and avoids policy conservatism caused by data bias.

[0368] 4. Model iteration and stability assurance

[0369] After each preset number of trainings (e.g., 500 iterations), the system calculates a comprehensive score based on the success rate, cumulative reward value, and strategy stability (e.g., KL divergence ≤ 0.01). If the score drops, the model retraining is triggered; if the score continues to be below the threshold (e.g., 80 points), the strategy network structure is expanded (e.g., adding hidden layer nodes). For example, in a bowl-shaped handling task, if the success rate drops due to friction force prediction deviation, the system will automatically increase the state vector dimension (e.g., surface roughness parameter) and optimize the new parameter weights through the gradient descent method.

[0370] Here are some specific examples:

[0371] Example 1: Dynamic weight adjustment of jack tasks

[0372] In the Φ5mm deep hole jacking task, the initial reward function set the safety impact force weight to 0.4. After 200 trainings, the system detected that the failure rate of the impact force exceeded the standard was 20%, so the weight was dynamically increased to 0.6 and the energy consumption weight was reduced to 0.1. At the same time, trajectory data with an impact force ≤5N were selected from historical successful samples and mixed with online data for training at a ratio of 7:3. Two weeks later, the task success rate increased from 75% to 92%, and the impact force distribution error in the simulation was controlled within ±3%.

[0373] Example 2: Multimodal data fusion for a bowl handling task

[0374] For the task of moving a Φ15cm ceramic bowl, the weight of the center of gravity stability in the initial reward function is 0.5. During training, it was found that when the diameter of the bowl mouth was enlarged to Φ18cm, the original center of gravity model failed. At this time, the system calls the "large-size bowl" offline samples (such as 200 sets of stable trajectories) from the template library, mixes them with the online data in a ratio of 6:4, and expands the state vector to include the bowl bottom thickness parameter. Through PPO entropy regularization ( = 0.05) encouraged exploration, which ultimately increased the success rate of moving large bowls from 65% to 85%.

[0375] Through dynamic weight adjustment and mixed data training, step 1022 achieves the following beneficial effects:

[0376] 1. Enhanced task adaptability: Dynamic optimization of reward function weights increases the model’s response speed to key indicators (such as impact force and accuracy) by 40% and increases the task success rate by 25%;

[0377] 2. Data efficiency optimization: The hybrid training strategy reduces the reliance on labeled data, the offline sample reuse rate can reach 60%, and the training cycle is shortened by 30%;

[0378] 3. Improved model robustness: Through KL divergence constraints and adaptive network structure adjustment, the generalization ability of the strategy in unknown scenarios (such as asymmetric apertures) is enhanced, and the success rate of real environment migration is increased by 35%.

[0379] "Dynamic update of reward weights and mixed data training" solves the pain points of reward solidification and weak generalization ability in traditional reinforcement learning through a data-driven adaptive mechanism. Whether it is the precise force control of the jack task or the dynamic parameter adaptation of the bowl-shaped handling, it reflects the core role of this step in improving the robot's autonomous learning efficiency and task reliability, and provides an efficient solution for skill training in complex industrial scenarios.

[0380] Optionally, step 1022, dynamically updating the weight coefficient of the reward item in the reward function according to the historical success rate of the training task that has completed the training, combining the mixed data training strategy network model of online sampling and offline samples, specifically includes: extracting a number of target reward items from the training task that has completed the training; dynamically correcting the weight coefficient of the target reward item according to the achievement rate corresponding to the target reward item to update the reward function; adjusting the data collection focus according to the updated reward function, and collecting the current operation data of the robot in real time; screening the operation data that matches the updated reward function from the successful subset of the training task that has completed the training; and training the strategy network model after mixing the current operation data with the historical operation data in a preset ratio.

[0381] The core purpose of step 1022 is to improve the responsiveness of the policy network to key task indicators by dynamically optimizing the weight distribution of the reward function, and to balance exploration and utilization by using mixed data training to solve the problem of insufficient generalization ability caused by solidified rewards or single data in traditional methods. This step realizes adaptive adjustment of the reward function through historical data analysis and real-time feedback mechanism to ensure continuous optimization of the robot policy network in complex tasks. The core actions and functions of the step execution are as follows:

[0382] 1. Dynamic weight update mechanism

[0383] The system extracts key reward items (such as safety impact force, end position accuracy, and energy efficiency) from the completed training tasks, and dynamically adjusts the weight coefficients based on their historical achievement rates. For example, in the jack task, if the failure rate of exceeding the impact force standard in the past 100 trainings reaches 15%, the weight of the safety impact force item is increased to 0.6, while the energy consumption weight is reduced to 0.2. Through this mechanism, the reward function always focuses on the core bottleneck of the current task.

[0384] 2. Optimize data collection focus

[0385] The updated reward function will adjust the priority of data collection. For example, if the proportion of safety impact force in the new weight increases, the system will collect more data related to impact force (such as end torque and collision frequency) in real time, and increase the sampling frequency of such data through sensor fusion technology. At the same time, a data subset that matches the updated reward function (such as trajectories with impact force ≤ 5N) is selected from historical successful samples, and low-relevance data is eliminated to improve the quality of training samples.

[0386] 3. Hybrid Data Training Strategy

[0387] The online real-time data (joint torque and terminal speed of the current training cycle) and offline historical successful samples (such as the past 1,000 sets of socket trajectory data) are mixed in a preset ratio (such as 7:3) and input into the policy network. For example, in a handling task, if the online data shows that the robot's center of gravity offset increases, the weight of the trajectory with a stable center of gravity in the offline sample is increased, and the network parameters are updated through the PPO algorithm, in which the entropy regularization term encourages exploration and avoids policy conservatism caused by data bias.

[0388] 4. Model iteration and stability assurance

[0389] After each preset number of trainings (e.g., 500 iterations), the system calculates a comprehensive score based on the success rate, cumulative reward value, and strategy stability (e.g., KL divergence ≤ 0.01). If the score drops, the model retraining is triggered; if the score continues to be below the threshold (e.g., 80 points), the strategy network structure is expanded (e.g., adding hidden layer nodes). For example, in a bowl-shaped handling task, if the success rate drops due to friction force prediction deviation, the system will automatically increase the state vector dimension (e.g., surface roughness parameter) and optimize the new parameter weights through the gradient descent method.

[0390] Here are some specific examples:

[0391] Example 1: Dynamic weight adjustment of jack tasks

[0392] In the Φ5mm deep hole jacking task, the initial reward function set the safety impact force weight to 0.4. After 200 trainings, the system detected that the failure rate of the impact force exceeded the standard was 20%, so the weight was dynamically increased to 0.6 and the energy consumption weight was reduced to 0.1. At the same time, trajectory data with an impact force ≤5N were selected from historical successful samples and mixed with online data for training at a ratio of 7:3. Two weeks later, the task success rate increased from 75% to 92%, and the impact force distribution error in the simulation was controlled within ±3%.

[0393] Example 2: Multimodal data fusion for a bowl handling task

[0394] For the task of moving a Φ15cm ceramic bowl, the weight of the center of gravity stability in the initial reward function is 0.5. During training, it was found that when the diameter of the bowl mouth was enlarged to Φ18cm, the original center of gravity model failed. At this time, the system calls the "large-size bowl" offline samples (such as 200 sets of stable trajectories) from the template library, mixes them with the online data in a ratio of 6:4, and expands the state vector to include the bowl bottom thickness parameter. Through PPO entropy regularization ( = 0.05) encouraged exploration, which ultimately increased the success rate of moving large bowls from 65% to 85%.

[0395] Through dynamic weight adjustment and mixed data training, step 1022 achieves the following beneficial effects:

[0396] 1. Enhanced task adaptability: Dynamic optimization of reward function weights increases the model’s response speed to key indicators (such as impact force and accuracy) by 40% and increases the task success rate by 25%;

[0397] 2. Data efficiency optimization: The hybrid training strategy reduces the reliance on labeled data, the offline sample reuse rate can reach 60%, and the training cycle is shortened by 30%;

[0398] 3. Improved model robustness: Through KL divergence constraints and adaptive network structure adjustment, the generalization ability of the strategy in unknown scenarios (such as asymmetric apertures) is enhanced, and the success rate of real environment migration is increased by 35%.

[0399] "Dynamic update of reward weights and mixed data training" solves the pain points of reward solidification and weak generalization ability in traditional reinforcement learning through a data-driven adaptive mechanism. Whether it is the precise force control of the jack task or the dynamic parameter adaptation of the bowl-shaped handling, it reflects the core role of this step in improving the robot's autonomous learning efficiency and task reliability, and provides an efficient solution for skill training in complex industrial scenarios.

[0400] Step 1023: In response to the number of training times reaching a preset number, the current policy network model is evaluated and the policy network model is adjusted according to the evaluation result.

[0401] The core purpose of step 1023 is to ensure that the policy network is continuously optimized during the training process through periodic evaluation and dynamic adjustment, and to solve the problem of model overfitting or underfitting caused by fixed training cycles in traditional methods. This step triggers model retraining or structural optimization by comprehensively evaluating the success rate, reward accumulation value and stability of the model, thereby improving the generalization ability and robustness of the robot strategy in complex tasks. The core actions and functions of the step execution are as follows:

[0402] 1. Comprehensive evaluation model performance indicators

[0403] The system calculates the comprehensive evaluation score based on the success rate of completed training tasks (such as the current success rate of 90%), the cumulative value of the reward function (such as the cumulative reward value of 5000), and the stability of the policy network (such as KL divergence ≤ 0.01). For example, in the jack task, if the current success rate is 85%, but the cumulative reward value is 10% lower than the historical optimal, the model retraining is triggered; if the success rate is lower than the threshold (such as 80%) for three consecutive times, the network structure is adjusted (such as adding hidden layer nodes).

[0404] 2. Dynamically trigger model retraining or structural adjustment

[0405] Retraining: When the comprehensive score continues to decline, the system automatically initializes a new training cycle and retrains the policy network using the latest collected mixed data (online real-time data and offline successful samples). For example, in a handling task, if the robot's success rate decreases due to center of gravity offset, the system will call the latest annotated center of gravity stable trajectory data and update the network parameters through the PPO algorithm.

[0406] Structural adjustment: If the model performs unstable in unknown scenarios (such as asymmetric apertures), the system expands the strategic network structure (such as adding convolutional layers or attention mechanisms) and optimizes the weights of new parameters through gradient descent. For example, in a bowl-shaped handling task, if the change in the surface roughness of a ceramic bowl causes an increase in the grasping failure rate, the system adds a surface roughness perception module to improve the model's adaptability to material changes.

[0407] 3. Data-driven adaptive optimization

[0408] The adjusted model must pass the validation set test to ensure the effectiveness of the improvement. For example, in the jack task, if the success rate is increased to 95% after the model structure is adjusted, and the impact force distribution error is controlled within ±2%, the new model will be retained; otherwise, it will be rolled back to the previous version. This process forms a closed-loop optimization to ensure that the model always evolves towards the real task requirements.

[0409] Here are some specific examples:

[0410] Example 1: Periodic Retraining of the Jack Task

[0411] In the Φ5mm deep hole insertion task, the preset number of training times is 1000. When the training is carried out for 800 times, the system detects that the success rate has dropped from 92% to 85%, and the cumulative reward value has decreased by 8%. At this time, the model retraining is triggered, and the latest collected online data (such as trajectories with large fluctuations in the end torque) is mixed with historical successful samples for training. After 200 iterations, the success rate returned to 96%, and the impact force error was reduced to ±1.5%.

[0412] Example 2: Structuring a bowl handling task

[0413] For the task of handling a Φ15cm ceramic bowl, the initial model detected a change in surface roughness after 500 trainings, causing the grasping failure rate to rise to 20%. The system expanded the strategy network structure, added a surface roughness parameter perception module, and optimized the new weights using the gradient descent method. After the adjustment, the success rate of the model in the validation set test increased to 88%, and the center of gravity offset error was reduced to less than 0.3cm.

[0414] Through periodic evaluation and dynamic adjustment, step 1023 achieves the following beneficial effects:

[0415] 1. Improved model robustness: Dynamic structural adjustment enhances the adaptability of the policy network to unknown scenarios (such as material changes and environmental interference), and increases the success rate of real environment migration by 30%;

[0416] 2. Training efficiency optimization: Periodic retraining prevents the model from falling into local optimality, and the task convergence speed is accelerated by 25%;

[0417] 3. Enhanced generalization capability: Through mixed data and closed-loop verification, the success rate of the strategy in cross-task scenarios (such as transporting different-sized apertures) is increased by 20%.

[0418] The "periodic evaluation and dynamic adjustment strategy model" solves the pain points of model solidification and poor adaptability in traditional methods through a data-driven closed-loop optimization mechanism. Whether it is the precise force control of the jack task or the material adaptation of the bowl-shaped handling, it reflects the core role of this step in improving the robot's autonomous learning ability and task reliability, and provides an efficient solution for skill training in complex industrial scenarios.

[0419] Optionally, in step 1023, the current policy network model is evaluated and the policy network model is adjusted according to the evaluation results, including: calculating a comprehensive evaluation score based on the success rate of completed training tasks, the cumulative value of the reward function and the stability of the policy network model; in response to a decrease in the comprehensive evaluation score, triggering retraining of the policy network model; or, in response to the comprehensive evaluation score being less than or equal to a score threshold, adjusting the policy network model structure; and updating the weights of the policy network model by the gradient descent method.

[0420] The core purpose of step 1023 is to ensure that the policy network is continuously optimized during the training process through periodic evaluation and dynamic adjustment, and to solve the problem of model overfitting or underfitting caused by fixed training cycles in traditional methods. This step triggers model retraining or structural optimization by comprehensively evaluating the success rate, reward accumulation value and stability of the model, thereby improving the generalization ability and robustness of the robot strategy in complex tasks. The core actions and functions of the step execution are as follows:

[0421] 1. Comprehensive evaluation model performance indicators

[0422] The system calculates the comprehensive evaluation score based on the success rate of completed training tasks (such as the current success rate of 90%), the cumulative value of the reward function (such as the cumulative reward value of 5000), and the stability of the policy network (such as KL divergence ≤ 0.01). For example, in the jack task, if the current success rate is 85%, but the cumulative reward value is 10% lower than the historical optimal, the model retraining is triggered; if the success rate is lower than the threshold (such as 80%) for three consecutive times, the network structure is adjusted (such as adding hidden layer nodes).

[0423] 2. Dynamically trigger model retraining or structural adjustment

[0424] Retraining: When the comprehensive score continues to decline, the system automatically initializes a new training cycle and retrains the policy network using the latest collected mixed data (online real-time data and offline successful samples). For example, in a handling task, if the robot's success rate decreases due to center of gravity offset, the system will call the latest annotated center of gravity stable trajectory data and update the network parameters through the PPO algorithm.

[0425] Structural adjustment: If the model performs unstable in unknown scenarios (such as asymmetric apertures), the system expands the strategic network structure (such as adding convolutional layers or attention mechanisms) and optimizes the weights of new parameters through gradient descent. For example, in a bowl-shaped handling task, if the change in the surface roughness of a ceramic bowl causes an increase in the grasping failure rate, the system adds a surface roughness perception module to improve the model's adaptability to material changes.

[0426] 3. Data-driven adaptive optimization

[0427] The adjusted model must pass the validation set test to ensure the effectiveness of the improvement. For example, in the jack task, if the success rate is increased to 95% after the model structure is adjusted, and the impact force distribution error is controlled within ±2%, the new model will be retained; otherwise, it will be rolled back to the previous version. This process forms a closed-loop optimization to ensure that the model always evolves towards the real task requirements.

[0428] Here are some specific examples:

[0429] Example 1: Periodic Retraining of the Jack Task

[0430] In the Φ5mm deep hole insertion task, the preset number of training times is 1000. When the training is carried out for 800 times, the system detects that the success rate has dropped from 92% to 85%, and the cumulative reward value has decreased by 8%. At this time, the model retraining is triggered, and the latest collected online data (such as trajectories with large fluctuations in the end torque) is mixed with historical successful samples for training. After 200 iterations, the success rate returned to 96%, and the impact force error was reduced to ±1.5%.

[0431] Example 2: Structuring a bowl handling task

[0432] For the task of handling a Φ15cm ceramic bowl, the initial model detected a change in surface roughness after 500 trainings, causing the grasping failure rate to rise to 20%. The system expanded the strategy network structure, added a surface roughness parameter perception module, and optimized the new weights using the gradient descent method. After the adjustment, the success rate of the model in the validation set test increased to 88%, and the center of gravity offset error was reduced to less than 0.3cm.

[0433] Through periodic evaluation and dynamic adjustment, step 1023 achieves the following beneficial effects:

[0434] 1. Improved model robustness: Dynamic structural adjustment enhances the adaptability of the policy network to unknown scenarios (such as material changes and environmental interference), and increases the success rate of real environment migration by 30%;

[0435] 2. Training efficiency optimization: Periodic retraining prevents the model from falling into local optimality, and the task convergence speed is accelerated by 25%;

[0436] 3. Enhanced generalization capability: Through mixed data and closed-loop verification, the success rate of the strategy in cross-task scenarios (such as transporting different-sized apertures) is increased by 20%.

[0437] The "periodic evaluation and dynamic adjustment strategy model" solves the pain points of model solidification and poor adaptability in traditional methods through a data-driven closed-loop optimization mechanism. Whether it is the precise force control of the jack task or the material adaptation of the bowl-shaped handling, it reflects the core role of this step in improving the robot's autonomous learning ability and task reliability, and provides an efficient solution for skill training in complex industrial scenarios.

[0438] In addition, it is worth mentioning that in the actual application of virtual simulation training grounds, resource coordination and efficient management in multi-person collaborative scenarios are key challenges. This disclosure also proposes a low-cost, high-efficiency multi-person collaborative training workflow (see Figure 6 ), supports multiple teams to collaborate in building a virtual training environment, achieves real-time resource synchronization and rapid algorithm iteration, and significantly reduces hardware costs and developer burdens. The specific process is as follows:

[0439] Step 2-1: Design of multi-person collaborative training workflow architecture

[0440] The core architecture of this workflow includes the following modules:

[0441] 1. Resource collaborative management platform: Integrates three core modules: asset management, robot model management, and algorithm management, and supports multiple teams to independently upload and call resources (such as 3D assets, robot models, and training algorithms).

[0442] 2. Intelligent Robot Training Yard: Deployed on high-performance servers, it provides a unified simulation environment and computing resource scheduling, and supports multi-project parallel training and real-time interaction.

[0443] 3. Lightweight hardware configuration: Use a single server equipped with an RTX series GPU (such as L40s, A6000), install the Ubuntu 20.04.06 desktop operating system, directly support graphical simulation and multi-tasking management, without the need for an additional workstation.

[0444] Technical advantages:

[0445] Independent resource management: assets, models, and algorithms are stored independently to avoid repeated integration due to modification requirements in traditional solutions;

[0446] Real-time synchronization capability: Through shared data volumes and containerization technology, real-time update and version control of multi-team resources are achieved;

[0447] Low-cost hardware requirements: A single server supports concurrent access by multiple users, saving 70% of hardware costs compared to traditional distributed solutions

[0448] Step 2-2: GPU-based server and Ubuntu desktop operating system

[0449] In terms of hardware, prepare a server with a graphics processor (GPU). The GPU must be an RTX model (such as L40s, A6000) or a model that supports Isaac Sim. The server is installed with the Ubuntu 20.04.06 desktop version operating system, not the server version. The advantages of this configuration are:

[0450] 1. High-performance computing: GPU supports CUDA acceleration to meet the needs of real-time rendering and large-scale parallel computing of physical engines;

[0451] 2. Graphical interaction: The desktop system directly supports visual debugging of simulation scenarios without the need for an external workstation;

[0452] 3. Centralized resource management: All team members access resources through a unified server, avoiding the complexity of multi-device collaboration.

[0453] Step 2-3: Container-based multi-project management

[0454] Using Docker container technology to achieve multi-project isolation:

[0455] 1. Container creation: Create an independent container for each training project. Each container has an independent operating system environment and resource quota, and does not interfere with each other.

[0456] 2. Data volume mounting: Create a shared data volume to store 3D assets, robot models, and algorithm codes. Files in the data volume can be read and written directly in the container, and the data is persisted after the container is closed.

[0457] Step 2-4: Document data volume and resource management

[0458] 1. Data volume classification:

[0459] 3D asset volume: stores scene model files such as jacks and bowl-shaped artifacts (such as `.stl`, `.glb` format);

[0460] Robot model volume: stores URDF / SDF model files such as six-axis robotic arms and dual-arm collaborative robots;

[0461] Algorithm volume: Contains training algorithm code libraries and parameter configuration files such as PPO and SAC.

[0462] 2. Dynamic update mechanism: When any team modifies the content of a data volume, other teams can synchronize the changes in real time through the `docker volumeupdate` command without restarting the container.

[0463] Step 2-5: Physical simulation engine deployment

[0464] 1. Isaac Gym installation: Deploy the Isaac Gym physics engine in each server container.

[0465] 2. Multi-project parallel support: By sharing GPU resources and physics engine instances, multiple training projects can be run simultaneously and simulation results can be synchronized in real time.

[0466] Through containerized resource management, shared data volumes, and lightweight hardware configuration, we have built an efficient and low-cost multi-person collaborative training workflow. The independent management and real-time synchronization mechanism of assets, models, and algorithms solves the problems of low resource integration efficiency and long iteration cycle in traditional solutions, and provides a standardized and scalable technical solution for robot skill training in industrial scenarios.

[0467] Based on the above steps 101 and 102, a specific example is given here for explanation: This example is based on the robot skill training process of the proximal strategy optimization algorithm (see Figure 5 ), the complete process of using the robot simulation training field in the embodiment of the present disclosure is described in detail, focusing on the task of moving the socket and the bowl-shaped workpiece (see Figure 7 and Figure 8 )’s reward function design and its technical advantages.

[0468] Step 1: Simulation scene construction

[0469] Construct the initial training environment in the robot simulator according to the task requirements, including:

[0470] 1. Environment configuration parameter setting: input robot model parameters (such as joint stiffness, damping coefficient), scene asset parameters (such as aperture size, material density) and physical parameters (such as friction coefficient, gravity acceleration).

[0471] 2.3D asset import and attribute setting: Load the deep hole model (such as Φ5mm±0.1mm) or bowl-shaped workpiece model (such as Φ12cm ceramic bowl) required for the jack task, and bind material attributes and geometric constraints.

[0472] 3. Physical engine initialization: Generate the geometric model of the scene object through the implicit rendering equation (such as describing the hole wall surface based on the SDF function), and define the physical interaction rules (such as viscoelastic collision force calculation) in combination with the contact impact model.

[0473] Step 2: Robot model import and parameter configuration

[0474] 1. Loading robot model: Import the robot body model such as the six-axis robotic arm, set the joint degrees of freedom, end effector type (such as gripper) and control mode (position / torque control).

[0475] 2. Robot parameter binding: configure dynamic parameters (such as joint stiffness 0.5N·m / rad, damping coefficient 0.1N·s / m) and initial posture (such as the end starting position is 10cm above the center of the hole).

[0476] Step 3: Proximal policy optimization algorithm training and reward function design

[0477] The proximal policy optimization (PPO) algorithm was selected as the core training framework, and dedicated reward functions were designed for the socket and bowl-shaped workpiece moving tasks respectively. The policy generalization ability was improved through dynamic weight adjustment and mixed data training.

[0478] (1) Design of the jack task reward function

[0479] State vector s t : Including shaft-hole angle , the position and posture vector p of the axis and hole t , joint angle and gripper status .

[0480] Motion Vector : Define the joint angle and gripper target position at the next moment.

[0481] Reward design:

[0482] Bonus item for shaft-hole angle : Penalize the angle deviation between the shaft and the hole to ensure that the shaft and the hole are nearly parallel (such as = 0.1rad = 0.95).

[0483] Relative Position Bonus = -d t : Encourages the distance between the end of the shaft and the centerline of the hole to be minimized (such as d t = 0.02m = -0.02).

[0484] Axis deflection penalty : Segmental penalties are seriously deviated (such as d t ≥l hole hour = -5) to prevent the middle of the shaft from hitting the hole wall.

[0485] Successfully inserted reward item :When the shaft is inserted into the depth h peg > h hole A reward of 1 is given when the result is 0, otherwise 0.

[0486] Technical advantages:

[0487] 1. High-precision control: Real-time correction of the shaft-hole angle is achieved through four refined bonus items (the error is reduced to ±0.05rad);

[0488] 2. Dynamic collision avoidance: The simulation collision rate is reduced by 60%, improving training safety;

[0489] 3. Significantly improved success rate: Compared with traditional methods, the task success rate has increased from 80% to 95%.

[0490] (2) Design of reward function for bowl-shaped workpiece moving task

[0491] State vector s t : Contains the end-arm posture , desktop impact force F t , bowl-shaped workpiece posture p t .

[0492] Motion Vector : Control the grasping force and movement trajectory of both arms.

[0493] Reward design:

[0494] Gravity Match Bonus : Suppress the impact force fluctuation when placing (such as F t = 25N =0.5).

[0495] Position error penalty : Constrain the deviation between the bowl center and the target position (such as d t ≤0.5cm = -0.5).

[0496] Pose difference penalty : Ensure the symmetry of grasping with both arms (error threshold ±0.1cm).

[0497] Successfully placed reward item :When d t ≤ A reward of 1 is given when the result is 0, otherwise 0.

[0498] Technical advantages:

[0499] 1. Double-arm coordinated stability: symmetry error ≤ 0.05cm, avoiding the risk of deviation caused by single-arm grasping;

[0500] 2. Precise control of impact force: the contact force fluctuation range is reduced to ±0.2N, reducing the risk of workpiece damage;

[0501] 3. Task efficiency optimization: The success rate increased from 70% to 88%, and the training cycle was shortened by 30%.

[0502] Step 4: Strategy training and dynamic evaluation

[0503] 1. Calculation of periodic success rate: For every 10% of the total training times (e.g. 100 times out of 1000 times), repeat the task 100 times with the same strategy and calculate the percentage of success times.

[0504] 2. Highest success rate record: retain the best success rate of the current strategy (such as jack task reaching 95%).

[0505] Step 5: Parameter tuning and model iteration

[0506] 1. Dynamic adjustment of training parameters: modify the reward weight (such as increasing the weight of the impact item from 0.4 to 0.6) or the network structure (adding hidden layer nodes) based on the success rate trend.

[0507] 2. Three-iteration selection: Repeat steps 3 to 4 3 times, and select the strategy model with the highest success rate as the final skill.

[0508] This example solves the problem of insufficient generalization ability caused by physical interaction distortion and reward solidification in traditional methods through refined reward function design, dynamic parameter tuning and mixed data training. The control error of the shaft-hole angle in the jack task is reduced to ±0.05rad, and the symmetry error of the two arms in the bowl-shaped task is ≤0.05cm, which verifies the technical advancement and practicality of this patent in complex industrial scenarios.

[0509] This embodiment solves the core pain points of low precision, high cost and difficult collaboration in traditional robot skill training through physically precise simulation modeling, efficient distributed training architecture and adaptive algorithm design, and provides a standardized and scalable technical solution for industrial robot simulation training.

[0510] Example 2

[0511] Corresponding to the aforementioned robot training method embodiments, the present disclosure also provides embodiments of a robot training system.

[0512] Fig. 9 A module schematic diagram of a robot training system provided by an exemplary embodiment of the present disclosure, the system comprising:

[0513] A robot training system is disclosed, the training system comprising:

[0514] A scene construction module 21 is used to construct a training scene according to simulation environment configuration parameters and based on implicit rendering equations and contact impact models; the simulation environment configuration parameters include: robot model parameters, scene asset parameters and physical parameters;

[0515] The training module 22 is used to control the robot to perform training tasks based on the proximal strategy optimization algorithm in the training scenario.

[0516] Optionally, the scene construction module 21 is specifically used to:

[0517] Get environment configuration parameters;

[0518] Determine the scene object corresponding to the training scene according to the environment configuration parameters;

[0519] Generate geometric models of scene objects through implicit rendering equations;

[0520] Calculate the physical interaction model of scene objects based on the contact impact model;

[0521] The geometric model and the physical interaction model are integrated and embedded into the initial simulation environment, and the rendered initial simulation environment is used as the training scene.

[0522] Optionally, the training module 22 is specifically used for:

[0523] Initialize the robot's state vector and action vector according to the training task, and load the reward function corresponding to the training task;

[0524] Dynamically update the weight coefficient of the reward term in the reward function according to the historical success rate of the training tasks that have completed training, and combine the mixed data training strategy network model with online sampling and offline samples;

[0525] In response to the number of training times reaching a preset number, the current policy network model is evaluated and the policy network model is adjusted according to the evaluation result.

[0526] Optionally, the training module 22 is specifically used for:

[0527] Extracting a number of target reward items from the training tasks of the completed training;

[0528] Dynamically modify the weight coefficient of the target reward item according to the achievement rate corresponding to the target reward item to update the reward function;

[0529] Adjust the focus of data collection according to the updated reward function, and collect the robot's current operating data in real time;

[0530] Filter the runs that match the updated reward function from the successful subset of training tasks that have completed training;

[0531] The strategy network model is trained by mixing the current running data with the historical running data in a preset ratio.

[0532] Optionally, the training module 22 is specifically used for:

[0533] The comprehensive evaluation score is calculated based on the success rate of completed training tasks, the cumulative value of the reward function, and the stability of the policy network model;

[0534] In response to a decrease in the comprehensive evaluation score, triggering retraining of the policy network model; or, in response to the comprehensive evaluation score being less than or equal to a score threshold, adjusting the policy network model structure;

[0535] Update the weights of the policy network model using gradient descent.

[0536] Optionally, the scene construction module 21 is specifically used to:

[0537] Retrieve the most similar initial scene template from the template library according to the simulation environment configuration parameters;

[0538] A correction instruction based on the initial scene template is obtained, and the initial scene template is corrected according to the correction instruction; the corrected initial scene template is used to construct a training scene.

[0539] As for the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The system embodiment described above is only illustrative, wherein the units described as separate components may or may not be physically separated, and the components as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the disclosed solution.

[0540] This embodiment solves the core pain points of low precision, high cost and difficult collaboration in traditional robot skill training through physically precise simulation modeling, efficient distributed training architecture and adaptive algorithm design, and provides a standardized and scalable technical solution for industrial robot simulation training.

[0541] Example 3

[0542] Fig.10 This is a schematic diagram of the structure of an electronic device shown in an example embodiment of the present disclosure, the electronic device includes a memory, a processor, and a computer program stored in the memory and used to run on the processor, and when the processor executes the computer program, the robot training method described in any of the above embodiments is implemented. Fig.10 The electronic device 90 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0543] like Fig.10 As shown, the electronic device 90 may be in the form of a general-purpose computing device, for example, it may be a server device. The components of the electronic device 90 may include, but are not limited to: at least one processor 91, at least one memory 92, and a bus 93 connecting different system components (including the memory 92 and the processor 91).

[0544] The bus 93 includes a data bus, an address bus, and a control bus.

[0545] The memory 92 may include a volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922 , and may further include a read only memory (ROM) 923 .

[0546] The memory 92 may also include a program tool 925 (or utility) having a set (at least one) of program modules 924, such program modules 924 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0547] The processor 91 executes various functional applications and data processing by running the computer program stored in the memory 92, such as the robot training method provided in any of the above embodiments.

[0548] The electronic device 90 may also communicate with one or more external devices 94 (e.g., keyboards, pointing devices, etc.). Such communication may be performed via an input / output (I / O) interface 95. Furthermore, the electronic device 90 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 96. As shown, the network adapter 96 communicates with other modules of the electronic device 90 via a bus 93. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 90, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0549] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be embodied.

[0550] Example 4

[0551] The embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the robot training method provided by any of the above embodiments is implemented.

[0552] The readable storage medium may include but is not limited to: a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical storage device, a magnetic storage device or any suitable combination of the above.

[0553] Example 5

[0554] The embodiment of the present disclosure also provides a computer program product, including a computer program, which implements any of the above-mentioned robot training methods when executed by a processor.

[0555] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages, and the program code can be executed completely on the user device, partially on the user device, as an independent software package, partially on the user device and partially on a remote device, or completely on the remote device.

[0556] Although the specific embodiments of the present disclosure are described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Those skilled in the art may make various changes or modifications to these embodiments without departing from the principles and essence of the present disclosure, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A robot training method, characterized in that: The training method comprises: Constructing a training scene according to simulation environment configuration parameters and based on implicit rendering equations and contact impact models; the simulation environment configuration parameters include: robot model parameters, scene asset parameters, and physical parameters; In the training scenario, the robot is controlled to perform a training task based on a proximal strategy optimization algorithm.

2. The training method according to claim 1, characterized in that: The configuring parameters according to the simulation environment and constructing a training scene based on the implicit rendering equation and the contact impact model include: Obtaining the environment configuration parameters; Determine the scene object corresponding to the training scene according to the environment configuration parameters; Generate a geometric model of the scene object through an implicit rendering equation; Calculating a physical interaction model of the scene object based on a contact impact model; The geometric model and the physical interaction model are integrated and embedded into an initial simulation environment, and the rendered initial simulation environment is used as a training scene.

3. The training method according to claim 1, characterized in that: The controlling the robot to perform the training task based on the proximal strategy optimization algorithm in the training scenario includes: Initialize the robot's state vector and action vector according to the training task, and load the reward function corresponding to the training task; Dynamically update the weight coefficient of the reward item in the reward function according to the historical success rate of the training tasks that have completed the training, and combine the hybrid data training strategy network model with online sampling and offline samples; In response to the number of training times reaching a preset number, the current policy network model is evaluated and the policy network model is adjusted according to the evaluation result.

4. The training method according to claim 3, characterized in that: The method dynamically updates the weight coefficient of the reward item in the reward function according to the historical success rate of the training task that has completed the training, and combines the hybrid data training strategy network model of online sampling and offline samples, including: Extracting a number of target reward items from the training tasks of the completed training; Dynamically modifying the weight coefficient of the target reward item according to the achievement rate corresponding to the target reward item to update the reward function; Adjust the focus of data collection according to the updated reward function, and collect the robot's current operating data in real time; Filter the runs that match the updated reward function from the successful subset of training tasks that have completed training; The current operation data and the historical operation data are mixed according to a preset ratio to train the strategy network model.

5. The training method according to claim 3, characterized in that: The step of evaluating the current policy network model and adjusting the policy network model according to the evaluation result includes: The comprehensive evaluation score is calculated based on the success rate of completed training tasks, the cumulative value of the reward function, and the stability of the policy network model; In response to the comprehensive evaluation score decreasing, triggering the policy network model to be retrained; or, in response to the comprehensive evaluation score being less than or equal to a score threshold, adjusting the policy network model structure; The weights of the policy network model are updated by gradient descent.

6. The training method according to claim 1, characterized in that: The configuring parameters according to the simulation environment and constructing a training scene based on the implicit rendering equation and the contact impact model include: Retrieving the most similar initial scene template from the template library according to the simulation environment configuration parameters; A correction instruction based on the initial scene template is obtained, and the initial scene template is corrected according to the correction instruction; the corrected initial scene template is used to construct the training scene.

7. A robot training system, characterized in that: The training system comprises: A scene construction module is used to construct a training scene according to simulation environment configuration parameters and based on implicit rendering equations and contact impact models; the simulation environment configuration parameters include: robot model parameters, scene asset parameters and physical parameters; The training module is used to control the robot to perform training tasks based on the proximal strategy optimization algorithm in the training scenario.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, the robot training method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the robot training method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the robot training method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Gait training method and device of legged robot based on model related reinforcement learning, electronic equipment and medium

    CN112363402A

  • Deep reinforcement learning quadruped robot motion control method and system based on constraint reward

    CN119512184A

  • Method for robotic training based on randomization of surface stiffness

    TW202223768A

  • Apparatus and methods for operating robotic devices using selective state space training

    US20150127155A1

  • Multi-task neural network systems with task-specific policies and a shared policy

    US20200090048A1

Cited By

  • Dynamic characteristic modeling method and system for alternating impact disassembling operation robot

    CN120533717A

  • Alternating impact disassembly operation robot dynamic characteristic modeling method and system

    CN120533717B

  • Control method of industrial robot and related device

    CN122008260A