Digital twin mechanical arm gangue poking training platform based on multi-sensor fusion and construction method thereof
By constructing a multi-sensor fusion digital twin robotic arm rock-moving training platform, the problems of high cost, safety risks, and the gap between simulation and reality in complex tasks of robotic arms have been solved, realizing efficient and safe intelligent control strategy training and physical control synchronization.
Patent Information
- Application Number
- CN202511745387.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-01-16
AI Technical Summary
Traditional robotic arms face challenges such as high cost, significant safety risks, difficulty in reproducing working conditions, and a gap between simulation and reality when dealing with unstructured or complex dynamic tasks. Furthermore, existing digital twin technology combined with reinforcement learning has failed to completely resolve these pain points.
A training platform for a digital twin robotic arm to remove coal gangue based on multi-sensor fusion was constructed. By fusing multi-sensor perception data, a high-precision digital twin model was established to interact in a closed loop with the real physical world, and reinforcement learning training was carried out in a high-fidelity perception-type digital twin environment.
It reduces training costs and security risks, improves training efficiency and policy robustness, enhances the perception accuracy and policy generalization ability of the real environment, and realizes the synchronous training of efficient intelligent control policies and physical control in the virtual environment.
Smart Images

Figure CN121340280A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of digital twins, robotics, and reinforcement learning, and in particular to a training platform for a digital twin robotic arm to remove coal gangue based on multi-sensor fusion and its construction method. Background Technology
[0002] In the mining process, sorting coal gangue is an essential step. With the development of industry and artificial intelligence, robotic arms are playing an increasingly important role in production lines, material sorting, and hazardous operations. However, when faced with unstructured or complex dynamic tasks, traditional control methods based on teach-and-write programming or motion planning often appear rigid, lacking sufficient adaptability and robustness, and also suffer from problems such as high scenario costs, difficulty in reproducing working conditions, and significant safety risks.
[0003] In recent years, reinforcement learning has provided a new paradigm for solving such complex control problems. By allowing robotic arms to optimize their control strategies through trial and error in their environment, reinforcement learning demonstrates strong adaptive capabilities. However, directly training robotic arms for reinforcement learning in real physical environments presents several challenges: 1) High training costs: Long-term, extensive trial-and-error training of real robotic arms incurs significant costs such as equipment wear and energy consumption; 2) High safety risks: Exploratory learning processes may lead to unexpected actions by the robotic arm, posing potential safety threats to the equipment itself, the surrounding environment, and even operators; 3) Difficulty in reproducing working conditions: Complex working conditions in real environments are difficult to reproduce accurately and on a large scale, limiting the diversity of training scenarios and the generalization ability of strategies; 4) The gap between simulation and reality: Due to parameter differences and modeling errors between simulation models and the real physical world, as well as the lag and distortion of real-world perception information in the simulation environment, strategies trained in simulation are often difficult to directly, efficiently, and stably transfer and deploy to real robotic arms.
[0004] In existing technologies, some attempts have been made to combine digital twin technology to optimize robot control, but these mainly focus on state monitoring, fault diagnosis, or offline planning. Few systematically utilize digital twins as a core platform to drive reinforcement learning training and continuous policy optimization. Traditional digital twin models may lack the accuracy to support the high-fidelity physical interaction required for reinforcement learning, and they lack effective closed-loop and feedback mechanisms for virtual and real data. This means that the integration of digital twins and reinforcement learning remains at a relatively superficial stage, failing to thoroughly address the aforementioned core pain points. Summary of the Invention
[0005] To address the problems of high cost, significant safety risks, difficulty in reproducing working conditions, and the gap between simulation and reality in existing reinforcement learning training for robotic arms, this invention provides a digital twin robotic arm training platform for rock removal based on multi-sensor fusion and its construction method. By deeply fusing multi-sensor perception data, a high-precision and highly robust digital twin model is established to interact in a closed loop with the real physical world, and reinforcement learning is used for efficient training in this high-fidelity perception-based digital twin environment.
[0006] This invention adopts the following technical solution: a method for constructing a training platform for a digital twin robotic arm to remove coal gangue based on multi-sensor fusion, comprising the following steps:
[0007] Step 1: After assembling the gripper file, add it to the URDF file of the robotic arm, and then use the URDFImporter tool to import the models of the robotic arm and gripper into the Unity engine.
[0008] Step 2: Construct a digital twin modeling module, discretize the conveyor belt, model and render it using 3ds Max, export it as an fbx file and import it into the Unity environment to build a digital environment that matches the actual environment.
[0009] Step 3: Construct a multi-sensor fusion module to collect visual data, robotic arm status data, and LiDAR data from real physical scenes. Through feature extraction and recognition, and multi-source data fusion based on deep learning, a unified environmental status is generated.
[0010] Step 4: Construct a reinforcement learning training module, use ML-Agent for reinforcement learning training, establish the state space, action space and reward function of the reinforcement learning algorithm, train the robotic arm's rock-moving strategy, and obtain the trained model.
[0011] Step 5: Set up a communication environment with ROS in the Unity environment;
[0012] Step 6: Construct a virtual-real interaction control module, export the trained model as ONNX format and deploy it to a computer device. Communicate with the controller of the real robotic arm through ROS, and listen to the joint_states topic of the real robotic arm to synchronize the joint states to the Unity environment and realize data visualization.
[0013] Preferably, step 1 includes the following sub-steps:
[0014] Step 1.1: Open the STP file of the gripper base and gripper fingers using SolidWorks, align the gripper fingers and assemble them on the gripper base to make the gripper fingers and gripper base a single unit;
[0015] Step 1.2: Install the sw_urdf_exporter plugin, export the gripper file as a URDF file, add it to the URDF file of the robotic arm, adjust the coordinate information to align it with the end effector of the robotic arm, and then assemble it.
[0016] Step 1.3: Install the URDF Importer plugin in Unity's Package Manager;
[0017] Step 1.4: Import the robotic arm's URDF file and resource files into Unity's Asset folder;
[0018] Step 1.5: In the Unity menu, select Import URDF under the URDF Importer path to load the complete robotic arm model into the Unity scene.
[0019] Preferably, step 2 includes the following sub-steps:
[0020] Step 2.1: Discretize the conveyor belt, dividing it into sections. There are three equal-length segments on a conveyor belt, with a speed of v and a time step of 0 for each segment. ;
[0021] Step 2.2, assuming the conveyor belt is... , , The coal quantities of each section are respectively , and Starting from section 1, the coal for section 2 flows in from section 1. The coal in the section is from the first Section inflow;
[0022] Step 2.3, after At time intervals, each segment of the coal flow on the conveyor belt transports its coal volume to the next segment while simultaneously receiving an equal amount of coal flowing in from the previous segment;
[0023] Step 2.4: Create a 3D model of the belt conveyor: Draw the main shape of the belt conveyor in 3ds Max, adjust the size, proportion and shape to match the actual belt conveyor, use editing tools to modify and refine the model to show the details of the belt conveyor, and add details and decorative elements according to the appearance and characteristics of the belt conveyor.
[0024] Step 2.5: Create a 3D model of the conveyor belt: Build a support frame, use cuboids to initialize and form the basic structure of the support frame, copy several identical structures to form the conveyor belt, and use the point welding command to optimize it; set up lighting, use the 3ds Max lighting tools to add light sources and adjust the brightness, color and shadow properties; set up rendering, adjust the resolution, shadows and reflections.
[0025] Step 2.6: In 3ds Max, select the created 3D model of the belt conveyor, export it as an FBX file, and place it in the Assets folder of the Unity project; in Unity, select to import the created 3D model of the conveyor belt, and create or assign materials to it in the Inspector panel;
[0026] Step 2.7: Create a gangue model in Unity, adjust its size, mass, and friction coefficient, and randomize the gangue shapes and initial distribution.
[0027] Step 2.8: Create a laboratory training environment and an area for placing gangue, adjust the position and physical properties of the conveyor belt, robotic arm, and gangue model, test whether they can contact and collide normally, and prepare for training.
[0028] Preferably, step 3 includes the following sub-steps:
[0029] Step 3.1: Select an RGB-D camera and fix it above the working area of the robotic arm to cover the area where the gangue is removed; connect it to the data interface of the robotic arm controller to read the joint angle and angular velocity data in real time; install a lidar to obtain the position of the gangue point cloud;
[0030] Step 3.2: Configure the sensor driver and perform time synchronization, perform distortion calibration on the RGB-D camera, and perform internal parameter calibration on the LiDAR;
[0031] Step 3.3: Collect raw data from the RGB-D camera, robotic arm controller, and LiDAR interface via an independent ROS node, publish it in a standardized message format, and attach a timestamp;
[0032] Step 3.4: Deploy the trained YOLOv8 model as an independent ROS node, subscribe to the preprocessed visual image stream, perform real-time inference, and output the 2D bounding box, confidence score, and category of each identified gangue; acquire the original state data of the robotic arm, convert the original encoder readings into joint angles and angular velocities; subscribe to the preprocessed LiDAR point cloud data, remove ground points, and use the K-Means algorithm to segment the remaining non-ground point cloud into multiple independent clusters, each cluster representing a potential gangue target; perform geometric analysis on the clustered point cloud clusters to calculate the centroid coordinates and principal direction;
[0033] Step 3.5: Construct a multi-layer fully connected deep neural network and train it. Deploy the trained network model as an independent ROS node, subscribe to feature inputs and perform real-time inference, and publish a unified environmental state data packet.
[0034] Step 3.6: The multi-sensor fusion entity layer transmits the data to the virtual-real interaction control module to update the virtual environment status.
[0035] Preferably, step 4 includes the following sub-steps:
[0036] Step 4.1: Install and configure the ML-Agents plugin in Unity, configure the Agent component for the virtual robotic arm model in the digital twin, and specify its behavior;
[0037] Step 4.2: Define the observation space and perform normalization processing; the observation space is represented by 16 continuous values, including: 6-dimensional joint angles, 4-dimensional information from the gripper to the target, 4-dimensional information from the target to the target area, and 2-dimensional task status;
[0038] Step 4.3: Define the motion space and control it through an agent; the motion space is 6-dimensional, corresponding to the 6 joints of the robotic arm;
[0039] Step 4.4: Define the reward function and use dense rewards to ensure that the Agent receives clear feedback signals at each step, and follow the hierarchical guidance principle to enable the Agent to learn complex tasks step by step.
[0040] Step 4.5: Randomize key parameters in the digital twin environment, randomly changing the friction coefficient, mass, and elasticity of the gangue; randomize the light source, texture, etc.
[0041] Step 4.6: Train ML-Agents by using reinforcement learning algorithms to train the control strategy of the robotic arm and optimize the performance of the Agent in different scenarios.
[0042] Step 4.7: Observe and debug the learning effect of the Agent during the training process, and gradually improve the control of the robotic arm by adjusting hyperparameters, reward functions or adding new constraints;
[0043] Step 4.8: Deploy the trained virtual robotic arm model in Unity for optimization, export the optimized model in ONNX format, and deploy the application on the ROS platform.
[0044] Preferably, step 5 includes the following sub-steps:
[0045] Step 5.1: Import the two components, ros-tcp-connector and unity robotics visualizations, into the Unity platform via the Package Manager under Windows;
[0046] Step 5.2: Configure ROS IP and ROS Port;
[0047] Step 5.3: Compile ROS-TCP-Endpoint in the ROS workspace.
[0048] Preferably, step 6 includes the following sub-steps:
[0049] Step 6.1: Configure the ROS communication plugin in Unity, subscribe to the joint_states topic, and obtain and parse real-time joint state data through the robotic arm control interface;
[0050] Step 6.2: Map the parsed joint angle values to the digital robotic arm model in Unity, convert the joint state data into angle information for each joint in Unity, update the joint state and end pose of the virtual robotic arm, and perform actual virtual control.
[0051] Step 6.3: When the digital twin environment receives a state update from the real world, it provides the received state as input to the Agent's neural network. The Agent's neural network outputs a 6-dimensional continuous action value based on the policy it has learned.
[0052] Step 6.4: Integrate the trajectory planner and use MoveIt to plan the path based on the received joint angles and target pose. Calculate the trajectory that will allow the robotic arm to move from the current position to the target pose, including the joint angles and pose changes at each time step.
[0053] Step 6.5: Send the calculated trajectory back to Unity as the response data of the ROS service, and control the digital robotic arm model in Unity to perform the corresponding movement according to the planned trajectory.
[0054] Step 6.6: Apply the same trajectory to the physical robotic arm to execute the same movements as the Unity digital robotic arm model, thereby achieving synchronized virtual control of the physical object's actions and completing the virtual-real interaction.
[0055] The present invention also provides a digital twin robotic arm training platform for removing coal gangue based on multi-sensor fusion, constructed using the above method, comprising:
[0056] Digital Twin Modeling Module: Constructs a digital twin model of the physical robotic arm and the working environment to reflect the response characteristics of the real robotic arm and its interaction with the environment, ensuring consistency between the digital twin model and the real physical entity;
[0057] Multi-sensor fusion module: used to collect visual data, robotic arm status data and LiDAR data of real physical scene in real time, extract and identify features of the collected data, and perform multi-source data fusion based on deep learning to output a unified environmental status.
[0058] Reinforcement learning training module: The digital twin environment is used as the trainer for reinforcement learning, and the robotic arm is used as the agent to learn through trial and error in the environment;
[0059] The virtual-real interaction control module is used for real-time mapping of the real physical environment to the digital twin environment, and at the same time, it converts the control strategies trained in the digital twin environment into control commands for the real robotic arm and issues them.
[0060] Compared with the prior art, the present invention, employing the above technical solution, has the following technical effects:
[0061] 1. Reduced training costs and improved security: This invention utilizes a high-fidelity digital twin environment as a reinforcement learning trainer. The agent conducts large-scale trial-and-error learning in a virtual environment, completely avoiding dependence on real devices, damage risks, and security issues. This greatly reduces training costs, shortens the training cycle, and improves the security of the training process.
[0062] 2. Improve training efficiency and strategy robustness: The platform of this invention provides clear and gradient feedback signals to the Agent at each step of the complex rock-removing task through a finely designed dense reward function and hierarchical guidance strategy. This effectively solves the problem of low exploration efficiency caused by sparse rewards in complex tasks, accelerates the learning convergence speed of the Agent, and enables it to form efficient intelligent control strategies more quickly.
[0063] 3. Improve the accuracy and completeness of perception of the real physical environment: This invention innovatively introduces a multi-sensor fusion entity layer, which uses deep learning technology to efficiently and intelligently fuse data from multiple heterogeneous sensors such as vision (YOLOv8), robotic arm status (encoder), and lidar (K-Means point cloud clustering). This overcomes the limitations of single sensors being susceptible to factors such as ambient light, occlusion, and noise, and provides a more accurate, complete, and reliable real-time unified environmental status than a single sensor. This improved perception capability is the foundation for supporting subsequent intelligent decision-making and control.
[0064] 4. Enhanced generalization ability of the strategy: The platform of this invention implements a deep and diverse domain randomization strategy, which allows the Agent to be exposed to almost an infinite number of working condition variations during the training phase. This greatly enhances the robustness and generalization ability of the learned control strategy, enabling it to perform tasks stably and efficiently when faced with various unknowns and changes in the real world.
[0065] 5. Easy to expand and migrate: The acquisition, inference, message bus, twin rendering and control distribution modules in the method of this invention are decoupled, and the model / protocol can be changed as needed and quickly reused to multiple conveyor lines or different mining areas, reducing the cost of modification and maintenance.
[0066] 6. The digital twin robotic arm training platform for removing coal gangue of this invention has broad industrial application prospects, can effectively improve the level of automation in coal mine production, and plays a positive role in promoting the development of the coal mining industry. Attached Figure Description
[0067] Figure 1 This is a flowchart of the construction method of the digital twin robotic arm rock-removing training platform of the present invention;
[0068] Figure 2 This is a model diagram of the robotic arm used in the embodiments of the present invention;
[0069] Figure 3 This is a model diagram of the gripper used in the embodiments of the present invention;
[0070] Figure 4 These are images of the digital twin environment updated in real time in embodiments of the present invention;
[0071] Figure 5 This is a flowchart of reinforcement learning training in an embodiment of the present invention;
[0072] Figure 6 This is the reward curve for reinforcement learning training in this embodiment of the invention. Detailed Implementation
[0073] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of the application will be further described in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in this invention. All non-innovative embodiments based on these embodiments by other researchers in the art are within the protection scope of this invention. Furthermore, the step numbers in the embodiments of this invention are only set for ease of explanation and do not limit the order of the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0074] Example 1
[0075] This embodiment provides a training platform for a digital twin robotic arm to remove rock based on multi-sensor fusion, including: a digital twin modeling module, a multi-sensor fusion module, a reinforcement learning training module, and a virtual-real interaction control module.
[0076] The digital twin modeling module is used to accurately construct a digital twin model of the physical robotic arm and the working environment. This model can accurately reflect the response characteristics of the real robotic arm and its interaction with the environment, ensuring a high degree of consistency between the digital twin model and the real physical entity.
[0077] The multi-sensor fusion module is used to collect visual data, robotic arm status data, and LiDAR data of real physical scenes in real time, extract and identify features from the collected data, and perform multi-source data fusion based on deep learning to output a unified environmental status.
[0078] The reinforcement learning training module uses a high-fidelity digital twin environment as the reinforcement learning trainer, enabling the robotic arm to perform large-scale trial-and-error learning in the environment as an agent, effectively solving the cost and safety issues of training in real environments.
[0079] The virtual-real interaction control module is used to realize the real-time mapping of the real physical environment to the digital twin environment, and at the same time convert the control strategy trained in the digital twin environment into control instructions that can be executed by the real robotic arm and issue them.
[0080] Example 2
[0081] This embodiment provides a method for constructing the digital twin robotic arm rock-removing training platform based on multi-sensor fusion as described in the above embodiments, such as... Figure 1 As shown, the details are as follows:
[0082] S1. After assembling the gripper file, add it to the URDF file of the robotic arm, and then import the robotic arm and gripper models into the Unity engine using the URDFImporter tool.
[0083] S2. Construct a digital twin modeling module, discretize the conveyor belt, model and render it using 3ds Max, export it as an FBX file, import it into the Unity environment, and build a digital environment that matches the actual environment.
[0084] S3. Construct a multi-sensor fusion module to collect visual data, robotic arm status data, and LiDAR data from real physical scenes. Through feature extraction and recognition, and multi-source data fusion based on deep learning, generate a unified environmental state.
[0085] S4. Construct a reinforcement learning training module, that is, use ML-Agent for reinforcement learning training, design the state space, action space and reward function of the reinforcement learning algorithm, and train the robotic arm's rock-moving strategy.
[0086] S5. Set up a communication environment with ROS in Unity.
[0087] S6. Construct a virtual-real interaction control module. Export the trained model as ONNX format, deploy it to a computer device, and communicate with the controller of the real robotic arm through ROS. At the same time, it can listen to the joint_states topic of the real robotic arm and synchronize the joint states to Unity to achieve data visualization.
[0088] As a preferred embodiment, step 1 specifically includes the following sub-steps:
[0089] S11, such as Figure 3 As shown, open the STP files of the gripper base and gripper fingers using SolidWorks, align the gripper fingers and assemble them onto the gripper base to form a single unit.
[0090] S12. Install the sw_urdf_exporter plugin to export the gripper file as a URDF file.
[0091] S13. Add the gripper's URDF file to the robotic arm's URDF file, such as... Figure 2 The image shows the model file of the robotic arm. Adjust the coordinate information of the gripper to align it with the end effector of the robotic arm and assemble it.
[0092] S14. Install the URDF Importer plugin in Unity's Package Manager.
[0093] S15. Import the robotic arm's URDF file and related resource files into Unity's Asset folder.
[0094] S16. In the Unity menu, select URDF Importer > Import URDF to load the complete robotic arm model into the Unity scene.
[0095] As a preferred embodiment, step 2 specifically includes the following sub-steps:
[0096] S21. Assume the coal flow rate on the belt conveyor is... Its size depends on changes in time and location.
[0097] Then the conveyor belt is divided into Discretize the conveyor belt using segments of equal length, with the conveyor belt running at speed v and the time step for each segment being . .
[0098] S22, Assuming the conveyor belt's first... , , The coal quantities of each section are respectively , and Starting from section 1, the coal for section 2 flows in from section 1. The coal in the section is from the first The flow of water into the section.
[0099] S23, Assume that at time... , No. The coal flow intensity of the section is So after After the time interval, the first The amount of coal in a section can be expressed in the following way:
[0100] ;
[0101] So in At any given time, the coal load for each segment can be expressed as follows: , and And there is a relationship
[0102] .
[0103] go through At time intervals, each segment of the coal flow on the conveyor belt will transport its coal volume to the next segment, while the current segment will receive coal of equal size flowing in from the previous segment.
[0104] S24. In 3ds Max, use the selected tools to draw the main shape of the belt conveyor, and adjust the size, scale, and shape of the geometry by adjusting parameters or dragging control points to ensure it matches the actual belt conveyor.
[0105] Use editing tools to further modify and refine the model, including moving, rotating, scaling, and deforming different parts, to accurately represent the details of the belt conveyor.
[0106] Add details and decorative elements, such as side panels and support frames, based on the appearance and characteristics of the belt conveyor.
[0107] S25. The design of the support frame first uses a cuboid for initialization to form a basic support structure. Then, the model is optimized by copying 36 identical structures and using the point welding command.
[0108] Next, set up the lighting for the belt conveyor model. Use the lighting tools in 3ds Max to add light sources in appropriate locations and adjust the brightness, color, and shadow properties.
[0109] Finally, configure the rendering settings by selecting a renderer, adjusting the resolution, shadows, and reflections.
[0110] S26. In 3ds Max, select the 3D model of the belt conveyor that you have created, export it as an FBX file, drag and drop the exported model file into the Assets folder of your Unity project, select the imported conveyor belt model in Unity, and then create or assign materials to it in the Inspector panel.
[0111] S27. Create a gangue model in Unity, adjust its physical properties such as size, mass, and friction coefficient, and use randomization techniques to generate diverse gangue shapes and initial distributions.
[0112] S28. Create the training environment and the area for placing gangue in the laboratory. Adjust the positions of the conveyor belt, robotic arm, and gangue model, adjust the physical properties, and test whether they can contact and collide normally in preparation for training.
[0113] As a preferred embodiment, step 3 specifically includes the following sub-steps:
[0114] S31. Select an RGB-D camera and fix it above the working area of the robotic arm to ensure that it can cover the entire area of the rock removal operation.
[0115] Connect to the data interface provided by the robotic arm controller to ensure that data such as joint angle and angular velocity can be read in real time.
[0116] Select a lidar that meets the scanning range of the work area and install it in a location that can effectively acquire the point cloud of gangue.
[0117] S32. Install and configure the corresponding driver for the sensor to ensure that it can acquire data streams normally through the ROS driver and perform time synchronization.
[0118] S33. Perform distortion calibration on the RGB camera and internal parameter calibration on the LiDAR to ensure the accuracy of distance measurement and angle scanning.
[0119] S34. Develop an independent ROS node to continuously collect raw data from vision sensors, robotic arm controllers, and LiDAR interfaces; publish the collected data in a standardized message format and attach a timestamp.
[0120] S35. Deploy the pre-trained YOLOv8 model as an independent ROS node. This node:
[0121] Subscribe to the preprocessed visual image stream, perform inference in real time, and output the 2D bounding box, confidence score, and category of each identified gangue;
[0122] Acquire the raw state data of the robotic arm and convert the raw encoder readings into precise joint angles and angular velocities;
[0123] Subscribe to preprocessed LiDAR point cloud data, remove ground points, and apply the K-Means algorithm to the remaining non-ground point cloud to divide it into multiple independent clusters. Each cluster represents a potential gangue target. Perform geometric analysis on the clustered point cloud clusters to calculate their centroid coordinates, principal direction, and other geometric features.
[0124] S36. Design a multi-layer fully connected deep neural network structure:
[0125] The input layer is for feature extraction, robotic arm joint features, and point cloud features; multiple fully connected layers are designed, including ReLU activation function and Batch Normalization to accelerate training; the output layer outputs a unified environment state, such as: the six joint angles of the robotic arm, and the (x, y, z, roll, pitch, yaw, class, size) of N stones, representing the X-axis coordinate, Y-axis coordinate, Z-axis coordinate, roll angle, pitch angle, yaw angle, class, and size, respectively.
[0126] S37. Deploy the trained model as a ROS node. This service subscribes to feature inputs, performs real-time inference, and publishes the final unified environment state data package.
[0127] S38. Transmit the multi-sensor fusion physical layer to the virtual-real interaction control module, such as... Figure 4 As shown, this is used to update the virtual environment status.
[0128] As a preferred embodiment, step 4 specifically includes the following sub-steps:
[0129] S41. Install and configure the ML-Agents plugin in Unity to configure the Agent component for the virtual robotic arm model in the digital twin and specify its behavior.
[0130] S42. Define the observation space:
[0131] In this embodiment, the observation space comprises a total of 16 continuous values, including: joint angles (6 dimensions); information from the gripper to the target (4 dimensions), including the direction vector (3 dimensions) and Euclidean distance (1 dimension); target-to-target region (4 dimensions); and task state (2 dimensions). These values are normalized to enable the neural network to learn more efficiently.
[0132] S43. Define the action space and adopt a continuous action space to allow the agent to perform fine and smooth control.
[0133] In this embodiment, the motion space is 6-dimensional, precisely corresponding to the 6 joints of the robotic arm.
[0134] S44. Define the reward function: Adopt a dense reward design to ensure that the Agent receives clear feedback signals at every step, avoids the exploration difficulties caused by sparse rewards, and greatly accelerates the learning speed.
[0135] In this embodiment, the total reward consists of the following components:
[0136] Total Reward = Grappling Grip Approaching Target Reward + Target Approaching Target Area Reward + Milestone Reward
[0137] + Contact reward + Time penalty + Termination reward;
[0138] It follows the principle of hierarchical guidance, enabling the agent to learn complex tasks step by step.
[0139] S45. In order to enhance the generalization ability of the training strategy and narrow the gap between simulation and reality, certain key parameters are randomized in the digital twin environment.
[0140] Specifically, by randomly changing the friction coefficient, mass, and elasticity of the gangue, and by randomizing the light source and texture, the agent can learn robust policies that are insensitive to changes in these parameters through training in diverse virtual environments.
[0141] S46. Run ML-Agents training, using reinforcement learning algorithms (such as PPO or SAC) to train the control strategy of the robotic arm, such as... Figure 5 As shown, the performance of the Agent is continuously optimized in different scenarios.
[0142] S47. Observe and adjust the learning effect of the Agent during training, such as Figure 6 As shown, by adjusting hyperparameters, reward functions, or adding new constraints, the control accuracy and stability of the robotic arm can be gradually improved.
[0143] S48. After training, deploy the trained model in Unity to verify its performance in a real simulation environment and ensure that it can stably and accurately complete the target action.
[0144] S49. Export the optimized model in ONNX format and deploy the application on the ROS platform.
[0145] As a preferred embodiment, step 5 specifically includes the following sub-steps:
[0146] S51. Import the two components, ros-tcp-connector and unity robotics visualizations, into the Unity platform via Window / Package Manager.
[0147] S52, Configure ROS IP and ROS Port.
[0148] S53. Compile ROS-TCP-Endpoint in the ROS workspace.
[0149] As a preferred embodiment, step 6 specifically includes the following sub-steps:
[0150] S61. Configure the ROS communication plugin in Unity, subscribe to the / joint_states topic, and obtain real-time joint state data through the robotic arm control interface;
[0151] S62. Map the parsed joint angle values to the digital robotic arm model in Unity, convert the joint state data into the angle information of each joint in Unity, and immediately update the joint state and end pose of the virtual robotic arm to complete the actual control virtual function.
[0152] S63. When the digital twin environment receives state updates from the real world, it provides these states as input to the agent's neural network.
[0153] In this embodiment, the Agent's neural network outputs a 6-dimensional continuous action value based on the policy it has learned.
[0154] S64. To ensure the smoothness and safety of the robotic arm's movement, a trajectory planner is integrated here to avoid impact movements of the robotic arm and reduce wear.
[0155] Furthermore, MoveIt is used to perform path planning on the received joint angles and target pose, and calculate a trajectory that allows the robotic arm to move from the current position to the target pose; this trajectory includes the joint angle and pose changes at each time step.
[0156] S65. The calculated trajectory is sent back to Unity as the response data of the ROS service, so that the digital robotic arm model in Unity can perform the corresponding movement according to the planned trajectory.
[0157] S66. Apply the same trajectory to the physical robotic arm so that it performs the same movements as the Unity digital robotic arm model, thereby achieving synchronized action of the virtual object and completing the virtual-real interaction function.
[0158] In summary, this invention achieves virtual-real interaction and linkage between the virtual and real environments by constructing a high-fidelity digital twin model of the robotic arm and its operating environment. Its core lies in utilizing the digital twin environment to provide efficient, safe, and diverse training drivers for the robotic arm's reinforcement learning strategy, enabling rapid pre-training of the strategy. Subsequently, the strategy is fine-tuned and optimized by combining a small amount of interaction data from the real physical environment. Furthermore, through multi-sensor data fusion and a data closed-loop mechanism, the digital twin model and reinforcement learning strategy are updated iteratively and in real time.
[0159] The digital twin robotic arm rock-moving training platform of this invention effectively solves the pain points of traditional robotic arm reinforcement learning training, such as high cost of real-world scenarios, high safety risks, difficulty in reproducing working conditions, and the gap between simulation and reality. It significantly improves the training efficiency, transfer success rate, control accuracy, and environmental adaptability of robotic arm control strategies, and has broad industrial application prospects. It has important application value in fields such as coal mining and intelligent manufacturing.
[0160] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for constructing a digital twin mechanical arm gob training platform based on multi-sensor fusion, characterized in that, Comprise the following steps: Step 1, the jaw file is assembled and added to the URDF file of the mechanical arm, the model of the mechanical arm and the jaw is imported into the Unity engine through the URDFImporter tool; Step 2, build a digital twin modeling module, discretize the conveyor belt, model and render using 3dsmax, export as fbx file and import into Unity environment, build a digital environment matching the actual environment; Step 3, build a multi-sensor fusion module, collect visual data, mechanical arm state data and laser radar data of the real physical scene, generate a unified environment state through feature extraction and recognition and multi-source data fusion based on deep learning; Step 4, build a reinforcement learning training module, use ML-Agent for reinforcement learning training, establish the state space, action space and reward function of the reinforcement learning algorithm, train the mechanical arm stone strategy, and get the trained model; Step 5, build a communication environment with ROS in the Unity environment; Step 6, build a virtual-real interaction control module, export the trained model to onnx format and deploy it to a computer device, communicate with the controller of the real mechanical arm through ROS, and listen to the joint_states topic of the real mechanical arm at the same time, synchronize the joint state to the Unity environment, realize the visualization of data.
2. The method for constructing a training platform for a digital twin robotic arm to remove rock based on multi-sensor fusion as described in claim 1, characterized in that, The step 1 comprises the following sub-steps: Step 1.1, open the STP file of the jaw base and the finger using SolidWorks, align the finger to assemble on the jaw base, and make the finger and the jaw base into one; Step 1.2, install the sw_urdf_exporter plug-in, export the jaw file to the URDF file, add it to the URDF file of the mechanical arm, adjust the coordinate information to align the end position of the mechanical arm, and assemble it; Step 1.3, install the URDF Importer plug-in in the Package Manager of Unity; Step 1.4, import the URDF file and resource file of the mechanical arm into the Asset folder of Unity; Step 1.5, select Import URDF under the path of URDF Importer in the Unity menu, load the complete mechanical arm model into the Unity scene.
3. The method of claim 1, wherein the method is characterized by: The step 2 comprises the following sub-steps: Step 2.1, discretize the conveyor belt, divide the conveyor belt into equal length segments, the conveyor belt running speed is v, and the time step of each segment is ; Step 2.2, assuming the conveyor belt is... , , The coal quantities of each section are respectively , and Starting from section 1, the coal for section 2 flows in from section 1. The coal in the section is from the first Section inflow; Step 2.3, passing Time interval, each segment of coal flow on the conveyor belt transports its coal quantity to the next segment while receiving an equal quantity of coal from the previous segment. Step 2.4, establish a three-dimensional model of the belt conveyor: draw the shape of the belt conveyor in 3dsMAX, adjust the size, proportion and shape to match the actual belt conveyor, modify and refine the model using editing tools to present the details of the belt conveyor, and add details and decoration elements according to the appearance and characteristics of the belt conveyor; Step 2.5, establish a three-dimensional model of the conveyor belt: build a support frame, use a cuboid to initialize the basic structure of the support frame, copy several identical structures to form a conveyor belt, and optimize it using the welding command of points; set the light, add light sources using the 3dsMAX light tool and adjust the brightness, color and shadow attributes; Set the rendering, adjust the resolution, shadow and reflection; Step 2.
6. Select the established three-dimensional model of the belt conveyor in 3dsMAX and export it in FBX file format, placing it in the Unity project Assets folder. In Unity, select the imported three-dimensional model of the conveyor and create or assign a material to it in the Inspector panel. Step 2.
7. Create a gangue model in Unity and randomly generate diverse gangue shapes and initial distributions. Step 2.
8. Create a laboratory training environment and an area for placing gangue. Adjust the positions and physical properties of the conveyor belt, robotic arm, and gangue model to test whether they can normally contact and collide, and prepare for training.
4. The method of claim 3, wherein the method is characterized by, In step 2.3, the time interval is passed after the first segment of coal is represented as: ; At the moment, the , , coal-carrying capacity of the section is respectively represented as , and , and there is a relationship: ; wherein is the time The the coal flow intensity.
5. The method of claim 1, wherein the method is characterized by: Step 3 includes the following sub-steps: Step 3.
1. Select an RGB-D camera to be fixedly installed above the working area of the robotic arm, covering the range of the gangue picking operation. Connect the robotic arm controller data interface to read joint angle and angular velocity data in real time. Install a laser radar to obtain the position of the gangue point cloud. Step 3.
2. Configure the sensor driver and perform time synchronization. Calibrate the RGB-D camera for distortion and the laser radar for internal parameters. Step 3.
3. Collect raw data from the RGB-D camera, robotic arm controller, and laser radar interface through independent ROS nodes, publish in standardized message format, and attach timestamps. Step 3.
4. Deploy the trained YOLOv8 model as an independent ROS node to subscribe to preprocessed visual image streams, perform real-time inference, and output two-dimensional bounding boxes, confidence, and class for each identified gangue. Obtain the original robotic arm state data, convert the original encoder readings to joint angles and angular velocities. Subscribe to preprocessed laser radar point cloud data, remove ground points, and segment the remaining non-ground point cloud into multiple independent clusters using the K-Means algorithm. Each cluster represents a potential gangue target. Perform geometric analysis on the clustered point cloud clusters to calculate the centroid coordinates and principal direction. Step 3.
5. Construct a multi-layer fully connected deep neural network and train it. Deploy the trained network model as an independent ROS node to subscribe to feature inputs and perform real-time inference, and publish unified environment state data packets. Step 3.
6. Transmit the multi-sensor fusion entity layer to the virtual-real interaction control module to update the virtual environment state.
6. The method of claim 5, wherein the method is characterized by: In Step 3.5, the multi-layer fully connected deep neural network: The input layer is feature extraction, robotic arm joint nodes, and point cloud features. There are multiple fully connected layers, including ReLU activation functions and Batch Normalization to accelerate training. The output layer outputs unified environment state, including but not limited to: 6 joint angles of the robotic arm, X-axis, Y-axis, and Z-axis coordinates of N gangues, roll angle, pitch angle, yaw angle, class, and size.
7. The method of claim 1, wherein the method is characterized by: Step 4 includes the following sub-steps: Step 4.
1. Install and configure the ML-Agents plugin in Unity. Configure the Agent component for the virtual robotic arm model in the digital twin and specify its behavior. Step 4.
2. Define the observation space and perform normalization. The observation space is represented as 16 continuous values, including: 6-dimensional joint angles, 4-dimensional gripper-to-target information, 4-dimensional target-to-target region information, and 2-dimensional task status; Step 4.3, define the action space and control it through the Agent; the action space is 6-dimensional, corresponding to the 6 joints of the robot arm; Step 4.4, define the reward function, use dense rewards to make the Agent receive explicit feedback signals at each step, and follow the hierarchical guidance principle to make the Agent learn complex tasks gradually; The total reward function is represented as: Total reward = gripper-to-target reward + target-to-target region reward + milestone reward + contact reward + time penalty + termination reward; Step 4.5, randomize key parameters in the digital twin environment, randomly change the friction coefficient, mass, and elasticity of the gangue; randomize the light source and texture; Step 4.6, perform ML-Agents training, use reinforcement learning algorithms to train the control strategy of the robot arm, and optimize the Agent's performance in different scenarios; Step 4.7, observe and debug the learning effect of the Agent during training, and gradually improve the control of the robot arm by adjusting hyperparameters, reward functions, or adding new constraints; Step 4.8, deploy the trained virtual robot arm model in Unity for optimization, and export the optimized model through the ONNX format, and deploy and apply it on the ROS side.
8. The method of claim 1, wherein the method is characterized by: The step 5 includes the following sub-steps: Step 5.1, import ros-tcp-connector and unity robotics visualizations components through Package Manager under Window in Unity platform; Step 5.2, configure ROS IP and ROS Port; Step 5.3, place ROS-TCP-Endpoint in the ROS workspace and compile it.
9. The method of claim 1, wherein the method is characterized by, The step 6 includes the following sub-steps: Step 6.1, configure the ROS communication plug-in in Unity, subscribe to the joint_states topic, and get real-time joint state data through the robot arm control interface and analyze it; Step 6.2, map the analyzed joint angle values to the digital robot arm model in Unity, convert the joint state data to angle information for each joint in Unity, update the joint state and end pose of the virtual robot, and perform actual control of the virtual; Step 6.3, when the digital twin environment receives state updates from the real world, provide the received state as input to the neural network of the Agent, and the Agent's neural network outputs 6-dimensional continuous action values according to its learned strategy; Step 6.4, integrate the trajectory planner, use MoveIt to plan the path for the received joint angles and target pose, calculate the trajectory for the robot arm to reach the target pose from the current position, including the joint angle and pose change at each time step; Step 6.
5. Send the calculated trajectory back to the Unity side as response data of the ROS service, control the digital robot model in Unity to perform corresponding motion according to the planned trajectory; Step 6.
6. Apply the same trajectory to the real robot, perform the same motion as the Unity digital robot model, realize the synchronous action of virtual control of the real, and complete the virtual-real interaction.
10. A digital twin mechanical arm gob training platform based on multi-sensor fusion, constructed by the method of any one of claims 1 to 9, characterized in that, Comprise: Digital twin modeling module: build a digital twin model of the physical robot and the working environment, which reflects the response characteristics of the real robot and the interaction with the environment, so as to maintain consistency between the digital twin model and the real physical entity; Multi-sensor fusion module: used for real-time collection of visual data, robot state data and laser radar data of real physical scene, feature extraction and recognition of collected data, and multi-source data fusion based on deep learning, output unified environment state; Reinforcement learning training module: use the digital twin environment as the trainer of reinforcement learning, and use the robot as the Agent to learn by trial and error in the environment; Virtual-real interaction control module: used for real-time mapping of real physical environment to digital twin environment, and converting the control strategy trained in digital twin environment to control instructions of real robot and issuing.
Citation Information
Cited By
Cooperative control method and system for man-machine cooperative assembly robot based on digital twinning
CN121893295A