Coal mine robot action level planning device and method based on digital expression and MR

Through the action-level planning method of coal mine robots based on digital cousins ​​and mixed reality, the problems of low efficiency and poor adaptability in the existing technology are solved, efficient and highly adaptable action planning is achieved, and pre-execution is performed in a virtual environment.

CN120056111APending Publication Date: 2025-05-30TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510251811.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing coal mine robot action planning methods are inefficient and poorly adaptable, cannot be pre-executed in a virtual environment, and lack intuitive visual performance, making it difficult to quickly verify the feasibility and effectiveness of action planning results.

Method used

The action-level planning device and method of coal mine robot based on digital cousins ​​and mixed reality is used to generate the final action planning results through the determination of digital cousins ​​scene information, the generation of imitation learning models, the adjustment of multimodal interactions, the time and space deduction and pre-execution.

Benefits of technology

It improves the efficiency and adaptability of action planning, can pre-execute action planning in a virtual environment, and provides intuitive visual performance, enhancing the feasibility and effect of planning results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120056111A_ABST
    Figure CN120056111A_ABST
Patent Text Reader

Abstract

The invention discloses a coal mine robot action level planning device and method based on digital expressions and MR, and relates to the technical field of man-machine hybrid enhanced intelligence. Acquiring scene information data through an information data acquisition module; the digital expression scene information determination module determines digital expression scene information; the action data determination module inputs the digital expression scene information into an imitation learning model to obtain action data; then, the action pre-planning module determines an action pre-planning result according to the action data; the adjustment module adopts a multi-modal interaction method to adjust the action pre-planning result to obtain an adjustment planning result; the space-time deduction module processes the adjustment planning result by adopting a preset action planning algorithm, and carries out space-time deduction according to a time dimension to obtain a space-time deduction result; and the action planning module determines an action planning result according to the spatio-temporal deduction result. The invention aims to improve the action planning efficiency and adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of human-machine hybrid enhanced intelligence technology, and particularly to a device and method for action-level planning of coal mine robots based on digital cousins and MR. Background Art

[0002] Currently, the research and development and application of coal mine robots have become an important part of the intelligent construction of coal mines. The original intention of the application of coal mine robots is to achieve "replacing humans with robots". However, with the in-depth research and application, limited by the intelligent level of coal mine robots, they can only complete tasks independently in a few simple scenarios. In more cases, the in-depth participation of coal mine operators is still required to cooperate with coal mine robots to complete operation tasks.

[0003] The working process of coal mine robots can be divided into three stages: perception, decision-making, and execution. Decision-making is the link between perceptual information and actual actions and is the core factor determining the working quality of coal mine robots. According to the principle from macro to micro, the operation goals of coal mine robots can be decomposed into three levels: task level, behavior level, and action level. Among them, action-level planning is an important link in the decision-making process, which can refine high-level task goals into specific action instructions to ensure that coal mine robots execute accurately according to the established task requirements.

[0004] Currently, most coal mine underground robots adopt the collaborative type of main body-manipulator. However, existing action planning schemes for coal mine robots are all planned separately for the movement path of the robot or the trajectory of the manipulator alone, without considering the robot main body and the manipulator as a whole for joint planning. The action planning efficiency is low and the adaptability is poor.

[0005] The vast majority of coal mine robot action planning schemes rely on pure algorithm calculations, do not have the ability of spatio-temporal deduction, cannot be pre-executed in a virtual environment, and lack an intuitive visualization performance of the planning results. They can neither consider complex spatio-temporal relationships in the planning nor quickly verify the feasibility and effect of the action planning results.

[0006] Currently, for the action planning of coal mine robots based on digital twins, the planning process is extended to the virtual space. However, establishing a digital twin environment depends on high-precision modeling, a large number of sensors, and accurate data updated in real time, with a long cycle, high cost, and no good cross-domain generalization ability.

[0007] Due to the dynamics and complexity of the underground environment and tasks, human-robot collaboration is the main operation mode of current coal mine robots. However, in the context of human-robot collaboration, existing action planning methods for coal mine robots do not consider the role of humans in the action planning of coal mine robots; in other fields besides the coal mine field, some technical solutions introduce the role of humans in robot action planning, but only limit to defining motion goals or providing solutions completely independent of the robot's autonomous planning, without fully leveraging the advantages of human's flexible intervention and adjustment, nor cleverly combining machine intelligence and human intelligence.

[0008] Therefore, how to improve the action planning efficiency and adaptability is crucial. Summary of the Invention

[0009] The purpose of this application is to provide a coal mine robot action-level planning device and method based on Digital Twin and MR, which can improve the action planning efficiency and adaptability.

[0010] To achieve the above purpose, this application provides the following solutions:

[0011] In the first aspect, this application provides a coal mine robot action-level planning device based on Digital Twin and MR, including:

[0012] An information data acquisition module, configured to acquire scene information data; the scene information data includes a single-frame RGB image collected underground in a coal mine;

[0013] A Digital Twin scene information determination module, connected to the information data acquisition module, configured to determine Digital Twin scene information; the Digital Twin scene information is determined based on a 3D asset library virtual model according to the scene information data; the 3D asset library virtual model is a preset model library covering operation objects, operation tools, terrain structures, mechanical equipment, and auxiliary equipment for underground coal mine operation tasks;

[0014] An action data determination module, connected to the Digital Twin scene information determination module, configured to input the Digital Twin scene information into an imitation learning model to obtain action data; the action data includes: coal mine robot motion pose information and robotic arm operation pose information; the imitation learning model is obtained through simulation training and zero-shot transfer based on a coal mine robot behavior model; the coal mine robot behavior model is obtained through reinforcement learning simulation training of a coal mine robot model; the coal mine robot model is a physical model constructed using a 3D modeling tool based on the entity parameters of the coal mine robot; the entity parameters include: size ratio, structure data, and shape data;

[0015] An action pre-planning module, connected to the action data determination module, configured to determine an action pre-planning result according to the action data;

[0016] An adjustment module, connected to the action pre-planning module, is configured to adjust the action pre-planning result by using a multi-modal interaction method to obtain an adjusted planning result;

[0017] A spatio-temporal deduction module, connected to the adjustment module, is configured to process the adjusted planning result by using a preset action planning algorithm and perform spatio-temporal deduction according to the time dimension to obtain a spatio-temporal deduction result; the spatio-temporal deduction result includes the movement path of the coal mine robot and the movement trajectory of the robotic arm;

[0018] An action planning module, connected to the spatio-temporal deduction module, is configured to determine an action planning result according to the spatio-temporal deduction result; the action planning result includes: position coordinates, attitude information, and path points.

[0019] In a second aspect, the present application provides a method for action-level planning of a coal mine robot based on digital cousins and MR. The method for action-level planning of a coal mine robot based on digital cousins and MR is implemented by using a device for action-level planning of a coal mine robot based on digital cousins and MR. The method for action-level planning of a coal mine robot based on digital cousins and MR includes:

[0020] Obtain scene information data; the scene information data includes a single-frame RGB image collected underground in a coal mine;

[0021] Determine digital cousin scene information; the digital cousin scene information is determined based on a 3D asset library virtual model according to the scene information data; the 3D asset library virtual model is a preset model library covering operation objects, operation tools, terrain structures, mechanical equipment, and auxiliary equipment for underground coal mine operation tasks;

[0022] Input the digital cousin scene information into an imitation learning model to obtain action data; the action data includes: the movement pose information of the coal mine robot and the operation pose information of the robotic arm; the imitation learning model is obtained through simulation training and zero-shot transfer based on a coal mine robot behavior model; the coal mine robot behavior model is obtained through reinforcement learning simulation training of a coal mine robot model; the coal mine robot model is a physical model constructed by using a 3D modeling tool based on the entity parameters of the coal mine robot; the entity parameters include: size ratio, structure data, and appearance data;

[0023] Determine an action pre-planning result according to the action data;

[0024] Adjust the action pre-planning result by using a multi-modal interaction method to obtain an adjusted planning result;

[0025] Process the adjusted planning result using a preset action planning algorithm, and perform spatio-temporal deduction according to the time dimension to obtain a spatio-temporal deduction result; the spatio-temporal deduction result includes the movement path of the coal mine robot and the movement trajectory of the robotic arm.

[0026] Determine the action planning result according to the spatio-temporal deduction result; the action planning result includes: position coordinates, attitude information, and path points.

[0027] According to the specific embodiments provided in this application, this application has the following technical effects:

[0028] This application provides a device and method for action-level planning of a coal mine robot based on digital cousins and MR. First, determine the digital cousin scenario. Then, after performing reinforcement learning simulation training and simulation training with zero-shot transfer, the generated imitation learning model generates an action pre-planning result. Further adjust the pre-planning result through multi-modal interaction; then, perform spatio-temporal deduction and pre-execution. Finally, generate the final action planning result, that is, the action planning result, so that the coal mine robot can execute the operation. Through the adjustment and deduction of the pre-planning result, the generalization ability of the action planning is improved, and a better planning result that takes into account both machine intelligence and human intelligence can be obtained based on the imitation learning model. Thus, this application can improve the action planning efficiency and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a structural diagram of a device for action-level planning of a coal mine robot based on digital cousins and MR;

[0031] Figure 2 It is a flowchart of the action pre-planning stage of a coal mine robot based on digital cousins;

[0032] Figure 3 It is a schematic diagram of the principle of generating and understanding the digital cousin scenario in step S1;

[0033] Figure 4 It is a schematic diagram of the principle of integrating the pre-definition of the behavior of a coal mine robot and the reinforcement learning model in step S2;

[0034] Figure 5 It is a schematic diagram of the principle of imitation learning and zero-shot transfer in step S3;

[0035] Figure 6It is a flow chart of the action optimization stage of a coal mine robot based on mixed reality;

[0036] Figure 7 It is a schematic diagram of the principle of presenting the action pre - planning result in step S5;

[0037] Figure 8 It is a schematic diagram of the principle of adjusting the planning result based on multi - modal interaction in step S6;

[0038] Figure 9 It is a schematic diagram of the principle of space - time deduction and pre - execution in step S7. Specific implementation mode

[0039] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0040] The related technology discloses "A path planning method for a coal mine search - and - rescue robot improved based on the A* algorithm", which optimizes the neighborhood search method, increases the search range, adjusts the weight of the heuristic function, applies B - spline to smooth the path, and finally optimizes the path by introducing the D algorithm. And in "A motion planning method and device for a robotic arm on a coal mine autonomous inspection platform", by voxelizing the obstacle objects and combining with a pre - trained motion planner based on the Q neural network, the moving operation of the robotic arm in a complex scene is planned in real - time.

[0041] The above two solutions separately provide the action planning methods for the coal mine robot body and the carried robotic arm. However, most current coal mine robots are in the form of body - robotic arm cooperation, and the above two solutions are difficult to meet their complete action planning requirements; in addition, both solutions are based on algorithm calculations and are not combined with the space - time deduction ability of the virtual space, resulting in insufficient ability to handle complex and dynamic scenarios.

[0042] The related technology discloses "A digital twin system for a multi - functional quadruped robot in coal mines and its operation method", which includes a robot virtual planning subsystem that can perform robot path planning and synchronize the results to the autonomous control module of the robot to complete walking according to the path planning results.

[0043] Although the digital twin technology is used to transfer the path planning of the coal mine robot from the physical space to the virtual space, the factor of people is not considered. Under the condition of co - existence of humans and machines, the decision - making ability of people can play a crucial role and should be fully utilized. In addition, the construction cost of the virtual space based on the digital twin technology is high and the generalization ability is weak, and it may not be applicable to new environments or conditions.

[0044] The related technology discloses "a human-robot collaborative motion planning method and system for a mobile assembly robot". In each control cycle, the load type and pose information are obtained through an identification and positioning device, and the motion of the assembly robot is reasonably allocated and the motion trajectory of the end load is planned according to the load information.

[0045] Although the role of humans is introduced in the motion planning of the mobile assembly robot, the motion of the robot is still essentially completed by the motion allocation and trajectory planning algorithms. The main role of humans in this process is to define the position and orientation of the end load of the robotic arm, and they do not really participate in the planning of the specific motion mode of the robot.

[0046] The related technology discloses "a human-robot collaborative path planning method in an unstructured environment for a mobile robot", which mixes and synthesizes the path planned by the operator and the path autonomously planned by the robot to form a path planning with humans in the loop.

[0047] However, the autonomous planning of the robot and the path planning of the operator are carried out simultaneously and independently, and the global computing advantages of the robot and the intuitive judgment ability of the operator are not fully utilized. A more ideal way should be to first obtain a preliminary optimal solution through the computing power of the robot, and then the operator makes flexible responses and fine-tuning based on rich intuition and experience.

[0048] Digital cousin is a new concept proposed by the team of Professor Li Feifei at Stanford University in the United States. It not only retains the advantages of digital twins but also greatly reduces the generation cost from the real to the simulation environment, while improving the generalization ability of learning. The mixed reality technology can be used as an interface for humans to deeply participate in the robot motion planning, and then realize the organic integration of machine intelligence and human intelligence in the robot motion planning. Combining the digital cousin concept with the mixed reality (MR) technology is expected to provide a new way for the motion-level planning of coal mine operation robots.

[0049] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0050] In an exemplary embodiment, as Figure 1 shown, a coal mine robot motion-level planning device based on digital cousin and MR is provided. The coal mine robot motion-level planning device based on digital cousin and MR includes: an information data acquisition module, a digital cousin scene information determination module, an action data determination module, an action pre-planning module, an adjustment module, a spatio-temporal deduction module, and an action planning module.

[0051] The information data acquisition module is used to acquire scene information data; the scene information data includes single-frame RGB images collected underground in coal mines.

[0052] The digital twin scene information determination module is connected to the information data acquisition module. The digital twin scene information determination module is used to determine digital twin scene information; the digital twin scene information is determined based on the virtual model in the 3D asset library according to the scene information data. The 3D asset library virtual model is a preset model library covering operation objects, operation tools, terrain structures, mechanical equipment, and auxiliary equipment for underground coal mine operation tasks.

[0053] The action data determination module is connected to the digital twin scene information determination module. The action data determination module is used to input the digital twin scene information into the imitation learning model to obtain action data; the action data includes: the motion pose information of the coal mine robot and the operation pose information of the robotic arm; the imitation learning model is obtained through simulation training and zero-shot transfer based on the coal mine robot behavior model; the coal mine robot behavior model is obtained through reinforcement learning simulation training of the coal mine robot model; the coal mine robot model is a physical model constructed using 3D modeling tools based on the entity parameters of the coal mine robot; the entity parameters include: size ratio, structure data, and appearance data.

[0054] The action pre-planning module is connected to the action data determination module. The action pre-planning module is used to determine the action pre-planning result according to the action data.

[0055] The adjustment module is connected to the action pre-planning module. The adjustment module is used to adjust the action pre-planning result by using the method of multi-modal interaction to obtain the adjusted planning result.

[0056] The spatio-temporal deduction module is connected to the adjustment module. The spatio-temporal deduction module is used to process the adjusted planning result by using a preset action planning algorithm and perform spatio-temporal deduction according to the time dimension to obtain the spatio-temporal deduction result; the spatio-temporal deduction result includes the movement path of the coal mine robot and the movement trajectory of the robotic arm.

[0057] The action planning module is connected to the spatio-temporal deduction module. The action planning module is used to determine the action planning result according to the spatio-temporal deduction result. The action planning result includes: position coordinates, attitude information, and path points.

[0058] In one embodiment, the digital twin scene information determination module includes: a semantic segmentation sub-module, a label annotation sub-module, a point cloud data determination sub-module, a semantic similarity determination sub-module, a scene construction sub-module, and a matching sub-module.

[0059] The semantic segmentation sub-module is connected to the information data acquisition module. The semantic segmentation sub-module is used to perform semantic segmentation on the scene information data to obtain the segmented information data.

[0060] The label annotation sub-module is connected to the semantic segmentation sub-module. The label annotation sub-module is used to perform object label annotation on the segmentation information data by using a large language model to obtain annotation information.

[0061] The point cloud data determination sub-module is connected to the label annotation sub-module. The point cloud data determination sub-module is used to determine point cloud data according to the annotation information; the point cloud data includes three-dimensional position information.

[0062] The semantic similarity determination sub-module is connected to the point cloud data determination sub-module. The semantic similarity determination sub-module is used to determine the semantic similarity according to the annotation information and the virtual model in the 3D asset library.

[0063] The scene construction sub-module is connected to the semantic similarity determination sub-module. The scene construction sub-module is used to construct a digital twin scene by using the Unity3d engine.

[0064] The matching sub-module is connected to the scene construction sub-module. The matching sub-module is used to perform digital twin matching in the digital twin scene according to the point cloud data and the semantic similarity to obtain digital twin scene information.

[0065] An embodiment of the present application also provides a method for action-level planning of a coal mine robot based on digital twins and MR. This method is implemented by using a device for action-level planning of a coal mine robot based on digital twins and MR. The method for action-level planning of a coal mine robot based on digital twins and MR includes:

[0066] Obtain scene information data. The scene information data includes a single-frame RGB image collected underground in a coal mine.

[0067] Determine digital twin scene information. The digital twin scene information is determined based on the virtual model in the 3D asset library according to the scene information data; the virtual model in the 3D asset library is a preset model library covering operation objects, operation tools, terrain structures, mechanical equipment, and auxiliary equipment for underground coal mine operation tasks.

[0068] As an optional implementation manner, determining digital twin scene information specifically includes:

[0069] Perform semantic segmentation on the scene information data to obtain segmentation information data; perform object label annotation on the segmentation information data by using a large language model to obtain annotation information; determine point cloud data according to the annotation information; the point cloud data includes three-dimensional position information.

[0070] Determine the semantic similarity according to the annotation information and the virtual model in the 3D asset library; construct a digital twin scene by using the Unity 3d engine; perform digital twin matching in the digital twin scene according to the point cloud data and the semantic similarity to obtain digital twin scene information.

[0071] Input the digital cousin scenario information into the imitation learning model to obtain action data. The action data includes: the motion pose information of the coal mine robot and the operation pose information of the robotic arm; the imitation learning model is obtained through simulation training and zero-shot transfer based on the coal mine robot behavior model; the coal mine robot behavior model is obtained through reinforcement learning simulation training of the coal mine robot model; the coal mine robot model is a physical model constructed using a 3D modeling tool based on the entity parameters of the coal mine robot; the entity parameters include: dimensional ratio, structural data, and appearance data.

[0072] A method for determining the coal mine robot behavior model includes:

[0073] Obtain task data; the task data is the operation data when the coal mine robot executes tasks. Use the ML-Agents toolkit to construct a reinforcement learning simulation environment; disassemble the task data to obtain basic behavior data; the basic behavior data includes: operation data corresponding to movement, rotation, grasping, and placing. Use a reinforcement learning algorithm, based on a set reinforcement learning reward and punishment mechanism, in the reinforcement learning simulation environment, and perform multi-task optimization training on the coal mine robot model according to the basic behavior data to obtain the coal mine robot behavior model.

[0074] As an optional implementation, a method for determining the imitation learning model includes:

[0075] Obtain a demonstration dataset; the demonstration dataset is the state, action, and reward data of the coal mine robot behavior model in different scenarios; screen the demonstration dataset to obtain screened data; based on the screened data and the coal mine robot behavior model, perform simulation training in Unity 3d based on adding domain randomization, and transfer the trained model to the coal mine robot to obtain the imitation learning model; domain randomization includes: randomly changing lighting, object color, and object friction coefficient.

[0076] Determine the action pre-planning result according to the action data.

[0077] Use the multi-modal interaction method to adjust the action pre-planning result to obtain the adjusted planning result.

[0078] In one embodiment, using the multi-modal interaction method to adjust the action pre-planning result to obtain the adjusted planning result specifically includes:

[0079] Perform interference analysis on the action pre-planning result through a mixed reality headset to determine whether to make adjustments and obtain a judgment result.

[0080] If the judgment result is yes, then use gaze interaction in multi-modal interaction to select waypoints in the space of the action pre-planning result as the path points for movement.

[0081] Based on gesture interaction in multimodal interaction, interpolate and mark the action pre-planning results to determine the key points for optimizing the manipulator trajectory, and obtain the manipulator trajectory adjustment points.

[0082] Based on voice interaction in multimodal interaction, perform action optimization confirmation according to the moving path points and the manipulator trajectory adjustment points to obtain the adjusted planning results.

[0083] Use a preset action planning algorithm to process the adjusted planning results, and perform spatio-temporal deduction according to the time dimension to obtain the spatio-temporal deduction results. The spatio-temporal deduction results include the moving path of the coal mine robot and the moving trajectory of the manipulator. The preset action planning algorithms include: A* algorithm and RRT-Connect algorithm; among them, the A* algorithm is used to determine the moving path of the coal mine robot, and the RRT-Connect algorithm is used to determine the moving trajectory of the manipulator.

[0084] In one embodiment, use a preset action planning algorithm to process the adjusted planning results, and perform spatio-temporal deduction according to the time dimension to obtain the spatio-temporal deduction results, specifically including:

[0085] Use a preset action planning algorithm to optimize the adjusted planning results to obtain the action optimization results; perform spatio-temporal deduction according to the action optimization results according to the time dimension to obtain the spatio-temporal deduction results.

[0086] Determine the action planning results according to the spatio-temporal deduction results. The action planning results include: position coordinates, attitude information, and path points.

[0087] Among them, determining the action planning results according to the spatio-temporal deduction results specifically includes:

[0088] Judge whether it conforms to the expected action information according to the spatio-temporal deduction results; if so, perform data conversion on the spatio-temporal deduction results according to the instruction format of the preset standard to obtain the conversion data; determine the action planning results according to the conversion data.

[0089] The present application also provides an application scenario, which applies the above-mentioned action-level planning method for coal mine robots based on digital cousins and MR. Specifically: The action-level planning method for coal mine robots based on digital cousins and MR provided in this embodiment can be applied to the coal mine underground transportation scenario. The coal mine underground transportation scenario includes an information acquisition link, an action planning link, and a transportation execution link; the environmental information (scenario information) collected underground in the coal mine enters the action planning link from the information acquisition link, and through the action planning link, reinforcement learning and imitation learning in the digital cousin environment are carried out to realize the action pre-planning of the coal mine robot, and the pre-planning result is optimized to obtain the corresponding action planning result, and then enters the downstream transportation execution link, and moves according to the action planning result to realize the coal mine underground transportation operation. The action-level planning method for coal mine robots based on digital cousins and MR provided in this embodiment belongs to the process of action pre-planning, adjustment, deduction and optimization in the action planning link.

[0090] As Figure 2 shown, the technical concept of the method adopted in the present application mainly includes two stages: action pre-planning of coal mine robots based on digital cousins and action optimization of coal mine robots based on mixed reality.

[0091] In the stage of action pre-planning of coal mine robots based on digital cousins, the following steps are included.

[0092] S1: Generation and understanding of digital cousin scenarios. The principle is as shown in the appendix Figure 3 shown.

[0093] S101: Collection of scene data, i.e., scene information data, and generation of object labels. A single-frame RGB image is collected from the real coal mine underground scene; a pre-trained semantic segmentation model is used to generate masks of objects related to the operation task in the image; the large language model is used to label each object in the image one by one, and the description information is passed to the semantic segmentation model to ensure that each object has a unique label.

[0094] The semantic segmentation model is Grounded SAM 2, which can combine language and visual inputs to identify and segment complex objects and is suitable for semantic segmentation in complex scenarios. The large language model is GPT-4o, which inherits the core capabilities of GPT-4 and has more advantages in terms of speed and cost.

[0095] S102: Generation of object point clouds and state estimation. A monocular depth estimation model is used to generate the depth map of the scene, so as to obtain the three-dimensional position information of each object; the depth map is converted into point cloud data, and combined with the mask information of the object, a three-dimensional point cloud representation of each object is generated; the pose and size of the object are estimated according to the point cloud data to prepare for the matching and placement of virtual objects in the future.

[0096] The monocular depth estimation model is DepthAnythingV2, which captures hierarchical and depth changes more accurately compared to other depth estimation models.

[0097] S103: Digital cousin matching for language-vision model fusion. Calculate the semantic similarity between the actual object labels (annotation information) and the virtual models in the 3D asset library through a multimodal model to complete the preliminary category screening of the 3D asset library; use a visual feature embedding model to calculate the geometric feature distance between each actual object and the virtual model, that is, determine the semantic similarity based on the annotation information and the virtual models in the 3D asset library; screen out the virtual object with the closest shape and semantics as the digital cousin object.

[0098] The 3D asset library is a pre-established model library for coal mine underground operation tasks, covering operation objects, operation tools, terrain structures, mechanical equipment, and auxiliary facilities involved in coal mine underground operation tasks.

[0099] The multimodal model is CLIP, which has the ability to understand text descriptions and establish associations with visual information; the visual feature embedding model is DINOv2, which has high-quality visual feature embedding ability and can accurately capture the geometric and semantic features of objects.

[0100] S104: Digital cousin scene generation and setting. Use the Unity 3d engine to create a digital cousin scene, place the corresponding digital cousin objects at the corresponding positions in the scene according to the depth information and spatial positions of each object, and adjust the size of the objects to match the real size; set appropriate physical properties for the digital cousin objects in the scene to ensure that their physical behaviors are consistent with those of the actual objects; that is, perform digital cousin matching in the digital cousin scene according to the point cloud data and semantic similarity to obtain digital cousin scene information.

[0101] The setting of physical properties is achieved through rigid bodies, colliders, and joint components in the Unity 3d engine.

[0102] S105: Enhance the generalization ability of the digital cousin scene. Randomize the appearance of objects using the Shader and material library of the Unity 3d engine to improve the visual generalization ability of the robot; moderately randomize the positions and sizes of objects in the Unity 3d engine to improve the task adaptability of the robot in different layouts and positions; add random obstacles or dynamic elements in Unity3d to further enhance the diversity of the scene.

[0103] S2: Integrate the behavior predefinition of the coal mine robot and the reinforcement learning model; the principle is as Figure 4 shown.

[0104] S201: Construction of a coal mine robot model. Use 3D modeling tools to create a coal mine robot model, ensuring that the size ratio, structure, and appearance of the model are exactly the same as the physical entity of the coal mine robot, and integrate it into the digital twin scenario; use the joint components of the Unity 3d engine to control the rotation and movement range of each moving part of the coal mine robot model; add a physics engine to the coal mine robot model to ensure that it conforms to the actual motion laws.

[0105] The 3D modeling tool is Blender, which integrates the modeling and rendering processes and supports FBX, OBJ, and STL formats, facilitating seamless integration with Unity3d.

[0106] The joint components include Configurable Joint, Hinge Joint, Fixed Joint, and Spring Joint. Among them, Configurable Joint is used to control the complex translation and rotation of the coal mine robot body, Hinge Joint is used to control the rotation of the robotic arm joints, Fixed Joint is used to fix the body to the robotic arm base, and Spring Joint is used to control the grasping of the end effector of the robotic arm.

[0107] S202: Construction of a reinforcement learning simulation environment. Install and configure the ML-Agents toolkit, and attach the core scripts to the coal mine robot model and the digital twin environment; in the training configuration YAML file, select an appropriate reinforcement learning algorithm and set the hyperparameters; establish Socket communication between the mlagents Python library and Unity3d to achieve real-time data interaction between the Python-side training and the Unity3d simulation environment.

[0108] The core scripts include the Agent script and the EnvironmentManager script. The Agent script is attached to the robot model and is responsible for the behavior and learning of the robot model; the EnvironmentManager script is attached to the digital twin environment and is responsible for the management and coordination of the overall environment.

[0109] The reinforcement learning algorithm is the Deep Deterministic Policy Gradient (DDPG) algorithm, which is suitable for continuous action spaces and has good results in robot training.

[0110] S203: Definition of Task Skills and Reward and Punishment Mechanisms. Decompose the complex tasks of coal mine robots into basic behaviors; for each behavior, set up a reinforcement learning reward and punishment mechanism and write it into the OnActionReceived() method in the form of code; call the OnActionReceived() method in the AddReward() or EndEpisode() method so that the robot can obtain corresponding feedback when it succeeds or fails each time. The basic behaviors include moving, rotating, grasping, and placing.

[0111] Specifically, for the reinforcement learning reward and punishment mechanism, rewards are given when the coal mine robot agent successfully completes a behavior, while punishments are imposed when the behavior fails to be completed or when a collision occurs with an obstacle.

[0112] S204: Training of Behavior Models and Optimization of Multi-Task Learning. Start training using a reinforcement learning framework, and continuously optimize the strategy of each behavior model through repeated execution and reward feedback; under the multi-task learning framework, use a scheduling mechanism to switch between different behaviors to ensure that the reinforcement learning model achieves good policy optimization effects on each behavior task. The reinforcement learning framework is TensorFlow or PyTorch.

[0113] S205: Generation and Storage of Demonstration Data. Run the trained model strategy in a simulation environment, place the coal mine robot model in diverse scenarios, record the state, actions, and reward data in each scenario to generate a complete demonstration dataset; store the generated dataset in a structured manner in a database to ensure that the demonstration data can be directly accessed and used during subsequent imitation learning or training adjustment.

[0114] The database is InfluxDB, which is a time-series database that can handle a large amount of timestamp data and is suitable for the continuous data stream generated by the robot agent during training.

[0115] S3: Imitation Learning and Zero-Shot Transfer; its principle is as Figure 5 shown.

[0116] S301: Screening and Optimization of the Demonstration Dataset. In the demonstration dataset, perform weighted scoring on the task completion rate, path length, action smoothness, and task completion time, set a scoring threshold, and only retain the data above this threshold; use data augmentation methods to optimize the demonstration data, including adding different environmental conditions or position changes, to ensure that the data covers more scenarios and boundary conditions and improve the generalization ability of the model.

[0117] S302: Imitation learning model training. Using the screened and optimized demonstration dataset, that is, training the robot imitation learning model with the screened data in Unity 3d. Input the state-action pairs into the imitation learning model to enable it to learn how to imitate the operations in the demonstration data. Add domain randomization in the simulation training to make the model adapt to a wider range of physical and visual conditions, so as to improve the robustness of the imitation learning model in the actual environment. That is, based on the screened data and the coal mine robot behavior model, perform simulation training in Unity 3d with domain randomization added. Domain randomization includes randomly varying lighting, object color, and object friction coefficient.

[0118] S303: Model performance evaluation and optimization. Test the performance of the model in the simulation environment to ensure that it can accurately execute tasks under different conditions, and focus on evaluating the adaptability of the model in the domain randomization environment. According to the test results, make necessary fine-tuning to the model to ensure its stability in key task steps and changing environments, so as to further optimize the generalization effect of the model.

[0119] S304: Model zero-shot transfer and domain adaptation. Directly transfer the model trained through simulation training to the coal mine robot entity without additional real-world environment training, and directly test the initial performance of the model. That is, transfer the trained model to the coal mine robot to obtain the imitation learning model. During the transfer process, observe the performance of the model in the actual environment. If there are deviations, domain adaptation can be applied to make adaptive adjustments to the model.

[0120] S4: Generation of action pre-planning results. S401: Model deployment and interface configuration. Load and deploy the trained imitation learning model to the upper computer of the coal mine robot so that the model can run in the actual environment. Select an appropriate communication protocol to configure the communication interface between the robot upper computer and the imitation learning model to ensure that the robot can obtain the instructions generated by the model in real time. The communication protocol is ROS.

[0121] S402: Environmental perception and data transmission test. Start and initialize the coal mine robot sensor system to ensure that the robot can obtain environmental information in real time and provide accurate environmental data for subsequent action planning. Before the actual task, conduct a data transmission test to ensure that the model can receive the coal mine robot sensor data and output the corresponding action planning results.

[0122] The coal mine robot sensor system includes IMU, 3D lidar, depth camera, and visible light camera.

[0123] S403: Action Pre - planning Result Generation and Formatting. Based on the input job task objectives and the perception of the operation objects, tools, and environmental conditions by the coal mine robot, a complete movement path and robotic arm trajectory required for the coal mine robot to complete the job task are generated based on the imitation learning model, and a series of action data is output, including the pose information of the movement of the coal mine robot and the operation of the robotic arm.

[0124] S404: Action Pre - planning Result Formatting and Storage. Organize the action data into an array, where each node contains position coordinates, direction angles, timestamps, and other auxiliary information; save the action data in the form of an array as a JSON - format file and save it to a specific location for transmission to the mixed reality program in the subsequent action optimization stage.

[0125] As Figure 6 shown, in the action optimization stage of the coal mine robot based on mixed reality, the following steps are included.

[0126] S5: Presentation of Action Pre - planning Results, the principle of which is as Figure 7 shown.

[0127] S501: Basic Development of Mixed Reality Program. Import and configure the MRTK toolkit in the Unity 3d engine, select OpenXR as the XR development plugin, and add the Microsoft HoloLens feature group to it; import the coal mine robot model built in S201 into the scene and modify it into a coal mine robot AR model with a semi - transparent material to ensure the virtual - reality fusion effect of the mixed reality program.

[0128] S502: Function Integration of Mixed Reality Program. Add gesture interaction, gaze interaction, and voice interaction functions through the interaction configuration file of OpenXR for subsequent optimization of the action planning results by coal mine operators; import the Vuforia Engine into the mixed reality program and use the coal mine robot model as the Model Target; integrate the LineRenderer component into the coal mine robot AR model to support the drawing of action trajectories.

[0129] Gesture interaction detects gestures by using the HandJointUtils() and GestureRecognizer() classes.

[0130] Gaze interaction enables objects to respond to gaze through the Focusable component. The voice interaction uses the SpeechCommand() class to define voice commands that can be recognized.

[0131] S503: Mixed Reality Program Deployment and Spatial Alignment. Package the developed Mixed Reality program into a UWP platform application and deploy it to the Mixed Reality headset through Visual Studio for the coal mine operator to wear; when the physical entity of the coal mine robot is detected in the field of view of the Mixed Reality headset, load the Model Target to accurately align the pose of the coal mine robot model with the physical entity of the coal mine robot; the Mixed Reality headset is Microsoft HoloLens2.

[0132] S504: Pre-planned Result Rendering and Interaction Optimization. Load the action data in the JSON format file stored in S404 into the Mixed Reality program, use the Line Renderer component to draw the pre-planned results of the moving path and manipulator trajectory of the coal mine robot, and mark the interpolation points at key positions; add a ManipulationHandler component to each interpolation point object to enable it to be dragged through gesture interaction.

[0133] S6: Planning Result Adjustment Based on Multimodal Interaction; the principle is as Figure 8 shown.

[0134] S601: Observation and Analysis of Pre-planned Results. The coal mine operator observes the pre-planned results of the actions of the coal mine robot drawn through the Mixed Reality headset, judges whether the coal mine robot can accurately execute the operation tasks according to the planning results, and whether the moving path and manipulator interfere with the obstacles in the environment or the next behavior of the coal mine operator, and then analyzes whether it is necessary to adjust the pre-planned results of the actions and how to adjust them.

[0135] S602: Selection of Moving Waypoints Based on Gaze Interaction. The coal mine operator selects multiple key points in space as the waypoints for optimizing the moving path of the coal mine robot through gaze interaction; to ensure the accuracy of point selection and avoid misoperation, the system is designed so that when the coal mine operator's line of sight stays at a certain target point for more than 3 seconds, that point is regarded as confirmed selection.

[0136] S603: Adjustment of Manipulator Trajectory Points Based on Gesture Interaction. The coal mine operator moves the interpolation points marked on the pre-planned results of the manipulator trajectory through gesture interaction as the key points for optimizing the manipulator trajectory; when the coal mine operator adjusts the interpolation points, the working space of the manipulator is displayed in the Mixed Reality program to prevent the points from being adjusted to positions that the manipulator cannot reach.

[0137] S604: Confirmation of Action Optimization Based on Voice Interaction. After completing the selection of the moving waypoints of the coal mine robot and the adjustment of the manipulator trajectory points, the coal mine operator issues a command for re-planning through a voice instruction; to avoid misrecognition, it is necessary to confirm again through voice after issuing the command for re-planning.

[0138] S7: Spatiotemporal Deduction and Pre-execution; The principle is as Figure 9 shown.

[0139] S701: Optimization of Coal Mine Robot Movements. Record the pose information of the movement waypoints and the robotic arm trajectory points of the coal mine robot after being set and adjusted by the coal mine operator, and transmit it to the auxiliary computing device. Re-plan the trajectory through the preset motion planning algorithm deployed in the auxiliary computing device, and present the re-planned result in the mixed reality headset.

[0140] The auxiliary computing device is an embedded edge computing box placed in an explosion-proof enclosure and can be carried by the coal mine operator on the back; the preset motion planning algorithm is a combination of the A* algorithm and the RRT-Connect algorithm, where the A* algorithm is used for mobile path planning and the RRT-Connect algorithm is used for robotic arm trajectory planning.

[0141] Further, the communication protocol between the auxiliary computing device and the mixed reality headset is Socket.

[0142] S702: Spatiotemporal Deduction of the Optimization Result. Perform spatiotemporal deduction in the auxiliary computing device according to the optimization result of the coal mine robot movements, that is, gradually pre-execute the movement process of the coal mine robot body and the robotic arm in the spatiotemporal deduction scenario according to the time dimension.

[0143] S703: Push of the Spatiotemporal Deduction Result. Send the spatiotemporal deduction result to the mixed reality program in the mixed reality headset to generate a visual preview of the movement path of the coal mine robot and the robotic arm trajectory, and display the expected pose of the coal mine robot and the robotic arm at each moment.

[0144] S8: Generation and Sending of the Planning Result.

[0145] S801: Analysis and Re-optimization of the Optimization Result. The coal mine operator judges whether the optimization result of the movements is reasonable and feasible and whether there are actions that do not meet the expectations according to the spatiotemporal deduction result in the mixed reality headset; if the optimization result of the movements still needs to be adjusted, the coal mine operator can loop through S602 - S703 until it is ensured that the path fully meets the requirements before execution.

[0146] S802: Conversion of Path and Trajectory Data. Export the finally confirmed movement path of the coal mine robot and the robotic arm trajectory data from the mixed display program, that is, the motion planning result, including position coordinates, pose information, and key path points, and convert it into an instruction format that conforms to the parsing standard of the robot control system.

[0147] S803: Instruction Transmission and Execution. The path of the coal mine robot and the manipulator trajectory data, i.e., the action planning result after formatting, are transmitted to the robot control system through the configured communication interface and further parsed into the movement instructions of the robot body and the fine operation instructions of the manipulator to ensure that the robot accurately and efficiently completes the operation task according to the action planning result and task requirements.

[0148] Beneficial effects of this application:

[0149] (1) The action-level planning of the coal mine robot is split into two stages: pre-planning of the coal mine robot's actions and optimization of the coal mine robot's actions. In the first stage, the pre-planning of the coal mine robot's actions is achieved through reinforcement learning and imitation learning in the digital cousin environment. In the second stage, based on the first stage, mixed reality is used as an interface to enable further optimization of the pre-planning result by the coal mine operator. This not only makes full use of the global computing and macro-planning capabilities of machine intelligence but also reflects the flexible response and local optimization advantages of human intelligence.

[0150] (2) In the pre-planning stage of the coal mine robot's actions, the concept of digital cousin is introduced, and the data in the physical space is extended to the simulation environment for learning. Compared with digital twin, digital cousin does not pursue perfect reconstruction of the physical space in all minute details but focuses on retaining higher-level details such as the spatial relationship and semantic information between objects. This not only greatly reduces the cost of generating the virtual environment but also helps improve the cross-domain generalization ability of the learning strategy.

[0151] (3) Using the reinforcement learning algorithm to establish a pre-planning model of the coal mine robot in the digital cousin environment can regard the states of the coal mine robot body and the manipulator as a joint state space for joint action planning. While simplifying the action planning process, it enhances the task adaptability and generality of the pre-planning model. In addition, effective action planning can be achieved through interactive training, thus eliminating the complex kinematic analysis and modeling process and improving the planning efficiency.

[0152] (4) Using mixed reality technology as the interface and channel for the coal mine operator to deeply participate in the action planning of the coal mine robot, making full use of its virtual-real fusion characteristics and its good support for multi-modal interaction, enhances the active role of humans in the action planning of the coal mine robot, strengthens the human-machine dynamic interaction and feedback in human-machine collaborative planning, and enables the final action planning result of the coal mine robot to be more flexible and reliable in the complex and dynamic environment of the coal mine underground.

[0153] (5) In the action optimization of coal mine robots based on mixed reality, the auxiliary computing device carried by the coal mine operator can perform spatio-temporal deduction on the action planning results after optimization by the coal mine operator in the virtual scenario, and intuitively present the deduction results in the mixed reality headset, which conforms to the intuition of the coal mine operator, reduces the understanding cost of the coal mine operator, and enables it to better focus on the action optimization process; in addition, the deployment location of the auxiliary computing device can reduce communication latency and speed up the planning process.

[0154] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0155] In this article, specific examples are used to elaborate on the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A coal mine robot action level planning device based on digital cousin and MR, characterized in that: The coal mine robot action level planning device based on digital cousin and MR includes: An information data acquisition module, used to acquire scene information data; the scene information data includes a single-frame RGB image collected underground in a coal mine; A digital cousin scene information determination module is connected to the information data acquisition module and is used to determine the digital cousin scene information; the digital cousin scene information is determined based on the 3D asset library virtual model according to the scene information data; the 3D asset library virtual model is a preset model library covering operation objects, operation tools, terrain structures, mechanical equipment and auxiliary equipment, and is oriented to underground coal mine operation tasks; The action data determination module is connected to the digital cousin scene information determination module, and is used to input the digital cousin scene information into the imitation learning model to obtain action data; the action data includes: coal mine robot motion posture information and manipulator arm operation posture information; the imitation learning model is obtained by simulation training and zero-sample transfer based on the coal mine robot behavior model; the coal mine robot behavior model is obtained by reinforcement learning simulation training of the coal mine robot model; the coal mine robot model is a physical model constructed by 3D modeling tools based on the entity parameters of the coal mine robot; the entity parameters include: size ratio, structure data and appearance data; an action pre-planning module, connected to the action data determination module, and used to determine the action pre-planning result according to the action data; An adjustment module, connected to the action pre-planning module, for adjusting the action pre-planning result by adopting a multimodal interaction method to obtain an adjusted planning result; A space-time deduction module, connected to the adjustment module, is used to process the adjustment planning result using a preset action planning algorithm, and perform space-time deduction according to the time dimension to obtain a space-time deduction result; the space-time deduction result includes a moving path of the coal mine robot and a moving trajectory of the robotic arm; The action planning module is connected to the space-time deduction module and is used to determine the action planning result according to the space-time deduction result; the action planning result includes: position coordinates, posture information and path points.

2. The coal mine robot action level planning device based on digital cousin and MR according to claim 1 is characterized in that: The digital cousin scene information determination module includes: A semantic segmentation submodule, connected to the information data acquisition module, for performing semantic segmentation on the scene information data to obtain segmentation information data; A labeling submodule, connected to the semantic segmentation submodule, is used to label the segmentation information data with object labels using a large language model to obtain labeling information; A point cloud data determination submodule, connected to the label marking submodule, for determining point cloud data according to the marking information; the point cloud data includes three-dimensional position information; A semantic similarity determination submodule, connected to the point cloud data determination submodule, for determining semantic similarity based on the annotation information and the 3D asset library virtual model; A scene construction submodule, connected to the semantic similarity determination submodule, for constructing a digital cousin scene using a Unity 3D engine; The matching submodule is connected to the scene construction submodule and is used to perform digital cousin matching in the digital cousin scene according to the point cloud data and the semantic similarity to obtain digital cousin scene information.

3. A coal mine robot action level planning method based on digital cousin and MR, characterized in that: The method for action-level planning of a coal mine robot based on digital cousins ​​and MR is implemented by using the device for action-level planning of a coal mine robot based on digital cousins ​​and MR as described in any one of claims 1-2; The coal mine robot action level planning method based on digital cousin and MR includes: Acquire scene information data; the scene information data includes a single-frame RGB image collected in a coal mine; Determine digital cousin scene information; the digital cousin scene information is determined based on the 3D asset library virtual model according to the scene information data; the 3D asset library virtual model is a preset model library covering work objects, work tools, terrain structures, mechanical equipment and auxiliary equipment, and is oriented to underground coal mine work tasks; Input the digital cousin scene information into the imitation learning model to obtain action data; the action data includes: coal mine robot motion posture information and manipulator arm operation posture information; the imitation learning model is obtained by simulation training and zero-sample transfer based on the coal mine robot behavior model; the coal mine robot behavior model is obtained by reinforcement learning simulation training of the coal mine robot model; the coal mine robot model is a physical model constructed by 3D modeling tools based on the entity parameters of the coal mine robot; the entity parameters include: size ratio, structure data and appearance data; Determine the action pre-planning result according to the action data; The action pre-planning result is adjusted by adopting a multimodal interaction method to obtain an adjusted planning result; The adjustment planning result is processed by using a preset action planning algorithm, and time-space deduction is performed according to the time dimension to obtain a time-space deduction result; the time-space deduction result includes a moving path of the coal mine robot and a moving trajectory of the robotic arm; The action planning result is determined according to the spatiotemporal deduction result; the action planning result includes: position coordinates, posture information and path points.

4. The method for action-level planning of coal mine robots based on digital cousin and MR according to claim 3 is characterized in that: Determine the digital cousin scene information, including: Performing semantic segmentation on the scene information data to obtain segmentation information data; Using a large language model to label the segmentation information data with object labels to obtain labeling information; Determine point cloud data according to the annotation information; the point cloud data includes three-dimensional position information; Determining semantic similarity based on the annotation information and the 3D asset library virtual model; Use Unity 3D engine to build digital cousin scenes; Digital cousin matching is performed in the digital cousin scene according to the point cloud data and the semantic similarity to obtain digital cousin scene information.

5. The method for action-level planning of coal mine robots based on digital cousin and MR according to claim 3 is characterized in that: The action pre-planning result is adjusted by adopting a multimodal interaction method to obtain an adjusted planning result, which specifically includes: Performing interference analysis on the pre-planned action results through a mixed reality head display to determine whether to make adjustments and obtain a determination result; If the judgment result is yes, then using gaze interaction in multimodal interaction to select a waypoint in the space of the action pre-planning result as a path point for movement; Based on the gesture interaction in the multimodal interaction, the pre-planned action result is interpolated and marked, the key points of the robot arm trajectory tuning are determined, and the robot arm trajectory adjustment points are obtained; Based on the voice interaction in multimodal interaction, the action tuning and confirmation are performed according to the moving path points and the robot arm trajectory adjustment points to obtain the adjustment planning results.

6. The method for action-level planning of coal mine robots based on digital cousin and MR according to claim 3 is characterized in that: The adjustment planning result is processed using a preset action planning algorithm, and spatiotemporal deduction is performed according to the time dimension to obtain a spatiotemporal deduction result, which specifically includes: Using a preset action planning algorithm to optimize the adjustment planning result to obtain an action optimization result; According to the action optimization result, space-time deduction is performed according to the time dimension to obtain the space-time deduction result.

7. The method for action-level planning of coal mine robots based on digital cousin and MR according to claim 3 is characterized in that: Determining the action planning result according to the spatiotemporal deduction result specifically includes: Judging whether the spatiotemporal deduction result conforms to the expected action information; If yes, data conversion is performed on the time-space deduction result according to a preset standard instruction format to obtain conversion data; The action planning result is determined according to the conversion data.

8. The method for action-level planning of coal mine robots based on digital cousin and MR according to claim 3 is characterized in that: The method for determining the coal mine robot behavior model comprises: Acquire task data; the task data is the operation data of the coal mine robot when performing the task; Use the ML-Agents toolkit to build a reinforcement learning simulation environment; Decomposing the task data to obtain basic behavior data; the basic behavior data includes: operation data corresponding to moving, rotating, grabbing and placing; A reinforcement learning algorithm is adopted, based on a set reinforcement learning reward and punishment mechanism, and in the reinforcement learning simulation environment, multi-task optimization training is performed on the coal mine robot model according to the basic behavior data to obtain the coal mine robot behavior model.

9. The method for action-level planning of coal mine robots based on digital cousin and MR according to claim 3 is characterized in that: The method for determining the imitation learning model includes: Acquire a demonstration data set; the demonstration data set is the state, action and reward data of the coal mine robot behavior model in different scenarios; Screening the demonstration data set to obtain screening data; Based on the screened data and the coal mine robot behavior model, domain randomization is added in Unity 3D to perform simulation training, and the trained model is transferred to the coal mine robot to obtain the imitation learning model; the domain randomization includes: randomly changing lighting, object color and object friction coefficient.

10. The method for action-level planning of coal mine robots based on digital cousin and MR according to claim 3, characterized in that: The preset motion planning algorithms include: A* algorithm and RRT-Connect algorithm; among them, A* algorithm is used to determine the moving path of the coal mine robot, and RRT-Connect algorithm is used to determine the moving trajectory of the robotic arm.