Visual servo double-arm robot migration simulation learning method from simulation to reality

By building and optimizing data sets in virtual simulation scenarios and designing deep imitation learning networks, the challenge of training and adapting to two-arm robots in the real world is solved, efficient and safe visual servo control is achieved, and the robot's execution ability in complex tasks is improved.

CN120244972APending Publication Date: 2025-07-04HANGZHOU INTERNATIONAL INNOVATION INSTITUTE OF BEIHANG UNIVERSITY

Patent Information

Application Number
CN202510555270.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, two-arm robots have shortcomings in adapting to task diversity and complexity. The lack of labeled data affects learning effects, the risk of pure reality experiments is high, visual servo control technology faces perception problems in real-world applications, and it is difficult to choose suitable deep learning models and algorithms.

Method used

By building virtual simulation scenarios, collecting and optimizing robot operation data, designing deep imitation learning networks, combining variational autoencoder and Transformer architectures, end-to-end visual servo control is realized, using the simulation environment for training and migration to real scenes.

Benefits of technology

It reduces the time cost and risk of real-life scenario data acquisition and training, improves the exploration efficiency and generalization capabilities of robot learning networks, ensures the quality and safety of operations, and adapts to the needs of modern intelligent manufacturing and automated production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244972A_ABST
    Figure CN120244972A_ABST
Patent Text Reader

Abstract

The invention discloses a visual servo double-arm robot migration simulation learning method from simulation to reality, which comprises the following steps: S110, constructing a virtual simulation scene according to a real task scene, and establishing a relation between the virtual simulation scene and the real task scene; s120, the collected teaching mechanical arm and executing mechanical arm operation data are replayed and optimized in the virtual simulation scene, so that a data set used for final training is constructed; s130, a deep imitation learning network is designed and achieved, input of the deep imitation learning network comprises visual data obtained in the real task scene and state data of all joint motors of the teaching mechanical arm, and output of the deep imitation learning network is predicted states of all joint motors of the execution mechanical arm at the next moment; and S140, migrating the fully trained and converged deep imitation learning network from a virtual simulation scene to a double-arm robot in a real task scene. According to the invention, the exploration efficiency and generalization ability of the two-arm robot learning network are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robotics and relates to a method for migrating and imitating learning of a vision servo two-armed robot from simulation to reality. Background Art

[0002] In the rapid progress of technology, robotics has become an indispensable part of modern industry and service. Two-armed robots have received particular attention due to their ability to simulate the collaborative work of human arms. They show unique advantages when performing complex tasks, especially in fields such as manufacturing, medical surgery, service, and exploration. The research and development of two-armed robots are of great significance for promoting the process of automation and intelligence. They can improve production efficiency and service quality and are the key to achieving technological breakthroughs.

[0003] To overcome the deficiencies of traditional control algorithms in adapting to task diversity and complexity, the research of two-armed robots has begun to focus on deep learning and visual perception technologies. Deep learning technology optimizes the behavior of robots by learning from a large amount of data, enabling them to better understand and adapt to the working environment. At the same time, the development of visual perception technology, especially visual servo control technology, has become a key direction for improving the autonomy and adaptability of two-armed robots. It guides the actions of robots through visual information, realizes the accurate recognition and positioning of target objects, and provides a reliable perceptual basis for the grasping and cooperative operation of two-armed robots. Studying how to effectively train deep learning models in a simulation environment and successfully migrate them to the real world is crucial for the future development of two-armed robots.

[0004] In the process of selecting and training deep learning models, the field of two-armed robots faces multiple challenges. First, the lack of labeled data limits the training effect of deep learning models and affects the learning and adaptation ability of robots in complex environments. Second, the risk of pure real experiments is high, and frequent physical interactions may cause damage to the robots, increasing economic costs and safety risks. In addition, the application of visual servo control technology in the real world faces perceptual problems, and it is necessary to enable robots to accurately understand and process visual information to achieve precise control. At the same time, choosing the appropriate deep learning model and algorithm is also a problem because different tasks and environments may require different learning strategies and network structures. The accuracy and robustness of visual perception are also key issues, and the visual system may be affected by various factors such as environmental light changes and object surface characteristics. Solving these problems is crucial for developing a visual servo control technology for two-armed robots that can be trained in a simulation environment and seamlessly migrated to the real world. This can not only promote the progress of robotics but also provide strong technical support for future automation and intelligent applications. Summary of the Invention

[0005] The present invention provides a method for transfer imitation learning of a vision - servo dual - arm robot from simulation to reality, aiming to solve the problems in environmental perception, training data, and learning methods in related technologies. It is intended to use a simulation scenario to collect a large amount of associated data on visual perception and manipulator motion, drive preliminary model training based on deep imitation learning, and then optimize based on a small amount of real data, transfer the visual - servo motion planning ability of the model to a real manipulator, effectively accelerating the training time of the dual - arm robot control algorithm in a real scenario, and reducing the time cost and collision damage risk of data collection and training in the real scenario.

[0006] The technical solution adopted by the present invention is as follows: A method for transfer imitation learning of a vision - servo dual - arm robot from simulation to reality includes the following steps:

[0007] S110, construct a virtual simulation scenario according to the real task scenario, and establish the connection between the real task scenario and the virtual simulation scenario;

[0008] S120, replay and optimize the collected operation data of the dual - arm robot in the virtual simulation scenario to construct a dataset for final training;

[0009] S130, design and implement a deep imitation learning network. This deep imitation learning network adopts an end - to - end strategy and combines a variational auto - encoder and a Transformer architecture. The input of the deep imitation learning network includes visual data obtained from the real task scenario and the state data of each joint motor of the teaching manipulator, and the output is the predicted state data of each joint motor of the executing manipulator at the next moment;

[0010] S140, transfer the well - trained and converged deep imitation learning network from the virtual simulation scenario to the dual - arm robot in the real task scenario.

[0011] The beneficial effects of the present invention are as follows:

[0012] Compared with the prior art, a solution for visual - servo motion planning of a dual - arm robot is provided. This solution makes full use of the achievements of semi - physical simulation and deep - learning technologies, can accurately simulate and optimize the action execution of the robot in the actual scenario, achieve efficient and accurate task completion, ensure the quality and safety of operations, and meet the requirements of modern intelligent manufacturing and automated production.

[0013] By improving the existing training data collection process and verification platform of the dual - arm robot, the present invention significantly reduces the damage risk of the pure - physical experimental platform and the time consumption of data collection on the pure - simulation experimental platform, and at the same time ensures the action coherence, representativeness, and transferability of the dataset, thereby improving the exploration efficiency and generalization ability of the dual - arm robot learning network. Description of the Drawings

[0014] Figure 1 Schematic diagram of physical and simulation scenarios in a simulation-to-real visual servo dual-arm robot transfer imitation learning method according to the present invention;

[0015] Figure 2 Effect diagram of smooth data generation and playback in a simulation-to-real visual servo dual-arm robot transfer imitation learning method according to the present invention;

[0016] Figure 3 Deep imitation learning network architecture diagram adopted by the present invention;

[0017] Figure 4 Action block description diagram;

[0018] Figure 5 Time grouping description diagram;

[0019] Figure 6 Application diagram of the model controlling a dual-arm robot to perform a specific task in the real world. Specific implementation manners

[0020] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations.

[0021] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention claimed, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0022] The present invention provides a simulation-to-real visual servo dual-arm robot transfer imitation learning method, including the following steps:

[0023] S110, constructing a virtual simulation scenario according to the real task scenario, and establishing a connection between the real task scenario and the virtual simulation scenario. Figure 1 Schematic diagram of physical and simulation scenarios in a simulation-to-real visual servo dual-arm robot transfer imitation learning method according to the present invention. As Figure 1As shown in the figure, data is collected by the left - hand data acquisition area in cooperation with the middle virtual simulation scenario, providing training data for the application of the right - hand actual scenario model. This step is the basis for realizing the vision - servo motion planning of the dual - arm robot. Among them, the robotic arms in the middle virtual simulation scenario and the right - hand real - task scenario are called the executing robotic arms, and the robotic arms of the dual - arm robot in the left - hand real - task scenario are called the teaching robotic arms. Specifically, this step includes the following key links:

[0024] S110 - 1, Obtain the 3D models of all objects in the real - task scenario and configure them as the virtual simulation scenario. As Figure 1 shown, the core task is to accurately obtain the 3D models of all objects in the real - task scenario according to the right - hand real - task scenario and configure them as Figure 1 the virtual simulation environment shown in the middle.

[0025] In one implementation, the 3D models of all objects can be obtained through various methods, including: through 3D modeling automation technology, using algorithms and software to automatically construct 3D models based on text, photos, and video information of the object from multiple angles. 3D modeling automation technologies include structured light scanning, stereo vision, and Structure from Motion (SFM), etc. Among them, structured light scanning obtains the 3D shape of the object by projecting a specific light pattern and analyzing its deformation on the object surface; stereo vision reconstructs the 3D structure of the object by analyzing the differences between images taken by two or more cameras; while the Structure from Motion technology infers the 3D structure and motion information of the object by analyzing the image sequence of the object from different perspectives. These automated methods can significantly improve the efficiency and accuracy of model acquisition, especially suitable for large - scale scene modeling. The algorithms and software used can include ContextCapture, Agisoft.

[0026] In addition to 3D modeling automation technology, professional modeling software such as AutoCAD, 3ds Max, Maya, etc. can also be used to manually construct 3D models by professionals according to the blueprints or photos of the objects. This method is particularly important in scenarios that require high precision and complex details. For example, in the construction and engineering fields, engineers and designers use software such as AutoCAD to accurately construct 3D models of buildings according to detailed engineering drawings and design specifications; while in film and television animation and game development, artists use software such as 3ds Max and Maya to design and produce complex character, prop, and scene models according to creative and visual effect requirements. Manual modeling can give full play to people's creativity and professional skills to achieve highly customized model design.

[0027] In addition to the above, a physical engine such as Mujoco (Multi-Joint dynamics with Contact) can be used to build a virtual simulation scenario. Mujoco is a powerful physical simulation software that can accurately simulate complex physical interactions and dynamic behaviors, and is very suitable for robot simulation. It provides rich API interfaces, which can facilitate interaction and control with external programs.

[0028] These 3D models will be used to build an accurate virtual simulation environment to simulate the physical properties and appearance in the real world. In the simulation environment, physical properties of objects such as mass, density, elasticity, etc. need to be accurately simulated to ensure the reliability and effectiveness of the simulation results. For example, in robot manipulation simulation, the weight and friction coefficient of an object will affect the strategy and effect of the robot grasping and carrying the object; in vehicle collision simulation, the strength and stiffness of the vehicle body material determine the deformation and damage conditions during the collision. At the same time, the appearance details of objects such as color, texture, and gloss also need to be finely reproduced to enhance the realism and immersion of the simulation. These accurate 3D models provide a solid foundation for virtual simulation, enabling it to effectively simulate and predict various scenarios and events in the real world.

[0029] S110-2, establish the connection between the teaching robotic arm in the real task scenario and the executing robotic arm in the virtual simulation scenario. The core of this process lies in ensuring that the virtual simulation environment can accurately simulate the actions of the teaching robotic arm in the real world, and enabling the virtual executing robotic arm to precisely replicate these actions. The establishment of this connection is usually achieved by combining advanced simulation software and programming languages.

[0030] The simulation software can use a physical engine such as the above-mentioned Mujoco (Multi-Joint dynamics with Contact).

[0031] Step S110-2 may include: S110-2a, through the software development kit (SDK) of the robotic arm, real-time acquisition and synchronization of the motion data of the teaching robotic arm are achieved by Python. As a widely used programming language, Python has powerful functions and flexibility in the fields of robot simulation and control. The SDK usually provides rich functions and interfaces, which can obtain the state information of the robotic arm such as position, speed, and acceleration, as well as various commands and parameters for controlling the motion of the robotic arm. S110-2b, Python transfers the acquired motion data to the Mujoco virtual simulation scenario. In one implementation, during the simulation process, the Python script can continuously obtain the motion data of the teaching robotic arm from the SDK and then transfer this data to the Mujoco simulation environment. S110-2c, Mujoco drives the execution robotic arm in the virtual world to perform corresponding operations according to the received motion data. In this way, the execution robotic arm can accurately replicate the actions of the teaching robotic arm in the real world, achieving a close connection and synchronization between the two. By combining the methods of Mujoco and Python, not only can the motion control strategies of the teaching robotic arm be effectively simulated and verified, but also various experiments and tests can be carried out in the virtual environment, thereby reducing risks and costs in the real world. In addition, this method can also provide strong support and assistance for the optimal design, path planning, task execution, etc. of the teaching robotic arm.

[0032] Optionally, in addition to using a real-world robotic arm to implement the teaching task, a general and simple method can also be adopted, that is, using a handheld gripper in combination with the SLAM (Simultaneous Localization and Mapping) algorithm or an infrared motion capture environment to complete the teaching task. In this case, the handheld gripper in combination with the SLAM algorithm or the infrared motion capture environment can be regarded as the teaching robotic arm in the real task scenario mentioned in step S110-2. In this case, step S110-2 may include; S110-2a', tracking the pose of the handheld gripper through the SLAM algorithm, or by installing multiple infrared marker points on the handheld gripper and using an infrared camera to capture the motion trajectory of the infrared marker points, thereby achieving the tracking of the pose of the handheld gripper; S110-2b', transferring the captured pose of the handheld gripper to the robotic arm of the end effector version implemented in Mujoco, and accurately simulating the motion and control behavior of the robotic arm by Mujoco according to the pose; S110-2c', realizing the precise control and task execution of the robotic arm by operating the end effector in the virtual simulation scenario.

[0033] The core of this optional implementation lies in obtaining the precise end - effector pose and gripper opening / closing values, and directly operating the robotic arm with the end - controller version implemented in Mujoco based on these data. When using a handheld gripper, the operator can manually control the movement and opening / closing of the gripper to simulate the actions of the end - effector of the robotic arm. The SLAM algorithm plays an important role in this process. It can perform real - time localization and mapping in an unknown environment, providing accurate spatial position information for the handheld gripper. Through the SLAM algorithm, the system can accurately track the position and pose of the gripper, maintaining a high positioning accuracy even in complex or dynamically changing environments. In addition, the SLAM algorithm can be combined with machine vision technology to further improve the perception ability of the gripper pose, enabling it to better adapt to different task requirements and scenario changes. The infrared motion capture environment achieves precise tracking of the gripper pose by installing multiple infrared marker points on the gripper and using infrared cameras to capture the motion trajectories of these marker points. Infrared motion capture technology has advantages such as high precision, high sampling rate, and low latency, capable of capturing the subtle movements of the gripper in real - time and providing accurate input data for the simulation system. At the same time, the infrared motion capture environment can also be combined with other sensors and devices, such as gyroscopes, accelerometers, etc., to further enhance the perception and understanding of the gripper's motion state. After obtaining the precise end - effector pose and gripper opening / closing values, these data will be directly transmitted to the robotic arm with the end - controller version implemented in Mujoco. As a powerful physical simulation software, Mujoco can accurately simulate the motion and control behavior of the robotic arm based on these data. By operating the end - controller in the simulation environment, precise control and task execution of the robotic arm can be achieved, thereby verifying and optimizing control strategies and improving the performance and reliability of the robotic arm. In addition, this method can also provide important references and guidance for the design, planning, and application of robotic arms, promoting the further development of robotics technology.

[0034] S110-3. After establishing the connection between the teaching robotic arm in the real task scenario and the executing robotic arm in the virtual simulation scenario, the operator uses the teaching robotic arm to precisely control the executing robotic arm in the virtual simulation environment to perform a series of specific tasks. The key to this process lies in the real-time and accurate acquisition of the motion data of each joint of the teaching robotic arm, including parameters such as position and speed. These data are transmitted to the simulation system in real time through sensors and data acquisition systems and used as control signals to drive the corresponding joints of the executing robotic arm, enabling it to perform the same actions and tasks as the teaching robotic arm in the simulation scenario. In this way, the operator can intuitively operate the teaching robotic arm in the real world, while the simulation system can simulate the actions and behaviors of the executing robotic arm in real time. This not only provides the operator with an intuitive feedback and control interface but also enables the testing and verification of various task execution strategies and operation skills in a safe simulation environment. During the simulation process, the system records and stores the joint data of the executing robotic arm, which contains rich task execution information and motion characteristics. After subsequent data processing and optimization steps, these data will be used to train and optimize the artificial intelligence model. By accumulating a large amount of task execution data in the simulation environment, the generalization ability and execution efficiency of the artificial intelligence model in the real environment can be improved.

[0035] S120. Replay and optimize the operation data of the teaching robotic arm and the executing robotic arm collected in the virtual simulation scenario to construct the dataset for final training. This step is a key link to ensure the accuracy and effectiveness of the robotic arm motion data, including the following sub-steps:

[0036] S120-1. In the virtual simulation scenario, replay the operation data of the teaching robotic arm and the executing robotic arm collected. This process precisely simulates the behavior performance of the robot in the actual environment. Through data replay, every action detail of the robot during task execution can be carefully observed, including its response to visual information, the accuracy of motion planning, and the coherence and fluency of task execution. This observation not only covers the robot's performance at different task stages but also reveals its adaptability and flexibility under complex environmental conditions. By analyzing the replayed data, the performance and efficiency of the robot can be comprehensively evaluated, and potential problems and deficiencies can be identified. Ultimately, the data replay process helps determine the feasibility of the collected data for the task, that is, to evaluate whether the data can truly and accurately reflect the robot's operation ability and task execution effect in the actual environment. This is of great significance for subsequent model training and algorithm optimization, ensuring that the developed robotic arm system can efficiently and reliably complete various tasks in practical applications. Through this first step of data verification, a solid foundation can be laid for further optimization of the data.

[0037] S120-2. In the operation data of the teaching robotic arm and the executing robotic arm in the virtual simulation scenario, key actions are automatically selected through the key action selection algorithm. This algorithm automatically identifies the key action nodes during task execution by analyzing the changes in the motor angle values of each axis of the teaching robotic arm. Specifically, the algorithm continuously monitors the motion trajectory and joint changes of the teaching robotic arm during task execution, and accurately locates the action moments that have a decisive impact on task completion based on features such as the change rate, amplitude of the motor angle, and the collaborative relationship with other joints. By automatically selecting key actions through the key action selection algorithm, the efficiency and accuracy of subsequent manual selection are significantly improved. Instead of viewing the entire operation process frame by frame, one can directly focus on these key action nodes for more in-depth analysis and evaluation. This not only saves a large amount of human and time costs but also makes the selection of key actions more objective and scientific, laying a solid foundation for subsequent data annotation, model training, and algorithm optimization. In practical applications, the key action selection algorithm can effectively improve the intelligent level of robot task execution and promote the wide application of robot technology in complex task scenarios. The key action selection algorithm can be implemented using the idea of frame difference method for actions.

[0038] Specifically, the process of the key action selection algorithm is as follows: First, select the first action node in the operation process as the key node. Then, judge the change direction of the motor angle value of the current key node action through adjacent action nodes. If the angle value becomes smaller, use the sliding window method to find the local minimum point as the key point, and after finding the local minimum, find the local maximum point as the key point, loop until the last node, and take the last node as the key point; if the angle value becomes larger, use the sliding window method to find the local maximum point as the key point, and after finding the local maximum, find the local minimum point as the key point, loop until the last node, and take the last node as the key point. This method of dynamically selecting key points based on the change direction of the motor angle value can accurately capture the key action nodes in the operation process, thus providing strong support for subsequent analysis and applications.

[0039] S120-3, based on the key actions selected in step S120-2, continue to manually select key actions. This process can ensure that the core features of the task operation are retained. Although the key action selection algorithm has automatically identified a series of key action nodes, manual inspection is still essential to ensure that the selection of these nodes is completely reasonable and meets the actual requirements of the task. This step includes: examining each key action node and evaluating whether it accurately captures the core links of the task in combination with the specific objectives and operation processes of the task. By manually selecting key actions, the smoothed and optimized action data can be carefully adjusted and supplemented to ensure that while retaining the core features of the task operation, it can also meet the needs of subsequent model training and algorithm optimization. This step not only improves the quality and usability of the data but also provides strong guarantee for the robot to accurately and efficiently execute tasks in practical applications. The method of combining manual and algorithm for key action selection gives full play to their respective advantages, making the data verification process more comprehensive and reliable.

[0040] S120-4, use the updated key action nodes manually selected in S120-3 to generate smooth operation data for each joint motor of the execution manipulator. The effects of smooth data generation and playback are as Figure 2 shown. This process involves in-depth processing of the original data, which can include: by adopting data smoothing algorithms and filtering techniques, effectively eliminating excessive jitter noise and outliers in the key action nodes. These noises and outliers may come from factors such as sensor errors, environmental interference, or instability of the mechanical system. If not processed, they will directly affect the accuracy and efficiency of model training, and further affect the performance of the robot in practical applications. After smoothing the data, the motion trajectory of the execution manipulator in the simulation scenario becomes more smooth and natural, and the transition of joint actions is also more smooth and coherent. Playing back these smooth operation data in the simulation environment can visually observe the whole process of the execution manipulator performing tasks, including its response speed to task instructions, action precision, and adaptability in complex environments, etc. This provides an important basis for evaluating the task execution effect of the execution manipulator and helps to discover and solve potential problems in a timely manner. The data smoothing algorithm can adopt the moving average method or the exponential average method, and the filtering technique can adopt the median filtering technique.

[0041] Step S120 may include: To ensure that the smoothed data can complete the execution of the task completely and accurately, the three steps of S120-2 to S120-4 are looped. In each loop, a detailed inspection and optimization are carried out on the selection of key actions and the effect of data smoothing until the marked key actions can fully achieve the execution of the task. The inspection and optimization can be to further select key action nodes based on the key action nodes selected in the previous time, or to reselect key action nodes. This iterative process not only improves the quality and reliability of the data, but also lays a solid foundation for the smooth progress of subsequent model training, enabling the trained model to better adapt to the task requirements in the actual environment and improving the execution efficiency and accuracy of the robot in actual applications.

[0042] S120-5, Based on the visual data obtained from the virtual simulation scenario and the smoothed data generated based on the key action nodes selected in step S120-3, construct a data set. By organically combining the visual data obtained from the virtual simulation scenario and the smoothed data generated based on the key action nodes selected in step S120-3, a data set containing all the necessary information for the robot operation can be created, providing comprehensive and accurate data support for the training of the artificial intelligence model. The visual data provides rich environmental perception information for the data set, including the scene images observed by the robot during the task execution, the features such as the shape, color, and texture of the objects, and the relative position and interaction relationship between the robot and the environment. This information is of great significance for the model to understand the task scenario, identify the target object, and plan a reasonable operation strategy. The smoothed data generated by the key action nodes optimized by humans provides accurate action execution information for the data set, including the motion trajectories, speeds, and other parameters of each joint of the teaching manipulator, as well as the coordination relationship and timing information between the actions. These data can help the model accurately learn and master the action rules and skills of the robot during the task execution, improving the control accuracy and execution efficiency of the model. By constructing such a comprehensive data set, a solid data foundation can be provided for the training of the artificial intelligence model. During the training process, the model can make full use of various information in the data set for in-depth learning and optimization, thus possessing stronger perception ability, decision-making ability, and execution ability. This has an important promoting effect on improving the intelligent level of the robot in actual applications and enabling it to complete various complex tasks more flexibly and efficiently. In addition, high-quality data sets can also promote the innovation and development of robot technology, providing rich resources and inspiration for future research and applications.

[0043] S130, Design and implement a deep imitation learning network, which adopts an end-to-end strategy and combines a variational autoencoder (VAE) and a Transformer architecture to achieve precise control of the dual-arm robot, such as Figure 3As shown. The network input includes visual data obtained from the real task scenario and the joint motor state data of the teaching manipulator, and the output is the predicted joint motor state of the execution manipulator at the next moment. By calculating the loss between the actual value and the predicted value, the network is trained according to the loss and the number of training epochs until convergence. During the training process, the model is regularly verified in the virtual simulation scenario to obtain the task completion rate, which helps to judge whether the network converges. The visual data can be sourced from any type of visual sensor. Here, RGB images from multiple perspectives are taken as an example for illustration, which specifically includes the following sub-steps:

[0044] S130-1, with the Transformer encoder as the core, constructs a feature extraction framework as a variational autoencoder. The input data of the encoder covers the joint motor states of the dual-arm robot at a specific moment t, including joint position information, and a series of action sequences from moment t to t + k, such as Figure 4 As shown, these action sequences contain the detailed changes in the target joint positions of the robot over a period of time. In addition, a [CLS] token is specifically added to the input of the encoder. This special token plays a role in aggregating information in the sequence, helping the model better understand and integrate the context information of the entire action sequence. To enable the Transformer encoder to effectively process this input data, the joint position information and the action sequences are first mapped to a high-dimensional space through a linear embedding layer. This mapping process converts the original low-dimensional data into an embedding dimension representation suitable for the operation of the Transformer model, enabling the model to capture more complex and subtle feature relationships. In the embedding dimension, the Transformer encoder utilizes the advantages of its self-attention mechanism to deeply learn the complex patterns and implicit features of the joint movements of the robot when performing fine operation tasks. These features cover various aspects of information such as the coordination of joint movements, temporal relationships, and interactions with other joints. Finally, the encoder encodes the rich feature information extracted from the input data into a latent representation vector, which is then projected into a formal variable z through a linear layer. This formal variable z not only contains the static features of the robot's actions but also implies the dynamic change laws and potential execution strategies of the actions, providing key guiding information for subsequent decoding and action generation. Through this step, the deep imitation learning network can effectively capture and understand the action features of the robot in complex tasks, laying a solid foundation for achieving accurate action imitation and generation.

[0045] S130-2. The design of the VAE decoder integrates the formal variable z, the joint motor states of the dual-arm robot at time t, and RGB images from multiple perspectives as visual data inputs to achieve accurate prediction and generation of robot actions. First, the input RGB images undergo deep feature extraction through the ResNet18 network, converting the rich visual information in the images into feature maps. This process can capture key information such as object shapes, colors, and textures in the images, providing important visual bases for subsequent action prediction. To retain the spatial information in the feature maps, the ResNet18 network adopts the sine position embedding technique to incorporate the position information into the feature maps. In this way, the feature maps not only contain the content information of the images but also retain the spatial relationships between pixels, enabling the model to better understand the layout and interrelationships of objects in space. Next, the feature maps are flattened and mapped into a high-dimensional hidden dimension space. At the same time, the joint motor states of the dual-arm robot at time t are projected through a linear layer. Then, the mapping results of the feature maps are concatenated with the linear layer projection results of the formal variable z and the joint motor states as the input vector of the Transformer encoder. The formal variable z carries the style features and dynamic laws of the robot actions, while the joint motor states provide the specific execution information of the current action. These data together constitute a comprehensive high-dimensional feature representation. These embedded data are fed into the Transformer encoder, and the encoder deeply captures the deep information and complex relationships in the data through the self-attention mechanism. The encoder can understand the interactions and influences between different features and their importance in action generation. Next, the data output by the Transformer encoder is passed to the Transformer decoder. The decoder uses the cross-attention mechanism, combines the output of the encoder and the input data, predicts the action sequence of the robot in the next period of time, and combines these actions to weighted obtain the joint motor states at time t+1, as Figure 5 shown. This prediction process not only considers the current action state and visual information but also fully incorporates the action laws and strategies contained in the formal variable z, enabling the prediction results to accurately reflect the action execution of the robot at the next moment, providing strong support for realizing coherent and natural action imitation and generation. Through this step, the deep imitation learning network can effectively combine visual information, action states, and style features to achieve accurate prediction and control of robot actions.

[0046] S130 - 3, Design the termination conditions for model training to ensure that the deep imitation learning network can fully learn and master the complex mapping relationship from visual data and the joint motor state data of the teaching manipulator to the predicted joint motor state of the executing manipulator at the next moment. First, the end - to - end action generation loss function dropping to a preset target threshold is one of the key stopping conditions. This threshold represents the minimum standard that the model needs to achieve in terms of action prediction accuracy. When the loss function drops below this threshold, it indicates that the model can already generate action sequences that meet expectations relatively accurately, and the prediction error is controlled within an acceptable range, thus meeting the basic requirements of the model in action generation. Second, the total number of alternating iterative training steps exceeding a set value is also an important stopping condition. During the training process, the model needs to continuously optimize parameters and adjust the network structure through a large number of iterations to better adapt to the training data and task requirements. Setting a reasonable upper limit for the number of iteration steps can prevent the model from falling into the state of over - training, while ensuring that the model has enough time and opportunities to learn and absorb the key features and laws in the data. When the number of training steps reaches or exceeds this set value, it means that the model has been fully trained and has good generalization ability and stability. Finally, verifying the convergence of the model through the task completion rate in a pure simulation environment is another key stopping condition. The simulation environment provides a safe and controllable test platform, enabling a comprehensive evaluation and verification of the model's performance without increasing risks and costs in the real world. The task completion rate is an important indicator to measure the performance of the model in practical applications, which reflects the success rate and efficiency of the model when executing tasks. When the task completion rate of the model in the simulation environment reaches a relatively high level and tends to be stable, it indicates that the model already has the ability to execute tasks in the real environment, achieving the goal of deep imitation learning. Through this design of stopping conditions that comprehensively consider the loss function, the number of training steps, and the task completion rate, the network can effectively learn the complex mapping relationship from visual data and the joint state data of the teaching manipulator to the predicted joint state of the executing manipulator at the next moment. This enables the robot to perform fine - operation tasks even under low - precision hardware conditions, greatly expanding the application scope and potential of robotics technology.

[0047] The end - to - end action generation loss function is described below. The end - to - end action generation loss function includes action prediction loss, cross - entropy loss, and KL - divergence loss, and the formula is as follows:

[0048] ,

[0049] ,

[0050] ,

[0051] ,

[0052] The action prediction loss is used to measure the difference between the joint motor states predicted by the VAE decoder and the true joint motor states. Here, the mean squared error is selected. , where is the true joint motor state, is the predicted joint motor state, is the joint motor number, is the number of data points; the loss of the cross-attention mechanism is used to measure the rationality of the attention weights of the VAE decoder for the input data during the prediction process. Here, the cross-entropy loss is adopted to optimize the attention weights and ensure that the model can focus on key information more accurately. Among them, is the true attention weight, is the predicted attention weight, is the number of data points, is the number of attention heads, is the data point number, is the attention head number; the KL divergence loss is used to ensure that the distribution of the latent representation vectors is close to the preset prior distribution, which helps the model learn a smoother and more continuous latent space. Among them, and are the mean and standard deviation of the latent representation vectors respectively, is the dimension of the latent representation vectors, is the dimension number of the latent representation vectors. Finally, the comprehensive loss function combines all the above loss functions to form an end-to-end action generation loss function to ensure that the model performs well in multiple aspects. The weighted sum form is used to construct the comprehensive loss function , where and are hyperparameters used to balance the contributions of different loss functions.

[0053] S140. Transfer the well-trained and converged deep imitation learning network from the virtual simulation scenario to the dual-arm robot in the real task scenario. This step is a key link to verify the ability of the model to control the dual-arm robot to perform specific tasks in the real world and the correctness of the policy, as Figure 6 shown. The following are the detailed operation steps:

[0054] S140-1, The deep imitation learning network completed all necessary training and testing processes in the virtual simulation scenario and achieved the expected performance metrics. This means that the model needs to demonstrate excellent visual servo motion planning capabilities in the simulation environment, be able to accurately generate and execute motion planning instructions to guide the robotic arm to complete various complex operation tasks. In addition, the model also needs to conduct extensive simulation task tests in the simulation environment to verify its adaptability and stability in different task scenarios. These simulation tasks may include various types such as object grasping and placement, precision assembly, path planning and obstacle avoidance. Through its performance in these tasks, the comprehensive performance of the model can be comprehensively evaluated. During the process of determining the convergence of the model, it is required to closely monitor the decline trend of the model's loss function, the accuracy changes on the training and validation sets, and the actual performance of the model in the simulation tasks. When the loss function of the model steadily decreases and tends to be stable, the accuracy on the training and validation sets reaches a relatively high level and the gap is small, and at the same time the execution effect of the model in the simulation tasks is stable and reliable, and it can successfully complete various task objectives, it can be considered that the model has converged and has the basic conditions for migrating from the simulation environment to the actual application environment. At this time, the model already has strong generalization ability and robustness, and can still maintain good performance when facing complex situations and uncertainties in the actual environment, laying a solid foundation for the subsequent actual application migration.

[0055] S140-2, hardware deployment is a crucial step in successfully applying an artificial intelligence model to an actual dual-arm robot system. First of all, the trained deep imitation learning network needs to be integrated into the robot's control software to ensure that the deep imitation learning network can run and execute tasks in the actual hardware environment. This usually involves appropriate optimization and adjustment of the model to adapt to the hardware resource limitations and real-time requirements of the robot system. Next, the configuration of the hardware interface is crucial. It is necessary to ensure that the deep imitation learning network can accurately receive the actual visual data from the robot's sensors. This includes docking with visual sensors such as cameras and depth sensors, and transmitting the image data collected by the sensors to the deep imitation learning network for processing in real time. At the same time, necessary preprocessing of the sensor data, such as image denoising, size adjustment, color space conversion, etc., is also required to meet the requirements of the deep imitation learning network for the input data. In addition, the deep imitation learning network also needs to be able to send the generated control instructions to the robot's actuators, such as motor drivers and servo controllers, etc. This involves docking with the robot's motion control system to ensure that the control signals output by the model can be accurately parsed and executed. Detailed configuration of the format, transmission protocol, and timing of the control signals, etc., is required to achieve efficient cooperation between the model and the robot's actuators. During the hardware deployment process, it is also necessary to fully test the stability and reliability of the entire system. By conducting multiple trial runs and debugging in the actual environment, problems such as hardware interface problems, data transmission delays, and control instruction execution errors that may occur are discovered and solved in a timely manner to ensure that the artificial intelligence model can run stably in the actual dual-arm robot system and accurately execute various tasks. Through this series of hardware deployment work, the artificial intelligence model will be successfully migrated from the simulation environment to the actual application, providing strong support for the intelligence and automation of the robot in complex tasks.

[0056] S140-3, actual scenario testing, verifies the performance of the deep imitation learning network in real task scenarios. The dual-arm robot is deployed to a real working environment to perform a series of specific tasks similar to those in the simulation environment for a direct comparison and evaluation of the performance of the deep imitation learning network. These tasks typically cover complex operations such as object grasping, handling, and assembly, which are common key task types in the practical application of robots. In the grasping task, the robot needs to accurately identify the position and shape of the target object, plan an appropriate grasping path and posture, and then control the execution of the robotic arm and gripper to precisely grasp the object. The handling task requires the robot to quickly and smoothly move the object along a predetermined path to the designated position while keeping the object stable. The assembly task is even more complex. The robot not only needs to accurately identify and grasp each assembly part but also, according to the assembly sequence and requirements, accurately place the parts in their corresponding positions and complete operations such as fastening and connection. By performing these tasks in the actual scenario, the performance of the robot in the real environment can be observed and recorded, including the accuracy, efficiency, stability of task completion, and the ability to respond to emergencies. At the same time, a large amount of actual operation data can be collected for subsequent model optimization and improvement. The actual scenario testing can not only verify the training results of the model in the simulation environment but also reveal the possible problems and deficiencies of the model in practical applications, providing an important reference basis for the further improvement and enhancement of the model.

[0057] S140-4 uses a deep imitation learning network to generate visual servo motion planning instructions. The deep imitation learning network will generate precise visual servo motion planning instructions based on the visually feedback information obtained in real time. These instructions will guide the dual-arm robot on how to adjust its actions to adapt to environmental changes and task requirements, so as to accurately complete the specified tasks. The model needs to comprehensively consider factors such as the relative position between the robot and the target object, the posture changes of the object, and the obstacles in the surrounding environment, and generate reasonable motion trajectories and action sequences to ensure that the robot can execute tasks smoothly and accurately. The execution and monitoring phase follows immediately, where the generated motion planning instructions are executed in the actual scenario. It is required to monitor the actions of the robot in real time, closely observe parameters such as its motion trajectory, speed, acceleration, and its interaction with the environment. The purpose of monitoring is to ensure that the actions of the robot conform to the expected plan and can execute tasks within a safe range. Through monitoring, abnormal situations that may occur, such as excessive deviation of the robot's actions or the risk of collision with obstacles, can be detected and handled in a timely manner, thus ensuring the smooth progress of the task and the safe operation of the equipment. Performance verification is a key link in evaluating the actual application effect of the model. By comparing the performance of the robot in actual task execution with its performance in the simulation environment, the correctness and feasibility of the model are comprehensively verified. This includes evaluating the accuracy of task completion, such as the success rate of object grasping, the precision of assembly, etc.; evaluating the efficiency of task execution, such as the time required to complete the task, the motion speed of the robot, etc.; and evaluating the safety of the robot's operation, such as whether it avoids dangerous collisions with the environment and whether it maintains stability during operation. Through performance verification, the advantages and disadvantages of the model in actual applications can be quantified, providing an important basis for subsequent model optimization and improvement, and ensuring that the artificial intelligence model can reliably and efficiently guide the dual-arm robot to complete various tasks in the actual scenario.

[0058] S140-5. Verify the correctness of the deep imitation learning network by comparing the performance of the robot in actual task execution with that in the simulation environment. Analyze the performance of the model based on the task execution results in the actual scenario and make necessary adjustments and optimizations. It is required to deeply analyze the performance indicators of the model, such as the accuracy, efficiency, stability of task completion, and the adaptability to environmental changes, according to the detailed data collected and the observed performance during the actual task execution of the dual-arm robot. Identify the possible deficiencies of the model in actual applications by comparing the gap between the actual performance and the expected goals, such as the decrease in perception accuracy in complex environments, the lack of flexibility in motion planning, the slow response to unexpected situations, etc. Make necessary adjustments and optimizations for these discovered problems. This may include fine-tuning the model parameters, such as adjusting the weights, learning rate, regularization terms, etc. of the neural network, to better fit the actual data and task requirements and improve the prediction accuracy and generalization ability of the model. At the same time, it may also involve improving the robot control strategy, such as optimizing the motion planning algorithm to generate smoother and more efficient motion trajectories; enhancing the robot's perception fusion ability to more accurately integrate information from different sensors; or introducing more advanced decision-making mechanisms to enable the robot to better handle complex and changing task scenarios and unexpected situations. In addition, the architecture and functions of the model can be extended and upgraded according to the feedback and requirements in actual applications. For example, adding support for new types of tasks to enable the robot to perform more kinds of operations; or improving the real-time performance of the model to enable it to respond faster to environmental changes and task instructions and meet the requirements for the robot's response speed in actual applications. Through continuous result analysis and optimization, the artificial intelligence model will be continuously improved and enhanced, and finally realize the efficient, stable and intelligent operation of the dual-arm robot in the actual scenario.

[0059] In summary, first, use the semi-physical simulation platform to execute the test tasks of the visual servo dual-arm robot to obtain the visual and manipulator data of the scenario tasks; second, replay the collected operation data of the dual-arm robot in the pure simulation experimental environment, select key actions and optimize action smoothing to obtain accurate and feasible data of the dual-arm robot in the simulation experimental scenario, and construct the data set required for model training; then, based on the constructed data set, train the artificial intelligence model in the visual servo deep imitation learning network, and evaluate the training effect of the artificial intelligence model through the loss function during the model training process and the accuracy of the simulation scenario verification task to obtain an artificial intelligence model with good generalization ability; finally, transfer the trained artificial intelligence model to the dual-arm robot in the actual scenario to realize the generation and execution of visual servo motion planning instructions for specific tasks.

[0060] First, construct a virtual simulation scenario according to the real task scenario. Use a real-world teaching robotic arm or gripper to control the execution robotic arm in the virtual world to perform tasks, and obtain the data of each joint of the robotic arm and the data of the vision sensor in the virtual world during task execution. Optionally, when constructing the virtual simulation scenario according to the real task scenario, the methods for obtaining the 3D model include: algorithms and software for automatically constructing a 3D model based on texts, photos, and videos of multiple angles of real objects, and independently constructing a 3D model using modeling software.

[0061] Secondly, replay the collected dual-arm robot operation data in a pure simulation experiment environment to observe the task completion situation. Use the key action selection algorithm to automatically select key action nodes according to the changes in the motor angle values of each axis of the robotic arm. Then, replay the smoothed data generated by the updated key action nodes in the pure simulation experiment environment, manually select key action nodes, and then replay the smoothed data generated by the manually optimized key action nodes in the pure simulation experiment environment to ensure the rationality of the actions for the task. Continuously repeat the manual selection and data replay to finally ensure that the actions generated by the selected key action nodes can complete the task, obtain the accurate and feasible data of the dual-arm robot in the simulation experiment scenario, combine with the visual data obtained from the pure simulation experiment environment, and construct the dataset required for model training.

[0062] Next, according to the end-to-end strategy, a deep imitation learning network is designed by combining the variational autoencoder (VAE) and the Transformer architecture. The network takes visual data and the motor state data of each joint of the robotic arm as inputs, and outputs the predicted motor state of each joint of the robotic arm at the next moment. The loss between the actual value and the predicted value is calculated, and the network is trained based on the loss and the number of training epochs until the network converges. Among them, the model is verified in a pure simulation experiment environment every certain number of training epochs to obtain the task completion rate to assist in judging whether the network converges. Further, the design of the deep imitation learning includes: designing the VAE encoder and the VAE decoder with the VAE architecture by combining the characteristics of the Transformer architecture to obtain context features. The VAE encoder is mainly composed of a Transformer encoder, and its inputs include the motor state (joint position) of each joint of the dual-arm robot at time t and the action sequence (target joint position) from time t to time t + k. The input of the encoder also contains a special [CLS] token, which is used to aggregate sequence information. In addition, the joint positions and the action sequence are mapped to a high-dimensional space through a linear embedding layer so that the Transformer encoder can process them. In this way, the encoder can learn the complex patterns and implicit features of the joint movements of the robot when performing fine operation tasks, and encode these features into a style variable z of the robotic arm operation. The VAE decoder receives the style variable z, the motor state of each joint of the dual-arm robot at time t, and RGB images from multiple perspectives as inputs of visual data. These visual data can be sourced from any type of visual sensor, and RGB images from multiple perspectives are used as an example for illustration here. The image data is processed through the ResNet18 network and the sine position embedding to extract spatial features and map the joint positions to a high-dimensional embedding space. These embedded data, together with the style variable z, are fed into the Transformer encoder to capture deep information. Subsequently, the data is passed to the Transformer decoder, which uses the cross-attention mechanism to predict the motor state of each joint of the dual-arm robot at time t + 1. The final stopping condition for model training is that the end-to-end action generation loss function drops to a preset target threshold, or the total number of alternating iterative training steps exceeds a set value, and the convergence of the model is verified through the task completion rate in a pure simulation environment.

[0063] Finally, the converged artificial intelligence model is migrated to the dual-arm robot in the actual scenario to realize the generation and execution of visual servo motion planning instructions for specific tasks, and verify the correctness of the model migrated to the real scenario to control the dual-arm robot to execute tasks.

[0064] Although the invention has been described in terms of a limited number of embodiments, those skilled in the art, having the benefit of the foregoing description, will appreciate that other embodiments can be devised within the scope of the invention as described herein. Additionally, it should be noted that the language used in this specification has been principally selected for readability and instructional purposes and not to limit or circumscribe the inventive subject matter.

Claims

1. A method for transfer imitation learning of a vision servoing dual-arm robot from simulation to reality, characterized in that It includes the following steps: S110. Construct a virtual simulation scenario according to the real task scenario and establish the connection between the real task scenario and the virtual simulation scenario; S120. Replay and optimize the collected teaching manipulator and execution manipulator operation data in the virtual simulation scenario to construct the dataset finally used for training; S130. Design and implement a deep imitation learning network. The deep imitation learning network adopts an end-to-end strategy and combines a variational autoencoder and a Transformer architecture. The input of the deep imitation learning network includes the visual data obtained from the real task scenario and the motor state data of each joint of the teaching manipulator, and the output is the predicted motor state of each joint of the execution manipulator at the next moment; S140. Transfer the well-trained and converged deep imitation learning network from the virtual simulation scenario to the dual-arm robot in the real task scenario.

2. The method for migrating and imitating learning of a vision servo dual-arm robot from simulation to reality according to claim 1, wherein S110 includes: S110-1. Obtain the 3D models of all objects in the real task scenario and configure them as the virtual simulation scenario; S110-2. Establish the connection between the teaching manipulator in the real task scenario and the execution manipulator in the virtual simulation scenario; S110-3. After establishing the connection between the teaching manipulator in the real task scenario and the execution manipulator in the virtual simulation scenario, the operator uses the teaching manipulator to precisely control the execution manipulator in the virtual simulation environment to perform a series of specific tasks.

3. The method for migrating and imitating learning of a vision servo dual-arm robot from simulation to reality according to claim 2, wherein The method for obtaining the 3D models of all objects includes: through 3D modeling automation technology, using algorithms and software to automatically construct 3D models according to the text, photos, and video information of the object from multiple angles. The 3D modeling automation technology includes one of structured light scanning, stereo vision, and structure from motion; or manually constructing 3D models by professional personnel using professional modeling software according to the blueprints or photos of the object; or using the Mujoco physics engine to build the virtual simulation scenario.

4. The method for migrating and imitating learning of a vision servo dual-arm robot from simulation to reality according to claim 2, characterized in that Step S110-2 includes: S110-2a. Through the software development kit (SDK) of the manipulator, use Python to realize the real-time acquisition and synchronization of the motion data of the teaching manipulator in the real task scenario; S110-2b. Python transfers the obtained motion data to the Mujoco virtual simulation scenario; S110-2c. Mujoco drives the execution manipulator in the virtual world to perform corresponding operations according to the received motion data.

5. The method for transfer imitation learning of a vision servoing dual-arm robot from simulation to reality according to claim 1, characterized in that S120 includes the following steps: S120-1. In the virtual simulation scenario, replay the collected teaching manipulator and execution manipulator operation data; S120-2. In the operation data of the teaching manipulator and the execution manipulator in the real task scenario, automatically select key actions through the key action selection algorithm; S120-3. On the basis of the key actions selected in step S120-2, continue to manually select key actions; S120-4. Use the updated key action nodes manually selected in S120-3 to generate smooth operation data for each joint motor of the teaching manipulator; S120-5. Construct a data set based on the visual data obtained in the virtual simulation scenario and the smoothed data generated based on the key action nodes selected in step S120-3.

6. The method for migrating and imitating learning of a vision servo two-arm robot from simulation to reality according to claim 5, characterized in that, In S120-2, the process of the key action selection algorithm includes: First, select the first action node in the operation process as the key node. Then, judge the change direction of the motor angle value when the current key node acts through adjacent action nodes. If the angle value becomes smaller, use the sliding window method to find the local minimum point as the key point. After finding the local minimum, find the local maximum point as the key point, loop to the last node, and take the last node as the key point; If the angle value becomes larger, use the sliding window method to find the local maximum point as the key point. After finding the local maximum, find the local minimum point as the key point, loop to the last node, and take the last node as the key point.

7. The method for transfer imitation learning of a vision servoing dual-arm robot from simulation to reality according to claim 1, wherein S130 includes: S130-1. With the Transformer encoder as the core, construct a feature extraction framework as a variational autoencoder. The input data of the variational encoder covers the joint motor states of the dual-arm robot at a specific moment t, including joint position information, and a series of action sequences from moment t to t + k. The joint position information and action sequences are first mapped to a high-dimensional space through a linear embedding layer. This mapping process converts the original low-dimensional data into an embedding dimension representation suitable for the operation of the Transformer encoder. In the embedding dimension, the Transformer encoder learns the complex patterns and implicit features of the joint movements of the robot when performing fine operation tasks. Finally, the Transformer encoder encodes the rich feature information extracted from the input data into a latent representation vector, and then projects it through a linear layer into a formal variable z. S130-2. The design of the VAE decoder integrates the formal variable z, the joint motor states of the dual-arm robot at moment t, and RGB images from multiple perspectives as visual data inputs: First, the input RGB images are subjected to deep feature extraction through the ResNet18 network to convert the rich visual information in the images into feature maps; Then, the feature maps are flattened and mapped to a high-dimensional hidden dimension space. At the same time, the joint motor states of the dual-arm robot at moment t are projected through a linear layer. Then, the mapping result of the feature map is concatenated with the linear layer projection results of the formal variable z and the joint motor states as the input vector of the transformer encoder. Then, the data output by the Transformer encoder is passed to the Transformer decoder. The Transformer decoder uses the cross-attention mechanism to combine the output and input data of the Transformer encoder to predict the action sequence of the robot in the future for a period of time, and combines these actions to weight and obtain the joint motor states at moment t + 1. S130-3, Determine the termination conditions for the design model training to ensure that the deep imitation learning network can fully learn and master the mapping relationship from visual data and the joint motor state data of the teaching robotic arm to the predicted joint motor state of the execution robotic arm at the next moment.

8. The method for migrating and imitating learning of a vision servo two-arm robot from simulation to reality according to claim 7, characterized in that, The termination conditions for the model training in S130-3 include that the end-to-end action generation loss function drops to a preset target threshold. The end-to-end action generation loss function includes action prediction loss, cross-entropy loss, and KL divergence loss, and the formula is as follows: , , , , The action prediction loss is used to measure the difference between the joint motor states predicted by the VAE decoder and the true joint motor states. Here, the mean squared error is selected. , where is the true joint motor state, is the predicted joint motor state, is the joint motor number, is the number of data points; the loss of the cross-attention mechanism is used to measure the rationality of the attention weights of the VAE decoder for the input data during the prediction process. Here, the cross-entropy loss is used to optimize the attention weights to ensure that the model can focus on key information more accurately. Among them, is the true attention weight, is the predicted attention weight, is the number of data points, is the number of attention heads, is the data point number, is the attention head number; the KL divergence loss is used to ensure that the distribution of the latent representation vectors is close to the preset prior distribution, which helps the model learn a smoother and more continuous latent space. Among them, and are the mean and standard deviation of the latent representation vectors respectively, is the number of dimensions of the latent representation vectors, is the dimension number of the latent representation vectors. Finally, the comprehensive loss function combines all the above loss functions to form an end-to-end action generation loss function, and the weighted sum form is used to construct the comprehensive loss function , where and are hyperparameters used to balance the contributions of different loss functions.

9. The method for migrating and imitating learning of a vision servo dual-arm robot from simulation to reality according to claim 1, characterized in that, S140 includes: S140-1, The deep imitation learning network completes all necessary training and testing processes in the virtual simulation scenario and achieves the expected performance indicators. S140-2, Hardware deployment: Integrate the trained deep imitation learning network into the control software of the robot to ensure that the deep imitation learning network can run and execute tasks in the actual hardware environment. S140-3, Actual scenario testing to verify the performance of the deep imitation learning network in the real task scenario: Deploy the dual-arm robot to the real working environment and execute a series of specific tasks similar to those in the simulation environment to directly compare and evaluate the performance of the deep imitation learning network. S140-4, Use the deep imitation learning network to generate visual servo motion planning instructions. S140-5, Verify the correctness of the deep imitation learning network by comparing the performance of the robot in the actual task execution with its performance in the simulation environment.

10. The method for migrating and imitating learning of a vision servoing dual-arm robot from simulation to reality according to claim 9, characterized in that, S140-5 includes: Analyze the performance indicators of the deep imitation learning network, including the accuracy, efficiency, stability of task completion, and the adaptability to environmental changes, based on the detailed data collected and the observed performance of the dual-arm robot during the actual task execution. Identify the possible deficiencies of the deep imitation learning network in the actual application by comparing the gap between the actual performance and the expected goal, and make necessary adjustments and optimizations, including: Fine-tuning the parameters of the deep imitation learning network, including adjusting the weights, learning rate, and regularization terms of the deep imitation learning network.

Citation Information

Patent Citations

  • Implementation method of deep sea fine remote control task based on imitation learning

    CN113119132A

  • Robot reinforcement learning assembly method based on visual teaching and virtual-real migration

    CN115990891A

  • Method for grabbing target object by mechanical arm in dense scene

    CN116330283A

  • Teaching system, robot system, teaching method for robot, and teaching program for robot

    CN118369191A

  • Mobile grabbing robot imitation learning training method based on teleoperation

    CN118721181A

Cited By

  • Robot control method, electronic equipment, readable storage medium and program product

    CN120773066A

  • Method and device for constructing data set, equipment and storage medium

    CN121290363A

  • Method, apparatus, device and storage medium for constructing dataset

    CN121290363B

  • Data set adaptive optimization method based on artificial intelligence

    CN121706872A

  • Artificial intelligence based data set self-adaptive optimization method

    CN121706872B