Simulation learning mechanical arm control method based on edge TPU deployment

By deploying the imitation learning method of the multimodal perception system and Transformer architecture on the edge TPU, the existing imitation learning methods have solved the problems of high computing resources, insufficient hardware adaptability and insufficient visual perception in the robot arm control, and efficient and real-time robot arm control is achieved, improving the stability and intelligence level of task execution.

CN120269566APending Publication Date: 2025-07-08FUDAN UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510630653.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing imitation learning methods have problems in robotic arm control with high computing resource requirements, insufficient hardware adaptability, and insufficient visual perception capabilities, resulting in poor real-time performance and low energy efficiency ratio, making it difficult to operate efficiently on edge computing devices.

Method used

Using a multimodal sensing system based on edge TPU, combined with a three-way camera module and a Transformer architecture, the model is converted into a TPU-adapted BMODEL format through quantitative compression and operator fusion to realize multi-time step action sequence prediction, and drive the robot arm to perform tasks through inverse kinematics solution and dynamic control scheduling.

Benefits of technology

It significantly improves the real-time response speed and control accuracy of the robotic arm, can complete complex tasks more smoothly and naturally, reduces development and operation and maintenance costs, supports more complex action pattern recognition and timing analysis, and improves adaptability and intelligence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120269566A_ABST
    Figure CN120269566A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and robot control, in particular to an imitation learning mechanical arm control method based on edge TPU deployment, which comprises the following steps: (1) constructing a multi-mode sensing system which comprises a three-path camera module, a master-slave mechanical arm module and a TPU hardware development board; (2) constructing a training data set through demonstration action execution and synchronous data acquisition; (3) training an imitation learning strategy model based on a Transform architecture, and converting the model into a BMODEL format adapted to the TPU; (4) deploying an optimized BMODEL model on a TPU hardware platform, generating an action sequence through real-time reasoning, and driving a mechanical arm to execute a task in combination with inverse kinematics solution and dynamic control scheduling; and (5) realizing action error real-time calibration and system closed-loop control through a multi-view visual tracking and state feedback mechanism. According to the method, the compatibility and integration problems of the TPU chip and an existing mechanical arm control system can be effectively solved, and the response performance and the intelligent level of the mechanical arm autonomous control system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and robot control, and in particular, to an imitation learning manipulator control method based on edge TPU deployment. Background Art

[0002] Although imitation learning (IL) and reinforcement learning (RL) share a theoretical basis within the framework of the Markov decision process (MDP), there are significant differences in their implementation methods and application characteristics. IL learns the optimal policy by directly imitating expert demonstrations. This approach not only avoids the high sample complexity problem associated with RL's reliance on trial-and-error exploration but also bypasses the difficulty of manually designing reward functions. This process often requires a large amount of domain expertise in complex tasks such as robot manipulation, autonomous driving, or human motion imitation and is difficult to precisely quantify. In contrast, IL's learning paradigm is closer to the natural learning method of humans: for example, when people learn to swim or ride a bicycle, they usually master the skills by observing and imitating the actions of others rather than relying on abstract mathematical formulas or repeated trial-and-error. This intuitiveness gives IL unique advantages in interpretability and human-computer interaction, allowing users without a technical background to participate in model training by providing demonstration data, significantly reducing the development threshold of artificial intelligence systems. In addition, IL can effectively integrate human preference feedback (such as through inverse reinforcement learning or preference comparison), enabling the agent to not only copy expert behavior but also understand the underlying intentions behind the behavior, thus showing stronger adaptability than RL in fields such as medical diagnosis and educational robots that require high-precision human-machine collaboration. From the perspective of computational efficiency, IL can usually achieve satisfactory performance within a shorter training cycle, which is particularly important for practical application scenarios with limited computing resources. Although RL still has an irreplaceable value in exploring unknown environments and discovering strategies that exceed human experts, IL, with its data-driven characteristics, efficient utilization of prior knowledge, and more natural interaction methods, is becoming one of the important technical paths to achieve trustworthy artificial intelligence.

[0003] Existing imitation learning methods usually adopt single-step action prediction (such as behavior cloning), but such methods have the problem of compound error: that is, the prediction error will accumulate over time, causing the robot to deviate from the expected trajectory and ultimately resulting in task failure. To alleviate the compound error problem, the Stanford team proposed Action Chunking with Transformers (ACT). This method reduces the effective time range of the task by predicting action sequences for multiple future time steps (instead of single-step actions), thereby reducing the impact of error accumulation. However, ACT still faces challenges in terms of hardware adaptability. For example, the computational resource requirements are high because ACT relies on the Transformer architecture. When inferring results on general computing devices and running locally, the data transmission latency is high, making it difficult to meet the real-time control requirements. There is also insufficient hardware adaptability because existing ACT models are not optimized for dedicated AI acceleration chips (such as TPU), resulting in low energy efficiency when deployed on edge computing devices and making it difficult to operate efficiently on low-cost hardware.

[0004] In addition, most existing systems only use a single camera (usually a global observation camera or an end-effector camera of the robotic arm) to provide visual feedback, lacking multi-angle and omni-directional environmental perception ability. They cannot capture the global scene, side dynamics, and operation details simultaneously, resulting in visual blind spots and misjudgments easily occurring in complex tasks (such as curved surface fitting and cleaning). For example, in the task of cleaning red wine stains, it is difficult to accurately judge the contact state between the cleaning tool and the curved surface only relying on the top camera.

[0005] In summary, there are still some limitations in the existing robotic arm control methods. Traditional imitation learning methods based on expert demonstrations, such as behavior cloning, can effectively simplify the policy learning process, but are easily affected by the accumulation of compound errors, resulting in the task execution deviating from the target trajectory. The ACT method proposed by Stanford alleviates this problem to a certain extent by predicting multi-step action sequences. However, due to its dependence on the Transformer structure, the computational resource overhead is large and the inference latency is high, making it difficult to meet the real-time control requirements. At the same time, the current methods are not well optimized in terms of hardware adaptability and are difficult to be efficiently deployed on resource-constrained edge devices. In addition, most systems only rely on visual input from a single perspective, lacking multi-angle and omni-directional environmental perception ability, and are prone to visual blind spots when performing complex operations, affecting the accuracy and stability of the task.

[0006] Therefore, there is an urgent need for a new technical solution to solve the above technical problems. Summary of the Invention

[0007] The object of the present invention is to overcome the above-mentioned problems of the prior art, and provide an imitation learning manipulator control method based on edge TPU deployment, so as to solve the technical problems of poor real-time performance and low energy efficiency ratio caused by low computational efficiency and insufficient hardware adaptability of the existing imitation learning model in manipulator control.

[0008] The above object is achieved by the following technical solutions: An imitation learning manipulator control method based on edge TPU deployment, comprising: Step (1) Construct a multi-modal perception system, including three camera modules, a master-slave manipulator module and a TPU hardware development board; the three camera modules include a top-down camera, a side-view camera and an end-following camera, which are respectively used to collect multi-view image data of the global scene, lateral dynamics and operation details; Step (2) Through demonstration action execution and synchronous data collection, obtain a multi-view image stream and manipulator joint state data with time-aligned, and construct a training data set; Step (3) Train an imitation learning policy model based on the Transformer architecture, the model supports multi-time-step action sequence prediction, and convert the model into a BMODEL format adapted to TPU through quantization compression, operator fusion and graph optimization; Step (4) Deploy the optimized BMODEL model on the TPU hardware platform, generate an action sequence through real-time inference, and combine inverse kinematics solution and dynamic control scheduling to drive the manipulator to execute tasks; Step (5) Through a multi-view visual tracking and state feedback mechanism, realize real-time calibration of action errors and closed-loop control of the system.

[0009] Further, in step (1), the three camera modules unify the frame rate through a synchronous trigger mechanism, and are aligned with the manipulator state data through timestamps to form multi-modal time-series data pairs.

[0010] Further, in step (2), the data collection includes synchronously obtaining 720p resolution images and manipulator joint angles and end pose information at a frequency of 30Hz.

[0011] Further, the training of the imitation learning policy model in step (3) includes: Perform rotation, cropping, brightness perturbation and normalization processing on the multi-view image data; Adopt a mixed-precision training technique, combined with a gradient clipping mechanism to prevent gradient explosion; Design a multi-objective loss function, including position accuracy loss, action amplitude loss and time consistency loss.

[0012] Further, the model conversion and optimization steps in step (3) specifically include: Convert the ONNX model to an intermediate representation through the MLIR tool and perform structural reorganization; Execute INT8 quantization operations to adapt to the TPU low-precision computing architecture; Optimize the computational graph through operator fusion, redundant node pruning, and static constant folding.

[0013] Furthermore, the deployment on the TPU in step (4) includes: Build an input tensor encapsulation and output parsing module based on the Sophon inference engine; Implement the conversion from OpenCV images to the BMImage format through the BMCV tool; Dynamically schedule the inference process to ensure strict synchronization between the action output and the system operation cycle.

[0014] Furthermore, the control of the robotic arm in step (4) includes: Solve the end trajectory into six-dimensional joint angle commands through the inverse kinematics model; Generate underlying control signals by using smooth interpolation and energy machine angle quantization processing; Combine visual tracking and target detection mechanisms to calibrate the execution error in real time.

[0015] Furthermore, the three-way camera module is connected to the TPU hardware development board through a USB2.0 port for UVC communication; the master-slave robotic arm module is connected to the TPU hardware development board through a type-c port for serial communication; and a personal laptop connects to the TPU hardware development board through the SSH protocol.

[0016] Furthermore, the peak computing power of the TPU hardware development board is 32 TOPS, supports INT8 quantization inference, and completes model cross-platform conversion through the Docker containerization toolchain.

[0017] Furthermore, the multi-modal perception system supports dual-mode switching between automatic control and manual intervention and has the capabilities of anomaly detection and policy interruption recovery.

[0018] A method for controlling a robotic arm based on edge TPU deployment provided by the present invention can effectively solve the compatibility and integration problems between the TPU chip and the existing robotic arm control system, and significantly improve the response performance and intelligent level of the robotic arm autonomous control system. Specifically, it includes: The TPU chip adopted is designed specifically for accelerating deep learning inference calculations, has powerful parallel processing capabilities, and can efficiently run the imitation learning model. By integrating the TPU chip with the master-slave robotic arm system, the system can process the image information from multiple cameras and the robotic arm joint state data in real time, and then quickly make action decisions.

[0019] This method not only improves the real-time response speed of the system, but also significantly enhances the accuracy of control commands, enabling the robotic arm to complete complex tasks, such as item sorting, liquid cleaning, elevator button operation, etc., more smoothly and naturally. With the computing power of the TPU, the system can also support more complex action pattern recognition and timing analysis, providing a computing power foundation for the subsequent introduction of functions such as feedback learning and environmental perception, and further enhancing the adaptive ability and intelligent level of the robotic arm.

[0020] In terms of model deployment, the present invention adopts the Docker containerization method to convert the deep learning model into the BModel format supported by the TPU on the x86 platform, and uses the MLIR toolchain for model quantization, structure reorganization, and operator optimization, so as to maximize the inference speed while maintaining the accuracy. The system reconstructs the data flow and control logic architecture based on the SAIL library to improve the overall operation efficiency, and successfully realizes the efficient deployment of the imitation learning model on the TPU.

[0021] In addition, traditional robotic arm control systems rely on CPUs or GPUs to execute control logic. If edge deployment is required, communication problems will be involved, resulting in high inference latency, high energy consumption, etc., and it is difficult to meet the dual requirements of real-time performance and low power consumption. The efficient computing architecture of the TPU chip can complete high-density inference tasks at low power consumption, effectively extending the device life and simplifying the heat dissipation design. At the same time, the modular and complete toolchain characteristics of the TPU chip make it easy to integrate into the existing control system without large-scale changes to the hardware architecture, greatly reducing the development and operation and maintenance costs. Based on the imitation learning deployment solution of the present invention, the robotic arm control system can achieve intelligent control effects with high response, low power consumption, and easy deployment. Brief Description of the Drawings

[0022] Figure 1 It is a flowchart of a real-time evaluation method for a dedicated chip based on FPGA according to the present invention; Figure 2 It is a schematic diagram of the system hardware connection in a real-time evaluation method for a dedicated chip based on FPGA according to the present invention; Figure 3 It is a schematic diagram of the system module operation in a real-time evaluation method for a dedicated chip based on FPGA according to the present invention; Figure 4 It is a schematic diagram of the model inference software process in a real-time evaluation method for a dedicated chip based on FPGA according to the present invention; Figure 5 It is a schematic diagram of the model structure in a real-time evaluation method for a dedicated chip based on FPGA according to the present invention. Detailed Embodiments

[0023] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0024] As Figure 1 shown, this solution provides an imitation learning robotic arm control method based on edge TPU deployment. Through model optimization for the TPU architecture (such as operator adaptation, quantization compression, etc.), the ACT model can achieve millisecond-level low-latency and high-energy-efficiency real-time inference on edge computing devices, thereby improving the stability and reliability of imitation learning in robotic control tasks. This method includes: Step (1): Construct a multi-modal perception system, including three camera modules, a master-slave robotic arm module, and a TPU hardware development board; the three camera modules include a top-down camera, a side-view camera, and an end-following camera, which are respectively used to collect multi-view image data of the global scene, lateral dynamics, and operation details. Step (2): Through demonstration action execution and synchronous data collection, obtain a multi-view image stream and robotic arm joint state data with time-aligned, and construct a training data set. Step (3): Train an imitation learning policy model based on the Transformer architecture. The model supports multi-time-step action sequence prediction, and converts the model into a TPU-adapted BMODEL format through quantization compression, operator fusion, and graph optimization. Step (4): Deploy the optimized BMODEL model on the TPU hardware platform, generate an action sequence through real-time inference, and combine inverse kinematics solution and dynamic control scheduling to drive the robotic arm to execute tasks. Step (5): Through a multi-view visual tracking and state feedback mechanism, achieve real-time calibration of action errors and system closed-loop control.

[0025] As Figure 2 shown, in this embodiment, the hardware composition of the multi-modal perception system includes three cameras at different positions, which respectively capture the top view, side view, and robotic arm view of the test platform, a pair of master-slave robotic arms, and a TPU hardware development board. Among them, the three camera modules are connected to the TPU hardware development board through USB2.0 ports for UVC communication; the master-slave robotic arm module is connected to the TPU hardware development board through type-c ports for serial communication; the personal laptop is connected to the TPU hardware development board through the SSH protocol.

[0026] The control system architecture includes a perception feedback layer, a policy decision layer, and an execution control layer, where: The perception feedback layer is used to collect multi-view visual images and robotic arm status data during the operation of the robotic arm, and perform unified encoding and time synchronization processing; it includes a multi-view visual acquisition module and a robotic arm status reading module. The multi-view visual acquisition module consists of a top-down camera, a side-view camera, and a first-person camera at the end, which obtain operation environment images from different spatial positions respectively; the system unifies the frame rates of the three cameras through a synchronous triggering mechanism to ensure strict alignment of visual information and control data. The robotic arm status reading module reads the joint angles, end poses, and actuator statuses fed back by the robotic arm main control board in real time, and performs standardized encoding at a fixed frequency (such as 30Hz) to construct a state vector with time sequence alignment.

[0027] The policy decision layer includes an imitation learning policy model and an inference scheduling module, which are used to process the data collected by the perception feedback layer and output specific action instructions. The imitation learning policy model adopts a multi-modal fusion architecture based on Transformer, inputs the multi-view images and joint state information of the current frame, and outputs the prediction of the action trajectory in the next T steps. The inference scheduling module is based on the TPU platform, and performs high-speed inference operations by deploying the quantized and optimized BModel model; it supports the dynamic scheduling of static input formats, fixed frame rate processing, and forward inference processes to ensure strict synchronization of action output and system operation cycles. This layer supports the dual-mode switching between automatic control and manual intervention, and has the capabilities of anomaly detection and policy interruption recovery.

[0028] The execution control layer includes an action encoding module and a robotic arm driving module, which are responsible for converting the action vector output by the policy model into low-level control instructions and driving the robotic arm to complete actual operations. The action encoding module, based on the inverse kinematics model of the robotic arm, resolves the end trajectory output by the policy layer into six-dimensional joint angle instructions, and performs smooth interpolation and servo angle quantization processing. The robotic arm driving module establishes a communication link with the slave robotic arm control board through the serial port or CAN communication protocol to achieve real-time instruction issuance and execution feedback monitoring; at the same time, it records the execution status through a feedback mechanism for the upper-level inference module to dynamically adjust the policy. During the execution process, the system combines visual tracking and target detection mechanisms to calibrate the target position in real time to ensure the stability and accuracy of the grasping operation.

[0029] The specific deployment of the model in this solution includes the following steps: Step 1: Configuration of the TPU platform environment Complete the configuration of the operating environment on the TPU hardware development board, specifically including: loading necessary driver modules, building a Python operating environment, and installing the software development toolchain provided by the TPU official (including the BModel compilation tool, SAIL inference engine, etc.). This step ensures that the TPU chip has the ability to run deep learning models and process peripheral data, and establishes the foundation for the subsequent system operation.

[0030] Step 2: Multi-modal Device Connection and Initialization of Data Acquisition Channels Initialize the following modules through the driver interface at the software level: (1) Three-way camera module, including: a main perspective camera directly above (used to capture the overall task process), an end vision camera (close to the end of the robotic arm to collect operation details), and an environmental side view camera (capturing the dynamics on the side of the operation area); (2) Master-slave robotic arm module, which receives the control instructions of the master robotic arm and the execution status of the slave robotic arm respectively; (3) Data synchronization mechanism to ensure the alignment of image frames and robotic arm angle data in the time dimension.

[0031] This step provides a technical foundation for the high-quality acquisition of subsequent training data, ensuring the consistency and alignability of images and action signals.

[0032] Step 3: Demonstration Action Execution and Data Acquisition The operator manually controls the master robotic arm to execute specified tasks (such as tidying up desktop items, cleaning red wine stains on the tablecloth, accurately clicking the elevator button, etc.). During the task execution: synchronously collect image data from the three-way camera (resolution and frame rate set to 720p 30fps); synchronously collect status data such as the joint angles, speeds, and end poses of the slave robotic arm; collect continuous time-series samples for each task process, and divide the acquisition unit by about 10 seconds of data in a group.

[0033] Finally, collect no less than 50 groups of complete data sets, including high-resolution image streams and time-aligned control signal sequences, to construct a high-quality data set for training the imitation learning model.

[0034] Step 4: Model Training and ONNX Export Use the collected data to train the imitation learning model on a platform with high-performance GPU computing capabilities. The model structure is based on the Transformer as the core architecture, supports multi-time-step action prediction, and has the ability to mitigate compound errors. After training, export the model to a standard ONNX format file for subsequent cross-platform conversion.

[0035] Step 5: Model Compilation and TPU Format Conversion In the Docker container on the x86 platform, use the TPU toolchain to perform model conversion operations, specifically including: (1) Use the MLIR intermediate representation format to restructure the ONNX model; (2) Perform model quantization operations to compress the weights from FP32 to INT8 to adapt to the low-precision computing architecture of the TPU; (3)Perform graph optimization and operator fusion on the computational graph to reduce redundant computing nodes; (4)Finally, compile the optimized model into the BMODEL format supported by TPU to ensure efficient operation on TPU hardware.

[0036] Step 6: Deploy the inference model to TPU and reconstruct the control architecture Deploy the BMODEL model to the TPU hardware system, and reconstruct the control code architecture based on the TPU SAIL inference library. Implement the real-time processing flow of the whole process such as tensor encapsulation of images and joint states, model call interfaces, output parsing, and control instruction generation, ensuring that the system has end-to-end inference capabilities.

[0037] Step 7: Online inference and robotic arm control The system executes the following process during operation: (1)Continuously collect image stream data by three cameras, preprocess it through the local VPU (Vision Processing Unit), and encapsulate it as a tensor input; (2)Simultaneously obtain the real-time joint angle information from the robotic arm and construct it into a state input tensor; (3)Jointly input the image tensor and the joint tensor into the deployed BMODEL imitation learning model for inference, and output the action prediction sequence for several future time steps; (4)Convert the prediction result into a robotic arm control instruction through the action decoding module and send it to the slave robotic arm actuator to achieve low-latency and high-precision action execution.

[0038] This method makes full use of the high performance and low power consumption advantages of TPU in edge-side deployment, overcomes the problems such as high power consumption, slow response, and complex hardware integration existing in the GPU solution, realizes the efficient operation of the imitation learning model on the embedded platform, and has good scalability and engineering application prospects.

[0039] The inference of the model in this solution specifically includes the following steps: Step 1: System initialization and configuration loading At the system startup stage, the program first obtains the paths of various key configuration files by parsing command line parameters, including: model configuration file (defining the model structure and input / output parameters); model weight file (BModel format, suitable for the TPU platform); robot configuration file (including robotic arm communication protocol, control parameters, etc.).

[0040] The system creates a model loading object, which is responsible for parsing the model configuration file, extracting necessary parameters such as the image input resolution, the dimension of the state vector, the number of steps and dimensions of the action output, etc., loading the BModel model file, initializing the Sophon inference engine, and completing the registration of the model input and output nodes. In addition, the system also synchronously creates an image preprocessing module as part of the subsequent image input pipeline.

[0041] If the system is running in the real robot control mode (i.e., non-simulation mode), the program will further load the robot configuration file and construct a robot control object accordingly. This object completes the communication initialization with the physical robot hardware to ensure that subsequent motion control commands can be stably sent and status feedback can be received.

[0042] Step 2: Execution Process of the Main Control Loop After the system enters the main control loop, it first collects the current frame image data from the connected multi-channel cameras. The program sequentially calls the acquisition interfaces of each camera and divides the images according to the camera naming rules into: top-down view (such as the main view camera); side view (capturing lateral motion features); end view (capturing fine operation details).

[0043] The collected original OpenCV images are then converted to the BMImage format supported by the TPU platform through the BMCV (BitMain Computer Vision) tool. All images are organized into a dictionary structure classified by perspective for unified preprocessing and model input construction.

[0044] At the same time, the program reads the status data from the robot object in real time, including: the angles of each joint of the robotic arm; the position and orientation information of the end effector, etc. The above status data is encapsulated as a NumPy vector representation and is padded or truncated in dimension according to the model configuration to ensure strict consistency with the model input requirements.

[0045] Step 3: Data Preprocessing and Model Inference Each perspective image will be processed by the image preprocessing module in turn, mainly including: resizing to the model input resolution; color space conversion (such as BGR to RGB); data arrangement and format standardization (such as channel order, normalization, etc.).

[0046] The image and status data are then organized into a dictionary structure organized by the model input node name as the inference input and passed into the Sophon inference engine. The system writes each data item into the corresponding input tensor and performs a forward inference once.

[0047] The model output is an action sequence (usually a three-dimensional array in the form of [1, T, D], where T represents the predicted number of time steps and D represents the action dimension). The system extracts the action at the first time step of the sequence as the control signal to be executed during the current control cycle.

[0048] Step 4: Action Execution and Control Frequency Adjustment The system passes the current action vector into the robot drive module, converts it into low-level control instructions, and sends them to the robotic arm actuator to drive it to complete the corresponding operations. The main loop contains a frequency control mechanism, and the program will dynamically adjust according to the elapsed time of the current loop and the target frame rate.

[0049] Step 5: Exit Detection and Resource Release In the main loop, the system continuously monitors the exit condition. If a specific key input (such as key q or capturing a keyboard interrupt) is detected, the system will trigger a safe exit process: terminate the inference and control loops; safely disconnect the connection to the robot control interface; release the inference engine, image buffer, and system resources to complete an orderly shutdown of the system.

[0050] As a specific embodiment of this solution, a system for implementing the above method is provided.

[0051] As Figure 2 shown, the composition of the system hardware architecture includes: 1. Vision Acquisition System This system consists of three cameras with complementary perspectives, which are respectively deployed at different spatial positions to capture multi-angle image data during the task process: Global Overhead Camera: Fixedly installed above the workbench, using a 30fps, 720p resolution CMOS sensor, the image is transmitted through a USB2.0 interface, and the camera installation height is about 0.8m, providing a global perspective of the task and the initial layout information of the objects; Lateral View Camera: Horizontally installed 0.1m above the operation tabletop, with an image resolution of 720p and a frame rate of 30fps. The lens pitch angle is set to 30°, and it has the ability to manually adjust the zoom, specifically for capturing lateral operation dynamics; End-Effector Follow-up Camera: Fixed on the third joint of the slave robotic arm, and realizes follow-up shooting through a motion linkage mechanism, simulating the first-person perspective of humans, and is used to accurately record operation details and micro-action characteristics.

[0052] 2. Master-Slave Robotic Arm Control System A pair of 6-DOF serial manipulators with symmetric structures are used as the master-slave execution platforms of the system. High-precision harmonic reducers are configured for each joint, and the end effector is a standard two-finger gripper mechanism with a built-in magnetic encoder with a resolution of 4096 levels for closed-loop control. The master manipulator generates demonstration actions through manual operation, and the slave manipulator mimics the motion trajectory of the master manipulator in real time. The manipulator control unit establishes a data link with the edge computing platform through a Type-C interface to achieve low-latency communication and control.

[0053] 3. Edge Computing Platform (TPU) The computing platform is built based on a TPU development board with peak computing powers of 32 TOPS and 16 TFLOPS, supporting efficient integer operations, suitable for high-throughput inference tasks, and also suitable for models requiring floating-point precision. It is adapted to mainstream deep learning frameworks such as TensorFlow, PyTorch, Caffe, MXNet, ONNX, and PaddlePaddle, and supports parallel data input of 2 USB3.0 and 2 USB2.0. The TPU is integrated with the console through a customized base and can stably carry multi-source video processing and model inference tasks.

[0054] As Figure 3 shown, the implementation process of the system includes: 1. Multi-modal Data Synchronous Acquisition Stage When the system starts, the data acquisition clocks of the three cameras are coordinated through a timing synchronization mechanism to ensure that each image is acquired at the same timestamp. The image acquisition frequency is uniformly set to 30Hz; the manipulator status information (joint angles θ1 to θ6, resolution 0.01°) is synchronously sampled; each frame of data packet is attached with a globally unique ID, acquisition frame rate, manipulator configuration identifier, original image, and status data. Through the above mechanism, high-time-precision data alignment is achieved, providing a consistent sample set for subsequent training. During the data acquisition stage, the operator runs the acquisition code, and the development board automatically detects the connection of the cameras and the manipulator and initializes. The master manipulator sends the servo angle values of each joint to the development board in real time, and the development board then sends them to the slave manipulator at the same rate, so that the slave manipulator mimics the actions of the master manipulator at a frame rate of 50HZ. The operator controls the slave manipulator through the master manipulator to perform complex task demonstrations. Taking the "red wine stain cleaning" task as an example, the action process is divided into three sub-tasks: tool grasping, curved surface fitting and cleaning, and waste liquid treatment. A total of 50 rounds of complete demonstration data are collected in this stage to ensure coverage of different scenarios and boundary conditions in the operation. The collected data will be used as the data basis for training the imitation learning model.

[0055] 2. Model Training Stage, as Figure 5 shown In the model training stage, this system adopts an offline training method to train the policy model on a high-performance GPU platform.

[0056] First, training data pairs are constructed using the multi-modal dataset obtained in the data acquisition stage (including images from three perspective cameras and the corresponding robotic arm joint state vectors), where the image sequence and the robotic arm state sequence are jointly used as model inputs. To improve the generalization ability and robustness of the model, various forms of data augmentation are performed on the image data during training, including operations such as image rotation, central and random cropping, brightness and hue perturbation, Gaussian blur, and horizontal flipping, and normalization is uniformly performed before all image inputs (such as standard mean-variance normalization or normalization to the [0, 1] interval).

[0057] In addition, to improve training efficiency and reduce video memory occupancy, the model training introduces the mixed-precision training technique, that is, converting some computational processes to FP16 precision while retaining the key gradient accumulation process as FP32 precision. To ensure the stability of the training process, the system introduces a gradient clipping mechanism (Gradient Clipping), which limits the maximum norm of the gradients before each backpropagation update to prevent training instability or model divergence caused by instantaneous gradient explosion.

[0058] In terms of the loss function, a multi-objective loss structure is designed according to the task attributes, comprehensively considering the position accuracy loss, action amplitude loss, and time consistency loss, and the Adam or Ranger optimizer is used to update the weights. The entire training process supports breakpoint recovery and training log recording. After the final training is completed, the policy model is exported as a general ONNX model format, providing a standardized input model for subsequent deployment and compilation on the TPU platform. The training data in this stage covers 50 rounds of human demonstration tasks, covering multiple typical operation scenarios, ensuring that the model has strong generalization ability and action transfer ability.

[0059] 3. Model Conversion Stage In the model conversion stage, the system uses the Docker container environment under the x86 architecture and uses the toolchain provided by the TPU official to optimize and convert the trained ONNX format policy model to adapt to the inference execution requirements of the TPU hardware platform based on the BM1684X chip.

[0060] This process first uses the MLIR tool to restructure the ONNX model, converting complex or redundant high-level representations in the model (such as loops, residual modules, etc.) into an intermediate-level representation suitable for static inference graphs, and disassembling them into basic low-order operator structures, laying the foundation for subsequent quantization and graph optimization. On this basis, the system performs model quantization operations, that is, compressing the original FP32 floating-point weights and activation tensors to INT8 integer precision, significantly reducing the model size while maximizing the adaptation to the low-precision computing architecture of the TPU chip, improving the inference efficiency and reducing power consumption.

[0061] To further optimize the model inference performance, the system also performs graph-level optimization on the entire computational graph, including common operator fusion (such as fusing Conv+BN+ReLU into a single operator), subgraph replacement, static constant folding, redundant node pruning, etc. technologies, effectively reducing unnecessary computational paths and intermediate storage, and reducing the memory bandwidth pressure.

[0062] After all optimization steps are completed, the system compiles the model into a BMODEL format file natively supported by the TPU. This format not only contains the quantized parameters and graph structure information, but also integrates the execution plan generated for the BM1684X hardware instruction set, thus ensuring the maximum throughput rate and the lowest inference latency on the TPU hardware. The entire conversion process supports batch processing and multi-model parallel conversion, has good scalability and repeatability, and provides node mapping relationships, quantization error evaluation, and performance analysis reports during the conversion process through log output, facilitating subsequent debugging and deployment verification. This stage lays a key foundation for the imitation learning control system to achieve real-time, low-power, and high-efficiency operation on edge devices.

[0063] 4. TPU inference deployment stage, as Figure 4 shown In the model deployment stage, the system first completes the running environment configuration on the target TPU platform (based on the BM1684X chip), including driver loading, Sophon Runtime environment initialization, installation and version verification of required dependency libraries (such as BMCV, SAIL, bmrt, etc.), to ensure compatibility and stability among various components. After completing the system initialization, the program parses the model configuration file to obtain the structural parameters of the policy model (such as image input size, state input dimension, action output dimension, time step length, etc.), and prepares the input and output tensor structures accordingly. Subsequently, the system loads the BModel model file that has completed quantization and structural optimization, initializes the Sophon inference engine (BMRuntime), and registers all input and output nodes, establishing the mapping relationship between the input tensor and the predefined input of the model, ensuring that the inference engine can correctly read external data and output valid inference results. During this process, the system also creates an image preprocessing module.

[0064] Actual Effect Evaluation In the experimental verification stage, the system was field-tested in multiple representative complex robot task scenarios to comprehensively evaluate the actual performance and application feasibility of the proposed multi-view imitation learning control method. The experimental tasks included three types of typical operations: The first type was the item sorting task, which simulated industrial handling or household storage behaviors. It required the robotic arm to complete the recognition, grasping, and classified placement of scattered items on the table under multi-view visual guidance. The system executed 50 rounds of operations in total, and successfully completed 41 tasks. The overall execution success rate reached 82%, indicating that the system had good perception and control coordination capabilities in the face of complex layouts such as occlusion and overlap. The second type was the red wine stain cleaning task, which consisted of multiple sub-action steps such as tool grasping, wiping in close contact, and waste liquid disposal. The system needed to accurately reproduce the cleaning path and force control strategy in the human demonstration. In the test, 70 rounds of tasks were executed, and about 89.3% were successfully completed, demonstrating the feasibility and robustness of the imitation learning strategy in flexible operations. The third type was the elevator button operation task, which tested the system's performance in high-precision spatial positioning and fast action execution. The system continuously executed 50 rounds of operations under different heights and button arrangements. The control system response delay was strictly controlled within 120 ms, and there were no mis-touch or missed-touch situations, indicating good operation stability.

[0065] All experiments were carried out in the actual physical environment. The cameras, robots, and electric control systems involved were in the actual deployed state, which fully tested the adaptability of the algorithm to perception noise, light changes, and object dynamic interference. The data records included the success and failure judgments, control response time delays, key frame images, joint angle trajectories, etc. of each round of tasks, which were subsequently used for system performance analysis and improvement. The experimental results as a whole showed that: The imitation learning control framework proposed in the present invention, which integrates multi-view visual input and TPU-accelerated inference, demonstrated significant advantages in multi-source perception fusion, action inference accuracy, and system real-time performance, verifying its practicability, deployability, and high execution efficiency in real robot control tasks, and providing technical support for intelligent robot operating systems facing complex environments.

[0066] The above description is only for illustrating the implementation manners of the present invention and is not intended to limit the present invention. For those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A control method for an imitation learning manipulator based on edge TPU deployment, characterized in that Including: Step (1): Construct a multi-modal perception system, including a three-way camera module, a master-slave robotic arm module, and a TPU hardware development board; the three-way camera module includes a top-down camera, a side-view camera, and an end-following camera, which are respectively used to collect multi-view image data of the global scene, horizontal dynamics, and operation details; Step (2): Through demonstration action execution and synchronous data collection, obtain temporally aligned multi-view image streams and robotic arm joint state data, and construct a training data set; Step (3): Train an imitation learning policy model based on the Transformer architecture. The model supports multi-time-step action sequence prediction, and converts the model into a BMODEL format adapted to the TPU through quantization compression, operator fusion, and graph optimization; Step (4): Deploy the optimized BMODEL model on the TPU hardware platform, generate an action sequence through real-time inference, and drive the robotic arm to execute tasks by combining inverse kinematics solution and dynamic control scheduling; Step (5): Through a multi-view visual tracking and state feedback mechanism, realize real-time calibration of action errors and system closed-loop control.

2. The imitation learning robotic arm control method based on edge TPU deployment according to claim 1, wherein In step (1), the three-way camera module unifies the frame rate through a synchronous trigger mechanism and aligns with the robotic arm state data through timestamps to form multi-modal temporal data pairs.

3. The imitation learning robotic arm control method based on edge TPU deployment according to claim 1, wherein, The data collection in step (2) includes synchronously obtaining 720p resolution images and robotic arm joint angles and end pose information at a frequency of 30 Hz.

4. A method for controlling an imitation learning manipulator based on edge TPU deployment according to claim 1, wherein The training of the imitation learning policy model in step (3) includes: Performing rotation, cropping, brightness perturbation, and normalization processing on the multi-view image data; Adopting a mixed-precision training technique and combining a gradient clipping mechanism to prevent gradient explosion; Designing a multi-objective loss function, including position accuracy loss, action amplitude loss, and time consistency loss.

5. The imitation learning manipulator control method based on edge TPU deployment according to claim 4, characterized in that, The model conversion and optimization steps in step (3) specifically include: Converting the ONNX model into an intermediate representation through the MLIR tool and performing structural reorganization; Performing INT8 quantization operations to adapt to the TPU low-precision computing architecture; Optimizing the computational graph through operator fusion, redundant node pruning, and static constant folding.

6. The imitation learning robotic arm control method based on edge TPU deployment according to claim 1, wherein The deployment on the TPU in step (4) includes: Constructing an input tensor encapsulation and output parsing module based on the Sophon inference engine; Implementing the conversion of OpenCV images to the BMImage format through the BMCV tool; Dynamically scheduling the inference process to ensure that the action output is strictly synchronized with the system operation cycle.

7. A method for controlling an imitation learning manipulator based on edge TPU deployment according to claim 6, characterized in that The control of the robotic arm in step (4) includes: Solving the end trajectory into six-dimensional joint angle commands through an inverse kinematics model; Adopting smooth interpolation and energy-machine angle quantization processing to generate low-level control signals; Combining visual tracking and target detection mechanisms to calibrate execution errors in real time.

8. The imitation learning manipulator control method based on edge TPU deployment according to claim 1, wherein, The three-way camera module is connected to the TPU hardware development board through a USB2.0 port for UVC communication; the master-slave robotic arm module is connected to the TPU hardware development board through a type-c port for serial communication; a personal laptop is connected to the TPU hardware development board through the SSH protocol.

9. The imitation learning robotic arm control method based on edge TPU deployment according to claim 1 or 8, characterized in that, The peak computing power of the TPU hardware development board is 32 TOPS, supports INT8 quantization inference, and completes model cross-platform conversion through the Docker containerization toolchain.

10. A method for controlling an imitation learning robotic arm based on edge TPU deployment according to claim 1 or 8, characterized in that, The multimodal perception system supports dual-mode switching between automatic control and manual intervention, and has the capabilities of anomaly detection and policy interruption recovery.

Citation Information

Cited By

  • Mechanical arm data collecting and labeling method and system

    CN120461485A