Automatic aeronautical part sorting method and system based on analog simulation and real teaching combined training

An automated sorting method for aerospace parts, developed through a combination of simulation and real-world teaching, integrates ACT technology and closed-loop control. This method addresses the issues of low sorting accuracy and difficulty in migrating simulation strategies in existing technologies, achieving high-precision and robust automated sorting, and significantly improving production efficiency and equipment reliability.

CN122008226APending Publication Date: 2026-05-12DONGHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DONGHUA UNIV
Filing Date
2026-03-20
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing automated sorting technologies for aerospace parts cannot effectively combine simulation and real teaching data for model training, resulting in insufficient automated sorting accuracy, inability to adapt to changes in lighting and partial occlusion, and difficulty in transferring simulation strategies to real robot systems, leading to low sorting accuracy and poor robustness.

Method used

A combined training method of simulation and real teaching is adopted. By constructing a module for simulation data construction and generation, a real data acquisition module, a data preprocessing module, a VLA model joint training module, and a task execution optimization module, and combining Action-Chunking-Transformer (ACT) technology with closed-loop control, end-to-end learning of visual information, task text, and action sequences is achieved.

Benefits of technology

It improves the accuracy and robustness of aerospace parts sorting, increases task success rate by 15%, execution accuracy by 20%, reduces equipment failure rate by 25%, increases production efficiency by 40%, and saves 30%-50% of labor costs annually.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122008226A_ABST
    Figure CN122008226A_ABST
Patent Text Reader

Abstract

The invention discloses an aeronautical part automatic sorting method and system based on analog simulation and real teaching combined training. A system framework composed of simulation data construction and generation, real data collection, data preprocessing, VLA model combined training and task execution optimization is constructed; firstly, a simulation scene is built on a MuJoCo platform to generate large-scale multi-modal training data, real teaching data are collected through a PIKA terminal, after unified and standardized preprocessing is conducted, a VLA model fusing DinoV2, SigLIP and Llama2 7B is trained in a simulation pre-training and real data fine tuning mode, and finally robot action execution is optimized by combining an ACT technology with closed-loop control. According to the method, end-to-end learning of aeronautical part recognition, grabbing planning and sorting action generation is achieved, the sorting precision, stability and generalization ability are improved, and the method can be widely applied to scenes such as aeronautical assembly, aeronautical material management and aeronautical automatic storage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent manufacturing and robot operation technology in aviation. Specifically, it relates to an automatic sorting method and system for aviation parts based on a combination of simulation and real teaching training. It can be widely used in aviation manufacturing-related scenarios such as aviation assembly, aviation material management, and aviation automated warehousing. Background Technology

[0002] In the aerospace manufacturing industry, parts are diverse, have high similarity in appearance, and are arranged in complex ways. Manual sorting suffers from high costs, low efficiency, and large errors, making automated sorting an industry trend. Currently, existing automated sorting technologies for aerospace parts fall into two categories: one relies on traditional image recognition or simple feature matching methods, which can only achieve basic part classification and cannot combine task semantics, visual information, and operational history for overall reasoning. It lacks end-to-end learning models and the ability to perform continuous actions such as grasping, moving, and alignment. The other category is based on Vision-Language-Motion (VLA) intelligent operation technology. Although it attempts to integrate multimodal information, it has failed to effectively combine simulation and real teaching data for model training. Furthermore, it has weak adaptability to changes in lighting, partial occlusion, and random part placement. Simulated action strategies are also difficult to transfer to real robot systems, resulting in insufficient automated sorting accuracy, process instability, and an inability to support the complex sorting tasks in aerospace manufacturing.

[0003] Existing patents related to automatic sorting of aviation parts or intelligent operation based on VLA include: "An end-to-end system of causal control VLA integrating perception and world model" by Guangzhou Intelligent Technology Co., Ltd. (Patent No.: CN202511171514.8), and "An intelligent sorting device for aviation parts before inspection" applied for by Shenyang Aircraft Industry (Group) Co., Ltd. (Application No.: CN202511156191.5).

[0004] The method disclosed in patent number CN202511171514.8 is as follows: S1. Acquire images, semantic information, and sensor data of the environment through multimodal sensors, and input them into the environment perception module to form a unified environment representation. S2. Construct an environmental dynamics model using a world model training module to predict future states under different conditions. S3. Reason about the environmental state based on a causal modeling module to obtain possible behavioral decisions. S4. Generate multiple candidate trajectories through a trajectory generation module. S5. Perform security verification on the candidate trajectories using a formal verification module, and select the optimal trajectory as the control output.

[0005] The method disclosed in application number CN202511156191.5 is as follows: S1. Parts are conveyed to the sorting station via a conveyor belt, and a camera captures images of the parts. S2. The data acquisition module transmits the images to the recognition unit, and the parts are classified using traditional image processing / recognition algorithms. S3. The control module sends instructions to the sorting mechanism based on the recognition results, controlling push rods or diversion mechanisms to push the parts to the corresponding positions. S4. The sorted parts enter subsequent inspection or assembly processes, achieving basic automated operation.

[0006] Existing patents show that the causal control VLA end-to-end system, which integrates perception and world models, only achieves environmental perception and trajectory planning. It is not optimized for the specific characteristics of aerospace parts sorting scenarios, nor does it incorporate simulation and real data training. Furthermore, intelligent sorting devices before aerospace parts inspection use traditional image processing algorithms, which can only perform simple parts sorting and lack the ability to execute continuous actions and adapt to complex scenarios. Therefore, there is an urgent need for an automated sorting method that can integrate simulation and real data, possess multimodal end-to-end learning capabilities, and adapt to the complex sorting scenarios of aerospace parts. Summary of the Invention

[0007] This invention addresses the shortcomings of existing technologies by providing an automated sorting method and system for aerospace parts based on joint training of simulation and real-world teaching. It solves the problems of scarce teaching data, insufficient sorting accuracy, difficulty in executing complex tasks, and difficulty in transferring simulation strategies to the real environment, thereby achieving automated sorting of aerospace parts with high precision, strong robustness, and high scalability.

[0008] This invention proposes an automated sorting method for aerospace parts based on joint training of simulation and real-world teaching. It constructs a system framework consisting of a simulation data construction and generation module, a real-world data acquisition module, a data preprocessing module, a VLA model joint training module, and a task execution optimization module. The VLA model is optimized through joint training using simulation data pre-training and real-world data fine-tuning. Combined with Action-Chunking-Transformer (ACT) technology and closed-loop control, it achieves precise execution and dynamic optimization of sorting actions. The specific steps are as follows: S1. Simulation Data Construction and Generation: An interactive simulation environment for aerospace parts sorting actions is built on the MuJoCo physical simulation platform. The 3D models of aerospace parts, robotic arms, grippers, and work areas, as well as a virtual RGB camera, are loaded and physical parameters are configured. A scene randomization strategy is introduced to enhance data diversity. The grasping-movement-placement action process is automatically executed. Visual data, end effector trajectory, joint angle state, gripper opening and closing state, and action execution result labels are recorded to generate a simulation dataset containing images, actions, and semantics. S2. Real data acquisition: A multi-sensor integrated acquisition solution based on the PIKA terminal is adopted. The operator holds the PIKA terminal to perform the sorting operation of aerospace parts and records the visual flow, depth flow, six-degree-of-freedom pose, gripper status, IMU inertial information and task semantic labels in real time. After the data is automatically synchronized in time, a structured real teaching dataset is formed. S3. Data Preprocessing: Unified and standardized processing of simulation and real datasets, including image size standardization, quality filtering and data augmentation, denoising, hole filling and pixel-level alignment of depth data, coordinate system alignment, filtering smoothing and fixed-length interpolation of trajectory data, normalization and binary encoding of gripper information; and time synchronization of multimodal data based on unified timestamps, encapsulating it into a structured training dataset containing multimodal input samples and corresponding action sequences. S4, VLA Model Joint Training: Construct a VLA model consisting of a visual encoder, a text encoder, a multimodal fusion network, and an action decoder. First, pre-train the model using a simulation dataset, and then fine-tune the model using a real teaching dataset to achieve end-to-end joint learning of visual information, task text, and action sequences. S5. Task Execution Optimization: Based on the trained VLA model, robot motion instructions are generated. The Action-Chunking-Transformer technology is used to break down complex motion sequences into motion blocks and process them through the Transformer architecture. Combined with closed-loop control, real-time feedback from sensors is received to dynamically adjust the robotic arm's motion and complete the high-precision automatic sorting of aerospace parts.

[0009] The present invention also provides an automated sorting system for aerospace parts, comprising: The simulation data construction and generation module is used to generate a simulation multimodal dataset for aerospace parts sorting on the MuJoCo platform; The real data acquisition module, based on the PIKA terminal, realizes multimodal data acquisition and structured processing of sorting operations in a real environment; The data preprocessing module cleans, standardizes, synchronizes, and encapsulates simulation and real data to generate a unified structured training dataset. The VLA model joint training module enables joint training of simulation data pre-training and real data fine-tuning, and outputs a VLA model that can generate robot motion commands. The task execution optimization module, based on ACT technology and closed-loop control, transforms the instructions generated by the VLA model into precise sorting actions of the robotic arm and dynamically optimizes them.

[0010] Compared with the prior art, the technical solution of the present invention has the following significant advantages: By combining simulation data pre-training with real data fine-tuning, end-to-end learning of the VLA model was achieved, solving the problem of transferring simulation strategies to the real environment. The visual encoders of DinoV2 and SigLIP were combined to improve the model's visual recognition ability of aerospace parts, and the adaptability to changes in lighting, partial occlusion, and random placement of parts was significantly enhanced. The combination of ACT technology and closed-loop control enabled the robot to handle long-term complex sorting tasks, improving the task success rate by 15%, the execution accuracy by 20%, and the equipment failure rate by 25%.

[0011] Economic benefits: It has enabled automated sorting of aerospace parts, effectively reducing labor and production costs, increasing production efficiency by 40%, saving 30%-50% of labor expenses annually, and improving the production efficiency of aerospace manufacturing enterprises.

[0012] Social impact: It has promoted the automation and intelligentization of the aviation manufacturing industry and facilitated the application and development of robotics, multimodal fusion, and adaptive optimization technologies in the field of intelligent aviation manufacturing. Attached Figure Description

[0013] Figure 1 This is the core processing flowchart of the system of the present invention, which sequentially shows the process connection relationship of simulation data construction and generation, real data acquisition, data preprocessing, VLA model joint training, and task execution optimization. Figure 2 This is a VLA module model architecture diagram of the present invention, showing the composition and data transmission relationship of the visual encoder (DinoV2, SigLIP, MLPProjector), text encoder (Llama Tokenizer), multimodal fusion network (Llama2 7B), and action decoder (Action De-Tokenizer), and finally outputting 7D robot action commands. Detailed Implementation

[0014] The technical solution of the present invention will be described in detail below with reference to specific embodiments.

[0015] An automated sorting method for aerospace parts based on joint training of simulation and real teaching, the specific process of which is as follows: Simulation data construction and generation: Based on the MuJoCo physical simulation platform, an interactive simulation environment for sorting aerospace parts was built. 3D models of aerospace parts such as bolts, connectors, and pins, as well as robotic arm models, gripper models, and work area models such as material boxes / desktops were loaded. Virtual RGB cameras and corresponding physical parameters were configured to ensure the repeatability and accuracy of the motion simulation.

[0016] Introducing a scene randomization strategy enhances data diversity: randomly changing the position, attitude, quantity, and stacking method of aerospace parts, perturbing the viewpoint and spatial position of the virtual camera, randomly adjusting the scene lighting intensity and illumination angle, and replacing the background texture to generate visual data that closely resembles the real scene.

[0017] The complete grasping-movement-placement motion process is automatically executed through a control script: the robotic arm is reset to its initial posture, parts are randomly arranged, the target pose of the end effector is generated and the robotic arm is driven to approach, the gripper closes to grasp, lifts, moves to the placement area, and then releases. The MuJoCo dynamics engine is used to simulate physical processes such as contact, slippage, and friction, recording visual data (image frame sequences), end effector trajectory, joint angle states, gripper opening and closing states, and motion execution result markers (success) during the motion execution process. Finally, a simulation dataset containing image-motion-semantic information is generated for the basic training of the VLA model.

[0018] Real-world data acquisition: A lightweight, multi-sensor integrated data acquisition solution based on the PIKA terminal is adopted. Compared with traditional remote-operated equipment, it features a compact structure, high data accuracy, and high acquisition efficiency. The operator holds the PIKA terminal to perform sorting operations such as grasping, classifying, and aligning aerospace parts. The PIKA system records multimodal data in real time, including RGB image sequences, depth maps, terminal six-DOF pose, gripper opening and closing amplitude, acceleration and angular velocity acquired by the IMU, and task semantic labels.

[0019] The collected data is automatically synchronized in time, requiring no post-processing. Furthermore, PIKA uses the same end effector for data acquisition and execution. Real-time data transmission and spatial positioning are achieved through a data backpack and positioning base station, ensuring that the collected operation trajectory and the robot's subsequent inference instructions have consistent dynamic characteristics, ultimately forming a structured real teaching dataset.

[0020] Data preprocessing: The simulation dataset and the real teaching dataset are standardized to ensure that multimodal data is input into the VLA model in a consistent format. Specific processing steps include: Visual data: Images acquired by MuJoCo and PIKA are normalized in size, and their brightness and color are corrected. Poor quality frames are filtered out through blur detection and occlusion detection. Data augmentation such as rotation, cropping, and color perturbation is performed when necessary. Depth data: The depth data collected by PIKA is denoised, hole-filled and normalized, and the depth information is pixel-level aligned with the RGB image and transformed to the same coordinate system; Trajectory and pose data: Coordinate system alignment, filtering and smoothing, and keyframe extraction are performed on the pose and trajectory data of the end effector. The variable-length trajectory is interpolated to a fixed length to adapt to the training requirements of the sequence model. Gripper and sensor data: Normalize the opening and closing amplitude of the gripper, perform binary encoding on the contact state, and perform noise reduction and normalization if force / tactile information is included; Multimodal time synchronization: Based on a unified timestamp, image, depth, trajectory, and IMU data are aligned, missing frames are interpolated and filled in, and finally encapsulated into a structured training dataset containing multimodal input samples and corresponding action sequences / grabbing target labels.

[0021] Joint training of VLA model: Construct a VLA model consisting of a visual encoder, a text encoder, a multimodal fusion network, and an action decoder. Employ a joint training method of pre-training with simulated data and fine-tuning with real data to achieve end-to-end learning of visual information, task text, and action sequences.

[0022] The visual encoder consists of DinoV2, SigLIP, and MLP Projector. DinoV2 extracts low-level spatial features of the image (color, texture, edges, local shape), SigLIP extracts high-level semantic features of the image (part category, function, spatial relationship), and MLP Projector maps the two types of features to a feature space compatible with the language model.

[0023] Text encoder: Employs Llama Tokenizer to convert natural language task instructions into tokens that the model can process.

[0024] Multimodal fusion network: The Llama2 7B model is used to jointly process visual features and text tokens to generate a multimodal fusion feature representation that simultaneously contains visual information and task semantics.

[0025] Action Decoder: Converts multimodal fusion feature representations into 7D robot action instructions, including displacement Δx, posture Δθ, and gripping force ΔGrip, through Action De-Tokenizer.

[0026] Joint training: First, the VLA model is pre-trained using a large-scale simulation dataset to enable the model to learn capabilities such as part recognition, posture prediction, and basic motion mapping; then, the model is fine-tuned using a real teaching dataset collected by PIKA, adjusting the model weights to adapt to the material, lighting, noise, and robotic arm dynamics characteristics in the real environment, thereby improving the model's stability, accuracy, and generalization ability.

[0027] Task execution optimization Based on the trained VLA model, robot sorting action instructions are generated. ACT technology combined with closed-loop control is used to achieve precise optimization of action execution. The specific process is as follows: Action Segmentation and Processing: Complex long-sequence sorting actions are broken down into multiple independent "action blocks" using ACT technology. Each action block corresponds to a single operation step such as grasping, moving, and placing. The self-attention mechanism of the Transformer architecture is used to capture the dependencies between action blocks and generate efficient control instructions. Closed-loop control execution: The instructions generated by the VLA model drive the robotic arm to perform sorting actions. During the execution, feedback data from sensors such as vision and force are received in real time. If environmental changes or action deviations occur, the action parameters and execution strategies are dynamically adjusted immediately. Dynamic optimization: Based on real-time feedback, each action block is fine-tuned to ensure that the robotic arm accurately completes operations such as gripping, moving, and placing parts, thereby realizing automated sorting of aerospace parts and improving the success rate and accuracy of task execution.

[0028] The automated sorting system for aerospace parts proposed in this invention, based on joint training of simulation and real-world teaching, mainly consists of the following modules: (1) Simulation data construction and generation module (2) Real data acquisition module (3) Data preprocessing module (4) VLA model joint training module (5) Task execution optimization module The core processing flow of the system is as follows Figure 1 As shown: 1. Simulation data construction and generation module: This module is used to collect interactive data on the grasping, moving, and placing of aerospace parts in a simulation environment, generating large-scale structured training data at low cost and high efficiency. Based on the MuJoCo physics simulation platform, the module simulates the interaction between the robotic arm and the parts. By automatically executing motion strategies and recording the state information, pose trajectory, and success markers throughout the process, it constructs an example dataset of motions for training and evaluating the motion generation capabilities of the VLA model.

[0029] (1) Construction of MuJoCo simulation environment This invention builds an interactive motion simulation environment in MuJoCo, loads simulation objects including aerospace parts models, robotic arm models, and gripper models, and sets corresponding physical parameters, including: 1) 3D models of aerospace parts (bolts, connectors, pins, etc.) 2) Work area model (material frame, desktop, platform) 3) Robotic arm model, end effector and gripper 4) Virtual RGB camera (for image generation) This part ensures that the dynamic simulation of the motion capture process is repeatable and accurate.

[0030] (2) Scene randomization and data diversity enhancement To improve the generalization ability of the generated data, this invention introduces multiple randomization strategies into the simulation environment. These include randomly varying the position and attitude of aerospace parts, and randomly adjusting the number and stacking method of the parts; simultaneously perturbing the viewpoint and spatial position of the virtual camera to simulate different imaging conditions; further, randomly adjusting the scene's illumination intensity and angle, and replacing the background texture to increase environmental diversity. Through these randomization strategies, the simulation environment can generate visual data that more closely resembles changes in real-world scenes, thereby significantly improving the model's generalization ability in real-world environments.

[0031] (3) Action execution and data acquisition This module automatically executes the grasping action through a control script and records the data of the entire process. First, the robotic arm is reset to a preset initial posture, and parts are randomly arranged within the acquisition area. Then, the control script selects the target part according to a predetermined strategy, generates the target pose corresponding to the end effector, and drives the robotic arm to perform an approach motion. After the robotic arm reaches the target position, the gripper performs a closing operation to attempt to grasp the part, followed by lifting, moving to the placement area, and releasing the gripper, thus completing a full grasping and placement process. Because the MuJoCo dynamics engine can realistically simulate physical processes such as contact, slippage, and friction, the collected trajectory data has high physical reliability and can be used to train and evaluate the motion generation model of this invention.

[0032] (4) Track and status recording This module automatically records various data throughout the entire action execution process: 1) Visual data: During execution, a sequence of image frames is recorded by a virtual camera:

[0033] 2) End effector trajectory:

[0034] 3) Joint angle status:

[0035] 4) Gripper opening / closing state:

[0036] 5) Mark the action execution result as "success":

[0037] (5) Simulation dataset format Finally, the module outputs a simulation dataset containing image-action-semantic (optional) elements:

[0038] in, It is an image sequence. It is the trajectory of the end effector. It is in gripper mode. This is the successful execution trajectory.

[0039] This dataset was used to train the VLA model, enabling it to predict action sequences from images and task descriptions.

[0040] 2. Real Data Acquisition Module: This module employs a real-world data acquisition solution based on the PIKA terminal. Through a lightweight, multi-sensor integrated data acquisition fixture, it achieves multimodal data acquisition during the handling of aerospace parts, including visual data, depth data, motion trajectory, gripper status, and inertial information. Compared to traditional teaching methods relying on bulky remote-controlled equipment, the PIKA solution features a compact structure, ease of use, high data accuracy, and high acquisition efficiency, significantly improving the speed and quality of acquiring real-world example data.

[0041] During real-world data acquisition, operators use a handheld PIKA terminal to perform tasks such as gripping, sorting, aligning, and inserting pins on aerospace parts. The PIKA system records multimodal data in real time throughout the entire operation, including key actions and complete trajectory sequences, interaction information between the parts and the fixture, gripper opening and closing amplitude, RGB image sequences, depth maps, the terminal's six-DOF pose, and acceleration and angular velocity information acquired by the IMU. The recorded data is automatically time-synchronized, requiring no post-processing, ultimately forming a structured data sample.

[0042] It includes visual flow, depth flow, trajectory information, gripper status, and task semantic labels.

[0043] Furthermore, PIKA uses the same end effector for data acquisition and execution. It achieves real-time data transmission and spatial positioning of the operation process through a data backpack and positioning base station. This ensures that the acquired operation trajectory and the subsequent inference instructions executed by the robot have consistent dynamic characteristics, thereby improving the transferability and execution accuracy of the trained model on real robots.

[0044] 3. Data Preprocessing Module: The data preprocessing module of this invention is used to uniformly standardize multimodal data obtained from simulation and real acquisition environments, enabling visual data, depth data, pose trajectories, and sensor information to be input into the Vision-Language-Motion Model (VLA) in a consistent format. Specifically, for image data from MuJoCo and PIKA, the system first performs size normalization and brightness and color correction, and filters out poor-quality frames using methods such as blur detection and occlusion detection; when necessary, it performs data augmentation such as rotation, cropping, and color perturbation to improve the model's robustness to complex visual conditions.

[0045] For the depth data collected by the PIKA terminal, the system performs depth denoising, hole filling, and normalization processing, and aligns the depth information with the RGB image at the pixel level, while uniformly transforming it to the same coordinate system for use in conjunction with trajectory data. For the end effector pose and trajectory data of MuJoCo and PIKA, the system reduces noise disturbances through coordinate system alignment, filtering smoothing, and keyframe extraction, and interpolates variable-length trajectories to a fixed length to adapt them to the training requirements of sequence models.

[0046] Regarding gripper and contact information, the system normalizes the gripper opening and closing amplitude and performs binary encoding on the contact state. If force or tactile information is included, further denoising and normalization processing is performed to help distinguish between successful and failed gripping modes. To achieve multimodal time synchronization, this module aligns image, depth, trajectory, and IMU data based on a unified timestamp, interpolates missing frames when necessary, and finally encapsulates all types of information into a unified input format.

[0047] in For image frames, For depth information, For position, In gripper mode, For inertial data, This is a semantic description of the task.

[0048] Finally, this module outputs a structured training dataset:

[0049] in For multimodal input samples, This dataset, which corresponds to the action sequence or the target label to be captured, can be directly used for joint training of subsequent VLA models.

[0050] 4. VLA Model Joint Training Module The VLA model joint training module of this invention aims to generate efficient sorting operation control instructions through multimodal learning of visual data and task instructions. This module consists of four core modules: a visual encoder, a text encoder, a multimodal fusion network, and an action decoder. It can convert input image and language information into specific action instructions, realizing an end-to-end learning process from perception to execution.

[0051] (1) Visual encoder The task of a visual encoder is to extract key features from an input image and generate a visual feature representation. In this invention, the visual encoder uses the DinoV2 and SigLIP models to extract self-supervised features and language-related features from the image, respectively. By concatenating these two types of features, the visual encoder provides high-quality visual information for subsequent multimodal fusion, ensuring that the model can understand the types of parts, spatial layout, and their environment in the image.

[0052] 1) DinoV2: DINOv2 (Low-Level Spatial Information): DINOv2 extracts low-level spatial features from an image, namely detailed information such as color, texture, edges, and local shape. These features describe the specific spatial structure of the image and help understand the basic features such as the specific location and shape of objects. This helps robots determine the position, orientation, and relationships between objects.

[0053] 2) SigLip: SigLIP (High-Level Semantic Information): SigLIP provides high-level semantic information, namely the understanding of objects and scenes in an image, such as the category and function of objects and the semantic relationships between them, helping the model understand the overall meaning of the image. High-level semantic information is crucial for performing complex tasks (such as voice command execution) because it enables robots to understand the semantic roles of different objects in an image.

[0054] MPL Projector: After extracting image features, the projector is responsible for mapping these visual features (from DinoV2 and SigLIP) into a space compatible with the language model, so that visual information and language instructions can be combined and processed.

[0055] (2) Text encoder In this invention, the Llama Tokenizer, acting as a text encoder, is responsible for converting natural language instructions into tokens that the model can process, allowing the Llama2 model to further process them. This component is responsible for converting language instructions into text tokens, i.e., a form that the language model can understand and process. In this way, the Llama Tokenizer achieves efficient conversion between natural language and model input in this invention, ensuring accurate understanding and automated execution of task instructions.

[0056] (3) Multimodal fusion network The multimodal fusion network of this invention uses the Llama 2 7B model to jointly process image data and language instructions to generate a joint multimodal feature representation that can simultaneously understand visual information and task semantics.

[0057] (4) Action decoder The action decoder module of this invention processes the multimodal feature representation generated by the Llama 2 7B model, transforming the fused image tokens and text tokens into actual operation commands for the robot. First, the Llama 2 7B uses a self-attention mechanism to jointly process visual information and task semantic information, generating a multimodal feature vector. This feature vector is then input into the Action De-Tokenizer, which converts it into specific 7D robot action commands, including displacement (Δx), pose (Δθ), and gripping force (ΔGrip). These commands control the robot to perform gripping, moving, and placing tasks, ensuring the robot can complete sorting operations accurately and efficiently. The overall model architecture of the VLA module is as follows: Figure 2 As shown: (5) Combining simulation pre-training and fine-tuning with real data The VLA model joint training module of this invention achieves end-to-end joint learning of visual data, task text information, and action examples by combining pre-training with simulation data and fine-tuning with real data. Through this training process, the model can automatically understand the task intent based on the input image and task description, and generate corresponding grasping and manipulation action sequences, thereby improving the model's performance in real-world environments. During training, pre-training is first performed using large-scale simulation data. The model learns basic capabilities such as part recognition and pose prediction through labeled images and task examples generated in the simulation environment. The use of simulation data helps the model quickly acquire visual understanding and preliminary action mapping capabilities, especially in the early stages of training, enabling efficient generation of basic feature representations. However, since there may be differences between the simulation environment and the actual environment, fine-tuning with real data becomes a crucial step in further optimizing model performance.

[0058] To better adapt the model to challenges in real-world environments, such as changes in materials, lighting, noise, and the dynamics of the robotic arm, this invention employs a PIKA-based real-world data acquisition scheme for model fine-tuning. During fine-tuning, the model receives operational data from the real environment and progressively adjusts its weights to cope with complex changes under real-world physical conditions. Real-world data fine-tuning not only improves the model's stability and accuracy but also enables it to maintain high adaptability and robustness across different tasks and environmental variations. Ultimately, by combining simulation data pre-training with real data fine-tuning, the VLA model can automatically generate complete operation instruction sequences, providing precise grasping and sorting strategies for subsequent action execution modules, realizing end-to-end learning from perception to action decision-making, and significantly improving the robot's performance in practical applications.

[0059] 5. Task Execution Optimization Module The task execution optimization module of this invention employs Action-Chunking-Transformer (ACT) technology to optimize the robot's motion generation and control strategies during task execution. Based on the Transformer architecture, ACT technology breaks down the task into a series of finer-grained operational units by chunking the taught data, thereby improving the accuracy and efficiency of task execution. Through ACT optimization, this module enables the robot to process each small stage of the task in real time and continuously adjust and optimize during task execution.

[0060] (1) Action execution During task execution, the VLA model controls the robotic arm or other actuators to perform operations such as grasping, moving, and placing by generating motion sequences. First, the model generates precise control commands based on visual data and task description, determining the target position, grasping posture, path planning, etc., and guiding the robotic arm to perform the task along a predetermined path and posture. During execution, the module receives real-time feedback from sensors, such as visual and force data, and adjusts its actions accordingly to ensure accurate task completion. Through closed-loop control, the module can dynamically adjust its operating strategy to ensure task success rate and execution stability, especially in the event of environmental changes or operational deviations, enabling timely correction of actions to ensure the successful achievement of the task objective.

[0061] (2) Task optimization This invention employs Action-Chunking-Transformer (ACT) technology to optimize motion generation and control strategies during robot task execution. The core idea of ​​ACT is to break down complex motion sequences into multiple "action chunks," each representing an independent step in task execution, thereby improving task efficiency and accuracy. First, the system divides long sequences of actions into smaller chunks, each including steps such as grasping, moving, and placing, ensuring precise execution of each step. Then, the Transformer architecture processes each action chunk, capturing dependencies between actions through a self-attention mechanism to generate efficient control commands. Finally, based on real-time feedback, ACT dynamically fine-tunes each action chunk, enabling the robot to adjust its strategy according to environmental changes during execution, coping with complex task environments. During training, ACT combines teaching data and simulation data for optimization. Pre-training with simulation data teaches visual understanding and motion sequence generation capabilities, while fine-tuning with real-world operational data collected by PIKA allows the model to adapt to real-world physical conditions. Furthermore, ACT incorporates closed-loop control technology, ensuring precise task execution through real-time feedback and adjustments, thereby improving task success rate and execution accuracy.

[0062] This embodiment uses the UR5e robotic arm as the execution carrier and bolts, connectors, and pins commonly used in aerospace manufacturing as sorting objects to realize the automatic sorting method for aerospace parts of the present invention. The specific implementation process is as follows: Simulation environment setup and data generation A sorting scenario was built on the MuJoCo physical simulation platform. The UR5e robotic arm model, pneumatic gripper model, 3D models of bolts / connectors / pins, and material frame working area model were loaded. A virtual RGB camera (1920×1080 resolution) was configured, and physical parameters such as part friction coefficient and gripper clamping force were set.

[0063] The scene randomization strategy is implemented as follows: the spatial position (±5cm) and placement posture (0-360°) of the parts are randomly adjusted, the number of parts in a single frame scene (1-5) and the stacking method are randomly set, the camera's field of view (±10°) and height (±3cm) are disturbed, the light intensity (0.5-1.5 times) and illumination angle (0-90°) are adjusted, and the background with 3 different textures is replaced.

[0064] The Python control script automatically executes 100,000 grasping-moving-placing actions, recording the image frame sequence, end effector trajectory, joint angle changes, gripper opening and closing amplitude, and action success / failure markers for each action, generating a simulation dataset containing images, actions, and semantics.

[0065] Real teaching data collection The UR5e robotic arm was configured with a PIKA data acquisition terminal. The operator held the PIKA terminal and performed sorting operations on bolts, connectors, and pins at a real sorting station, collecting a total of 10,000 valid operation data. The PIKA terminal recorded RGB image sequences, depth maps, six-DOF poses, gripper opening and closing amplitudes, and IMU inertial information in real time, and added task semantic tags such as "grab bolts" and "place connectors to material frame 1". The collected data was automatically synchronized in time to form a structured real teaching dataset.

[0066] Data preprocessing Preprocessing of simulation and real data: Images are uniformly scaled to 224×224, and data augmentation with brightness correction and random flipping is performed; Gaussian denoising and hole filling are performed on depth data, and pixel-level alignment with RGB images is achieved; Kalman filtering is applied to smooth trajectory data, and variable-length trajectories are interpolated to a fixed length of 128 frames; the opening and closing amplitude of the grippers is normalized to the [0,1] interval, and the contact state is encoded using 0 / 1 binary encoding; multimodal data synchronization is achieved based on timestamps, and finally, the data is packaged into a structured training dataset.

[0067] VLA model joint training A training environment was built using an NVIDIA GeForce RTX 4090 graphics card. A VLA model was constructed and joint training was performed: First, the model was pre-trained using a simulation dataset for 50 epochs with a batch size of 32, enabling the model to recognize parts and map basic movements. Then, the model was fine-tuned using a real teaching dataset for 20 epochs with a batch size of 16, adjusting the model weights to adapt to the real environment. After training, the model can automatically generate 7D robot movement commands based on the input image and natural language task instructions.

[0068] Sorting task execution and optimization The trained VLA model is deployed to the RTDE control system of the UR5e robotic arm. The task instruction "sort the bolts, connectors, and pins in the material boxes to their corresponding material boxes" is input. The robotic arm executes the sorting operation based on the motion instructions generated by the model. During execution, the status of the parts is collected in real time by a vision sensor, and the gripping force of the grippers is detected by a force sensor. ACT technology is used to break down the sorting action into action blocks of "identifying parts - approaching parts - gripping - moving - placing." Each action block is dynamically fine-tuned based on sensor feedback to achieve closed-loop control.

[0069] In this embodiment, the success rate of automatic sorting of aviation parts reaches over 95%, and the sorting time per part is reduced by 60% compared to manual sorting. This effectively achieves high-precision and high-efficiency automated sorting of aviation parts, verifying the feasibility and superiority of the method of the present invention.

[0070] This invention presents an automated sorting method for aerospace parts based on joint training of simulation and real-world teaching. This method automates the sorting of multi-category parts with complex placement in the aerospace manufacturing field, solving the problems of low sorting accuracy, poor robustness, and difficulty in transferring simulation to real-world environments in existing technologies. It significantly improves the efficiency and accuracy of aerospace parts sorting. The method can be directly deployed to industrial robot systems, adapting to various aerospace manufacturing scenarios such as aerospace assembly, aerospace material management, and automated aerospace warehousing. Furthermore, its core simulation-real-world joint training strategy can be extended to automated parts operations in other fields such as mechanical manufacturing and automotive assembly, demonstrating excellent industrial applicability and promotional value.

Claims

1. An automated sorting method for aerospace parts based on a combination of simulation and real-world teaching training, characterized in that, Includes the following steps: S1. Simulation Data Construction and Generation: An interactive simulation environment for aerospace parts sorting actions is built on the MuJoCo physical simulation platform. The 3D models of aerospace parts, robotic arms, grippers, and work areas, as well as a virtual RGB camera, are loaded and physical parameters are configured. A scene randomization strategy is introduced to enhance data diversity. The grasping-movement-placement action process is automatically executed. Visual data, end effector trajectory, joint angle state, gripper opening and closing state, and action execution result labels are recorded to generate a simulation dataset containing images, actions, and semantics. S2. Real data acquisition: A multi-sensor integrated acquisition solution based on the PIKA terminal is adopted. The operator holds the PIKA terminal to perform the sorting operation of aerospace parts and records the visual flow, depth flow, six-degree-of-freedom pose, gripper status, IMU inertial information and task semantic labels in real time. After the data is automatically synchronized in time, a structured real teaching dataset is formed. S3. Data preprocessing: Unify and standardize the simulation and real datasets, including image size standardization, quality filtering and data augmentation, denoising, hole filling and pixel-level alignment of depth data, coordinate system alignment, filtering smoothing and fixed-length interpolation of trajectory data, and normalization and binary encoding of gripper information. Multimodal data time synchronization is achieved based on a unified timestamp, and it is encapsulated into a structured training dataset containing multimodal input samples and corresponding action sequences; S4, VLA Model Joint Training: Construct a VLA model consisting of a visual encoder, a text encoder, a multimodal fusion network, and an action decoder. First, pre-train the model using a simulation dataset, and then fine-tune the model using a real teaching dataset to achieve end-to-end joint learning of visual information, task text, and action sequences. S5. Task Execution Optimization: Based on the trained VLA model, robot motion instructions are generated. The Action-Chunking-Transformer technology is used to break down complex motion sequences into motion blocks and process them through the Transformer architecture. Combined with closed-loop control, real-time feedback from sensors is received to dynamically adjust the robotic arm's motion and complete the high-precision automatic sorting of aerospace parts.

2. The method according to claim 1, characterized in that, The scene randomization strategy described in step S1 includes: randomly changing the position, attitude, quantity and stacking method of aviation parts, perturbing the viewpoint and spatial position of the virtual camera, randomly adjusting the scene lighting intensity and illumination angle and replacing the background texture.

3. The method according to claim 1, characterized in that, The visual encoder described in step S4 consists of DinoV2, SigLIP, and MLP Projector. DinoV2 extracts low-level spatial features of the image, SigLIP extracts high-level semantic features of the image, and MLP Projector maps the two types of features to a feature space compatible with the language model.

4. The method according to claim 1, characterized in that, The text encoder in step S4 is LlamaTokenizer, which is used to convert natural language task instructions into tokens that the model can process; the multimodal fusion network is Llama2 7B model, which is used to jointly process visual features and text tokens to generate multimodal fusion feature representations.

5. The method according to claim 1, characterized in that, The action decoder in step S4 includes an Action De-Tokenizer, which is used to convert the multimodal fusion feature representation into 7D robot action instructions, wherein the 7D robot action instructions include displacement Δx, posture Δθ, and gripping force ΔGrip.

6. The method according to claim 1, characterized in that, The optimization process of the Action-Chunking-Transformer technology in step S5 is as follows: the long sequence of actions is divided into action blocks containing independent steps such as grasping, moving, and placing; the dependency relationship between action blocks is captured through a self-attention mechanism to generate control instructions; and each action block is dynamically fine-tuned based on real-time sensor feedback.

7. The method according to claim 1, characterized in that, The aerospace parts mentioned in step S1 include bolts, connectors, and pins. The actions performed by the robotic arm in step S5 are achieved through closed-loop control, receiving feedback from visual and force sensors in real time and correcting action deviations.

8. An automated sorting system for aerospace parts implementing the method of any one of claims 1-7, characterized in that, include: The simulation data construction and generation module is used to generate a simulation multimodal dataset for aerospace parts sorting on the MuJoCo platform; The real data acquisition module, based on the PIKA terminal, realizes multimodal data acquisition and structured processing of sorting operations in a real environment; The data preprocessing module cleans, standardizes, synchronizes, and encapsulates simulation and real data to generate a unified structured training dataset. The VLA model joint training module enables joint training of simulation data pre-training and real data fine-tuning, and outputs a VLA model that can generate robot motion commands. The task execution optimization module, based on ACT technology and closed-loop control, transforms the instructions generated by the VLA model into precise sorting actions of the robotic arm and dynamically optimizes them.