A multi-modal adaptive grasping system for key component assembly of humanoid robots and a control method thereof

By using a multimodal adaptive gripping system that combines multimodal sensing, edge computing, and cloud services, the rapid and stable assembly of multiple joint types in humanoid robots has been achieved. This solves the problems of long changeover time and poor stability in traditional humanoid robots, and improves assembly efficiency and accuracy.

CN120645216BActive Publication Date: 2025-12-09广州里工实业有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510822988.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-12-09
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Traditional humanoid robot assembly technology suffers from problems such as long model changeover time, high dependence on fixtures, and low efficiency and poor stability due to limited sensing methods, making it difficult to meet the requirements of high precision and rapid switching between multiple models.

Method used

A multimodal adaptive gripping system is adopted, which collects assembly parameters through a multimodal sensing module, generates control commands through an edge computing module, processes distillation data through a cloud service module, and realizes gripping and movement through an adaptive gripping module and a programmable logic control module. It also combines digital twin technology for simulation and online optimization.

Benefits of technology

It enables rapid, stop-and-go joint changes, sub-millimeter alignment, and millinews-level compliant force control for multiple joint types of humanoid robots without changing fixtures, thus improving assembly efficiency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120645216B_ABST
    Figure CN120645216B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal adaptive gripping systems for humanoid robot key components assembly and control method thereof, system includes: multi-modal sensing module, for collecting assembly parameter;Edge computing module, for according to the assembly parameter, control instruction is generated using multi-modal deep learning model;Cloud service module, for according to the assembly parameter, distillation data is generated;Adaptive gripping module, for gripping and moving humanoid robot key components;Programmable logic control module, for according to the control instruction and the distillation data, control the adaptive gripping module executes component assembly operation.The application realizes multi-modal adaptive gripping system, improves efficiency and gripping stability.The application can be widely applied to industrial robot technical field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of industrial robots, and in particular to a multi-modal adaptive gripping system for assembling key components of humanoid robots and a control method thereof. BACKGROUND

[0002] With the breakthroughs in bionic control, electric drive and speed reduction integration, and high energy density batteries, humanoid robots are rapidly moving towards mass production and multi-scene applications. The key components of humanoid robots, such as joint modules, lightweight skeletons, end effectors, and multi-modal perception heads, have high iteration frequency and large model differences, and require "sub-millimeter coaxiality + millinewton level compliant clamping" dual precision for assembly. At the same time, humanoid robots often use a combination of carbon fiber-magnesium-aluminum alloy skeletons and silica gel outer layer materials. Excessive clamping pressure may cause permanent deformation or internal micro-cracks, further increasing the difficulty of force control for assembly fixtures. Traditional methods use "rigid special fixture + PLC solidification program + manual teaching", but each time a joint or peripheral device is replaced, the fixture and the line need to be replaced and leveled, which takes a long time. Furthermore, the process data and PLC ladder diagram are converted manually or by script, and the entire program needs to be recompiled and verified after design changes, which is inefficient. At the same time, the perception means mainly rely on a single RGB camera or offline measuring head, and manual teaching is performed, which can easily lead to quality problems such as bruising, slipping, and wire kinking, resulting in low clamping stability.

[0003] In summary, the technical problems in the related art need to be improved. SUMMARY

[0004] The present application provides a multi-modal adaptive gripping system for assembling key components of humanoid robots and a control method thereof, which effectively improves efficiency and clamping stability.

[0005] In one aspect, the present application provides a multi-modal adaptive gripping system for assembling key components of humanoid robots, comprising:

[0006] A multi-modal sensing module for collecting assembly parameters;

[0007] An edge computing module for generating control instructions using a multi-modal deep learning model based on the assembly parameters;

[0008] A cloud service module for generating distilled data based on the assembly parameters;

[0009] An adaptive gripping module for gripping and moving key components of humanoid robots;

[0010] A programmable logic control module for controlling the adaptive gripping module to perform component assembly operations based on the control instructions and the distilled data.

[0011] In some embodiments, the assembly parameters include color images, depth information, three-directional forces, three-directional force moments, position slip amounts of clamp silica gel shells, and temperature change amounts of the clamp silica gel shells, and the multi-modal sensing module comprises:

[0012] An industrial-grade RGB-D camera is configured to collect color images and depth information.

[0013] A six-dimensional force-displacement sensor is configured to collect three-directional forces and three-directional force moments.

[0014] A miniature inertial measurement unit is configured to collect position slip amounts of clamp silica gel shells.

[0015] A flexible temperature patch is configured to collect temperature change amounts of the clamp silica gel shells.

[0016] In some embodiments, the cloud service module comprises:

[0017] A digital twin cluster is configured to perform assembly simulation using a digital twin framework according to the assembly parameters to obtain simulation data.

[0018] A cloud server is configured to perform model training to generate the distillation data according to the simulation data.

[0019] In some embodiments, the adaptive gripping module comprises:

[0020] An adaptive clamp is configured to grip humanoid robot key components.

[0021] A collaborative robot arm is configured to move the adaptive clamp to a predetermined position.

[0022] In some embodiments, the adaptive clamp comprises:

[0023] A mechanical fingertip sub-component, which is composed of a parallel three-finger structure and a miniature ball screw, is configured to grip the humanoid robot key components by width adjustment.

[0024] A vacuum adsorption sub-component is configured to adsorb the humanoid robot key components by atmospheric pressure.

[0025] An electro-permanent sub-component is configured to adsorb humanoid robot key components with metal shells by magnetic force.

[0026] The present application has the following beneficial effects:

[0027] The multi-modal adaptive gripping system for assembling key components of a humanoid robot provided by the embodiment of the present application comprises a multi-modal sensing module, an edge computing module, a cloud service module, an adaptive gripping module and a programmable logic control module. The multi-modal sensing module is used to collect assembly parameters; the edge computing module is used to generate control instructions by using a multi-modal deep learning model according to the assembly parameters; the cloud service module is used to generate distilled data according to the assembly parameters; the adaptive gripping module is used to grip and move key components of the humanoid robot; and the programmable logic control module is used to control the adaptive gripping module to perform component assembly operations according to the control instructions and the distilled data, so as to realize the multi-modal adaptive gripping system and improve the efficiency and gripping stability.

[0028] In another aspect, the embodiment of the present application provides a control method applied to the system, comprising the following steps:

[0029] Obtaining humanoid robot data, wherein the humanoid robot data comprises a humanoid robot joint CAD model and a bill of materials;

[0030] Generating a design semantic vector and a virtual production line scene according to the humanoid robot data;

[0031] In the virtual production line scene, performing offline simulation and distillation according to the design semantic vector to generate a gripper control model and a programmable logic controller ladder diagram;

[0032] Performing adaptive gripping control according to the gripper control model and the programmable logic controller ladder diagram to obtain component assembly records, wherein the component assembly records comprise coaxiality, maximum stress, temperature rise and retry count.

[0033] In some embodiments, the generating a design semantic vector and a virtual production line scene according to the humanoid robot data comprises:

[0034] Performing analysis and processing on the humanoid robot joint CAD model by using a CAD parser to obtain basic elements and adjacency relationships, wherein the basic elements comprise faces, edges and vertices;

[0035] Generating a geometric topology graph according to the basic elements and the adjacency relationships;

[0036] Generating a geometric vector by using a graph convolution network according to the geometric topology graph;

[0037] Generating a first semantic vector according to the bill of materials, an assembly sequence and a moment curve;

[0038] Splicing the geometric vector and the first semantic vector to obtain the design semantic vector;

[0039] According to the design semantic vector, the virtual production line scene is generated.

[0040] In some embodiments, according to the design semantic vector, off-line simulation and distillation are performed to generate a fixture control model and a programmable logic controller ladder diagram, including:

[0041] According to the design semantic vector, incoming tolerance, temperature and humidity, and assembly posture, model training is performed by using reinforcement learning to obtain the fixture control model;

[0042] The fixture control model is compressed by using a contrast distillation strategy to obtain light network weights;

[0043] According to the light network weights, the programmable logic controller ladder diagram is generated.

[0044] In some embodiments, according to the fixture control model and the programmable logic controller ladder diagram, adaptive gripping control is performed to obtain a component assembly record, including:

[0045] After calibrating the multi-modal sensing module, assembly parameters are collected;

[0046] The assembly parameters are aligned and fused by using a cross-modal multi-head attention network to obtain an assembly semantic vector;

[0047] According to the assembly semantic vector, instruction parameters are generated by using a graph neural network and the fixture control model, the instruction parameters including posture adjustment amount, clamping force size, insertion speed curve, and motion trajectory;

[0048] According to the instruction parameters and the programmable logic controller ladder diagram, the adaptive gripping module is controlled to perform component assembly operations to obtain the component assembly record.

[0049] In some embodiments, the method further includes:

[0050] When online quality monitoring is performed, the fixture control model is parameter-tuned according to the component assembly record, and the fixture control model is updated;

[0051] When full-day data aggregation is performed, the fixture control model is parameter-aggregated according to the component assembly record and extreme working condition simulation results, and the fixture control model is updated;

[0052] When model changeover is performed, the fixture control model is updated according to the component assembly record and small-scale simulation results.

[0053] In another aspect, an embodiment of the present application provides a computer device, including:

[0054] at least one processor;

[0055] at least one memory for storing at least one program;

[0056] When the at least one program is executed by the at least one processor, the at least one processor implements the method.

[0057] In another aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method.

[0058] The present application has the following beneficial effects:

[0059] The embodiment of the present application first acquires humanoid robot data, then generates design semantic vectors and virtual production line scenes according to the humanoid robot data, then performs offline simulation and distillation according to the design semantic vectors, generates fixture control models and programmable logic controller ladder diagrams, and finally performs adaptive gripping control according to the fixture control models and the programmable logic controller ladder diagrams to obtain component assembly records, thereby realizing adaptive gripping control and improving efficiency and gripping stability.

[0060] Other features and advantages of the present application will be described in the following description, and some will become apparent from the description, or will be learned through practice of the present application. The objects and other advantages of the present application will be achieved and obtained by the structure particularly pointed out in the description and the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0062] Figure 1 It is a structural schematic diagram of a multi-modal adaptive gripping system for key component assembly of a humanoid robot according to an embodiment of the present application;

[0063] Figure 2 It is a flowchart of a control method applied to a multi-modal adaptive gripping system according to an embodiment of the present application;

[0064] Figure 3 It is a schematic diagram of a fixture control running process according to an embodiment of the present application;

[0065] Figure 4 It is a hardware structure schematic diagram of a computer device according to an embodiment of the present application. DETAILED DESCRIPTION

[0066] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with embodiments of the present application. They are merely examples of apparatuses and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.

[0067] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining".

[0068] The terms "at least one", "multiple", "each", "any", and the like used in the present application include one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any refers to any one of the multiple.

[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0070] Before the embodiments of the present application are described in detail, first, some nouns and terms involved in the embodiments of the present application are described, and the nouns and terms involved in the embodiments of the present application are applicable to the following explanations.

[0071] Programmable Logic Controller (PLC): is a digital operation electronic system specially designed for application in industrial environment. It uses a programmable memory to store instructions for performing logic operations, sequential control, timing, counting and arithmetic operations, etc. in its internal storage, and controls various types of mechanical equipment or production processes through digital or analog input and output.

[0072] In the related art, with the breakthroughs in bionic control, electric drive and deceleration integration, and high energy density batteries, humanoid robots are rapidly moving towards mass production and multi-scene applications. The high iteration frequency and large model difference of the core components such as joint modules, lightweight skeletons, end effectors, and multi-modal perception heads of humanoid robots put forward the dual precision requirements of "sub-millimeter coaxiality + millinewton compliant clamping" for assembly. At the same time, humanoid robots often use a mixed material structure of carbon fiber-magnesium alloy skeleton and silica gel outer layer, and slight overpressure may cause permanent deformation or internal micro-cracks, further increasing the difficulty of force control of the assembly fixture. The existing production lines mostly follow the traditional mode of "rigid special fixture + manual teaching + PLC solidification program": on the one hand, the fixture and the line need to be replaced and leveled every time a joint or peripheral device is replaced, and the replacement time is 30-60 minutes; on the other hand, the sensing means mainly rely on a single RGB camera or an offline probe, lacking real-time fusion with force, temperature, and other data, resulting in quality problems such as bruising, slipping, and wire kinking. Even if the industry starts to introduce parallel / collaborative robots and offline simulation automatic programming, it is still subject to fixture rigidity dependence, sensing information island, and PLC program reprogramming costs, making it difficult to support the assembly tempo and yield requirements of humanoid robots "high heterogeneity, small batch, and fast iteration".

[0073] The problems of the traditional method include: (1) rigid fixture dependence, high replacement cost. The traditional humanoid robot production line uses rigid fixtures and tooling customized for a single part. Whenever the joint specifications, shell material, or wire harness orientation change slightly, the production line must replace the fixture or add temporary shims, and then manually level, align, and re-calibrate. The average downtime is 30-60 minutes, which not only delays delivery, but also increases fixture inventory and maintenance costs.(2) single modal sensing is fragmented and cannot be compliant force controlled. Existing solutions mostly rely on single RGB / 2D vision or offline laser probes to obtain pose information, lacking deep integration with real-time data such as force, temperature, and IMU. When the fixture grabs soft and hard mixed parts such as carbon fiber-silica gel, it is difficult to detect overpressure or slipping in time, resulting in quality problems such as bruising, wire entanglement, and micro-cracks, affecting the overall machine gait and safety.(3) PLC program solidification, lack of self-learning and hot replacement. Process data (CAD / BOM) and PLC ladder diagrams still use manual transcription or script conversion, and after design changes, the entire program needs to be reprogrammed and repeatedly verified; at the same time, the assembly strategy cannot be adaptively optimized based on real-time sensor feedback, resulting in the need to stop the line and adjust the parameters when the incoming tolerance and environmental drift exceed the boundaries.

[0074] Therefore, there is an urgent need for an adaptive gripping system that can integrate design priors (CAD / BOM) and real-time multimodal perception (vision-force-temperature-IMU) within a unified algorithm framework, and possess online self-learning and PLC hot-swap capabilities. This system should be like a human hand—"seeing clearly, touching accurately, and learning quickly"—enabling rapid switching between different humanoid robot joints and materials without production line downtime, while balancing high-precision alignment and compliant force control, fundamentally improving production line flexibility, yield, and economy. The aim is to address the following issues: how to achieve rapid, stop-and-go changeover and alignment of humanoid robots by quickly iterating on dimensions, materials, and assembly interfaces without changing fixtures; how to unify CAD / BOM design priors with real-time data from multiple modalities such as vision, force, temperature, and IMU into a single algorithm framework to complete attitude-force control dual closed-loop decisions in milliseconds to meet sub-millimeter coaxiality and millinews compliant clamping requirements; and how to introduce online self-learning and PLC hot-swap mechanisms to enable assembly strategies to continuously evolve with part batches, tolerance fluctuations, and environmental changes without manual teaching or downtime for reprogramming.

[0075] In view of this, this invention proposes a four-domain closed-loop panoramic multimodal adaptive gripping system for humanoid robot joint-actuator assembly, encompassing "design-perception-control-learning". The system is based on a digital twin-driven cross-modal Transformer algorithm stack, integrating CAD / BOM design priors, RGB-D vision, and multi-source real-time data from force, temperature, and IMU. Through a GNN-RL decision network, it outputs gripper posture and clamping force commands at the 20ms level, and utilizes a PLC hot-loading interface to achieve millisecond-level motion execution and online fine-tuning. Through simulation-real-machine comparative distillation and edge fine-tuning mechanisms, the assembly strategy can self-evolve under continuous operation, adapting to changes in part type, tolerance, and environment. This enables continuous, rapid joint changeover, sub-millimeter alignment, and millinews-level compliant force control for multiple humanoid robot joint types on a single gripper platform. This solution simultaneously considers high precision, flexibility, and economy, overcoming the technical bottlenecks of traditional rigid grippers, single-modal perception, and fixed PLC programs. It provides a replicable and sustainably optimized intelligent assembly path for the large-scale production of humanoid robots.

[0076] The embodiments of this application will be explained in detail below with reference to the accompanying drawings:

[0077] like Figure 1 As shown, this embodiment of the invention provides a multimodal adaptive gripping system for assembling key components of a humanoid robot, comprising:

[0078] Multimodal sensing module 101 is used to collect assembly parameters;

[0079] Edge computing module 102 is used to generate control commands based on assembly parameters using a multimodal deep learning model;

[0080] The cloud service module 103 is configured to generate distillation data according to the assembly parameters.

[0081] The adaptive gripping module 104 is configured to grip and move the key components of the humanoid robot.

[0082] The programmable logic control module 105 is configured to control the adaptive gripping module to perform component assembly operations according to the control instructions and the distillation data.

[0083] In some embodiments, the multi-modal adaptive gripping system for key component assembly of a humanoid robot includes a multi-modal sensing module 101, an edge computing module 102, a cloud service module 103, an adaptive gripping module 104, and a programmable logic control module 105. The assembly parameters can be collected by the multi-modal sensing module 101. For example, color images and depth information can be collected using an industrial-grade binocular RGB-D camera (120 frames per second), three-way force and three-way torque can be collected using a six-axis force-displacement sensor (accuracy 0.1 Newton), the position slip of the silicone shell of the gripper can be collected using a micro inertial measurement unit (1 kilohertz sampling), and the temperature change of the silicone shell of the gripper can be collected using a flexible temperature patch. All devices use PTP (Precision Time Protocol) for clock calibration to ensure consistent data timestamps.

[0084] The control instructions can be generated by the edge computing module 102 in combination with a multi-modal deep learning model according to the assembly parameters. The edge computing module is built-in with a high-performance GPU that can be used to run a multi-modal deep learning model, supporting a 20-millisecond-level inference cycle to generate control instructions. For example, the edge computing module can use an edge computing box, which is the "brain" of the entire production line. It uses an embedded GPU card with an inference power of about twenty TOPS (trillion operations per second), which can complete multi-modal data fusion and strategy calculation in about twenty milliseconds. To ensure industrial reliability, the motherboard is designed for wide temperature, the power supply supports DC 24-volt input, and an uninterruptible power supply module is integrated to prevent model parameter loss caused by instantaneous power failure. The edge computing box is pre-installed with real-time Linux and a container management system. Each version of the deep learning model is distributed in the form of a Docker image, which can hot-load new weights without interrupting device operation. The edge computing box communicates with the upload server through ten gigabit Ethernet and quickly sends instructions to the on-site PLC through a dual-port TSN (Time Sensitive Network) network card to ensure that the end-to-end delay of the decision link is within the sub-millisecond level.

[0085] The distillation data can be generated by the cloud service module 103 according to the assembly parameters. For example, the cloud service module can run a digital twin simulation to virtually create various humanoid robot assembly scenarios using a Unity-ROS framework, and generate training and distillation data in batches. The adaptive gripping module 104 can grip and move key components of the humanoid robot. For example, the adaptive gripping module can be connected to a PLC in real time through an EtherCAT bus, and an integrated gripper can be installed at the end of the module. The gripper can grip objects with mechanical fingertips, and can switch between vacuum or electric permanent magnet adsorption modes. The programmable logic control module 105 can control the adaptive gripping module to perform component assembly operations according to control instructions and distillation data. For example, the programmable logic control module can be responsible for 1 kHz (i.e., once per millisecond) servo closed-loop control and provide a "ladder diagram" program hot replacement interface. The production line communication uses a two-level network. The motion control layer uses EtherCAT with a bus cycle of one millisecond, and is responsible for high-speed interaction of mechanical arm joint driving, servo motors, and force sensors. The information fusion layer uses a gigabit TSN network, and is responsible for uploading large amounts of visual stream, model reasoning instructions, and logs. Both network cards support PTP (Precision Time Protocol) hardware time stamping, all nodes are synchronized by a master clock, and the maximum clock drift is kept within one microsecond, thereby ensuring strict alignment of multi-modal data on the time axis. The network switch integrates dual power supply and industrial-grade fans, supports port-level power failure detection and ring network redundancy. When any link packet loss or delay exceeds the standard, the switch immediately switches to the standby path and sends an alarm event to the edge box and PLC, ensuring that the production line does not stop due to communication failure.

[0086] It can be understood that the embodiment follows the "design-sensing-control-learning" four-link closed loop work. The uppermost layer is the digital twin and process management layer: it receives CAD three-dimensional drawings, BOM material lists and tolerance requirements, and builds a virtual production line in the cloud that synchronizes with the real production line to verify and optimize the assembly process. The middle layer is the edge inference and motion control layer: multi-modal inference models run on industrial computers close to the equipment, fusing "design information" and "on-site sensing information" in real time to calculate the pose (position and angle) and clamping force of the fixture, and send the results to the programmable logic controller (PLC). The bottom layer is the execution layer: composed of collaborative robots, adjustable fixtures, sensor arrays and PLCs, responsible for specific grabbing, handling and insertion actions. The three layers communicate at high speed through gigabit industrial Ethernet and Time-Sensitive Networking (TSN), ensuring sub-millisecond closed-loop communication between instructions and feedback. On the information flow, process data flows from the digital twin layer to the edge layer to guide the assembly strategy; sensing data (vision, force, temperature, inertia, etc.) flows from the execution layer to the edge layer, and then to the cloud for historical analysis; and strategy update flows from the cloud to the edge layer to complete model hot replacement without stopping the line, allowing the assembly strategy to be continuously iterated.

[0087] In some embodiments, the assembly parameters include color images, depth information, three-directional forces, three-directional force moments, position slip amounts of the fixture silica gel shell, and temperature change amounts of the fixture silica gel shell, and the multi-modal sensing module includes:

[0088] An industrial-grade RGB-D camera for collecting color images and depth information;

[0089] A six-dimensional force-displacement sensor for collecting three-directional forces and three-directional force moments;

[0090] A miniature inertial measurement unit for collecting position slip amounts of the fixture silica gel shell;

[0091] A flexible temperature patch for collecting temperature change amounts of the fixture silica gel shell.

[0092] In some embodiments, the multi-modal sensing module includes an industrial-grade RGB-D camera, a six-dimensional force-displacement sensor, a miniature inertial measurement unit, and a flexible temperature patch. Color images and depth information can be acquired by the industrial-grade RGB-D camera. Exemplarily, a pair of industrial-grade RGB-D cameras can be used, with a resolution of 1280x960 pixels and a frame rate of 120 frames per second, which can stably acquire color images and depth information under different lighting environments. The industrial-grade RGB-D camera can be fixed on the end of the mechanical arm and above the work station through an all-aluminum alloy support, and can be matched with a four-point LED strip light to reduce the measurement error of reflective materials. Three-directional force and three-directional torque can be acquired by the six-dimensional force-displacement sensor. Exemplarily, the six-dimensional force-displacement sensor can be installed between the clamp and the flange, which can simultaneously output three-directional force and three-directional torque. The nominal resolution is 0.1 Newton, and the full-scale overload is 150 times without damage, which meets the requirements of soft and hard mixed material compliant assembly. The position slip of the clamp silicone shell can be acquired by the miniature inertial measurement unit, and the temperature change of the clamp silicone shell can be acquired by the flexible temperature patch. Exemplarily, the miniature inertial measurement unit (IMU) can be used to monitor the fine slip, i.e., the position slip of the clamp silicone shell, and the flexible temperature patch can be used to monitor the surface heat change, i.e., the temperature change of the clamp silicone shell. When the temperature or slip inertia deviates from the normal curve, the system is adjusted in time to avoid the silicone shell being rubbed and heated.

[0093] In some embodiments, the cloud service module includes:

[0094] a digital twin cluster, configured to perform assembly simulation based on the assembly parameters by using a digital twin framework to obtain simulation data;

[0095] a cloud server, configured to perform model training based on the simulation data to generate distilled data.

[0096] In some embodiments, the cloud service module includes a digital twin cluster and a cloud server. The simulation data can be obtained by assembling simulation through the digital twin cluster combined with a digital twin framework according to the assembly parameters. Exemplarily, the digital twin framework adopts a Unity-ROS dual-engine, Unity is responsible for high-precision physical rendering and collision detection, and ROS is responsible for robot kinematics and message communication. The simulation data (including strategy and log) obtained by simulation is first stored in object storage, and then pushed to the edge computing module through a version control system. The cloud-edge is encrypted through a VPN tunnel to ensure the security and integrity of the simulation data and model parameters during transmission. The distillation data can be generated by model training through the cloud server according to the simulation data. Exemplarily, the cloud service module is composed of a group of high-performance servers, each cloud server is equipped with multiple professional GPU cards for large-scale parallel simulation and model training to generate distillation data. The cloud server is interconnected through Infiniband, and the reinforcement learning iteration of thousands of virtual assembly tasks can be completed within a few hours. It can be understood that InfiniBand (directly translated as "infinite bandwidth" technology, abbreviated as IB) is a computer network communication standard for high-performance computing, which has extremely high throughput and extremely low delay for data interconnection between computers. InfiniBand is also used as direct or exchange interconnection between servers and storage systems, as well as interconnection between storage systems.

[0097] In some embodiments, the adaptive gripping module includes:

[0098] An adaptive gripper for gripping key components of the humanoid robot;

[0099] A collaborative robot arm for moving the adaptive gripper to a predetermined position.

[0100] In some embodiments, the adaptive gripping module includes an adaptive gripper and a collaborative robot arm. The key components of the humanoid robot can be gripped by the adaptive gripper. Exemplarily, the adaptive gripper can quickly switch modes and combines mechanical fingertips, vacuum suction and electro-permanent magnetic gripping methods. The adaptive gripper is connected to the collaborative robot arm through a quick change interface, and the disassembly and reassembly time is not more than one minute. The adaptive gripper can be moved to a predetermined position by the collaborative robot arm. Exemplarily, the collaborative robot arm has six degrees of freedom and is loaded with a safety torque sensor. The maximum rated torque of a single axis can cover the assembly needs of the joints of the humanoid robot. The joint module integrates a servo drive and an encoder, and the highest repeat positioning accuracy is ± zero point zero three millimeters.

[0101] In some embodiments, the adaptive gripper includes:

[0102] A mechanical finger sub-component, which is composed of a parallel three-finger structure and a micro ball screw, is used to clamp the key components of the humanoid robot through width adjustment;

[0103] A vacuum suction sub-component is used to suction the key components of the humanoid robot through atmospheric pressure;

[0104] An electro-permanent magnetic sub-component is used to suction the key components of the humanoid robot with a metal housing through magnetic force.

[0105] In some embodiments, the adaptive clamp includes a mechanical finger sub-component, a vacuum suction sub-component, and an electro-permanent magnetic sub-component. The mechanical finger sub-component is composed of a parallel three-finger structure and a micro ball screw, and can clamp the key components of the humanoid robot through the mechanical finger sub-component combined with width adjustment. The mechanical finger sub-component can automatically adjust the width within three seconds. The vacuum suction sub-component can suction the key components of the humanoid robot through atmospheric pressure, and the vacuum suction has good fitting effect on carbon fiber or plastic parts. The electro-permanent magnetic sub-component can suction the key components of the humanoid robot with a metal housing through magnetic force, and can specially process metal housing, which can be quickly started and stopped, and has no continuous energy consumption. It can be understood that the metal housing generally refers to the housing or protective cover of the equipment or machine.

[0106] The beneficial effects of implementing the embodiments of the present application include that the multi-modal adaptive clamping system for assembling key components of a humanoid robot provided by the embodiments of the present application includes a multi-modal sensing module, an edge computing module, a cloud service module, an adaptive clamping module, and a programmable logic control module. The multi-modal sensing module is used to collect assembly parameters; the edge computing module is used to generate control instructions using a multi-modal deep learning model according to the assembly parameters; the cloud service module is used to generate distilled data according to the assembly parameters; the adaptive clamping module is used to clamp and move the key components of the humanoid robot; and the programmable logic control module is used to control the adaptive clamping module to perform component assembly operations according to the control instructions and the distilled data, thereby realizing the multi-modal adaptive clamping system and improving the efficiency and clamping stability.

[0107] Figure 2 is an optional flowchart of a control method applied to the system of the embodiments of the present application, Figure 2 The method in can include but is not limited to steps S201 to S204.

[0108] Step S201, acquiring humanoid robot data, the humanoid robot data including a humanoid robot joint CAD model and a bill of materials;

[0109] Step S202, generating a design semantic vector and a virtual production line scene according to the humanoid robot data;

[0110] In step S203, in the virtual production line scene, off-line simulation and distillation are performed according to the design semantic vector to generate a fixture control model and a programmable logic controller ladder diagram.

[0111] In step S204, adaptive fixture control is performed according to the fixture control model and the programmable logic controller ladder diagram to obtain a component assembly record, which includes coaxiality, maximum stress, temperature rise, and retry count.

[0112] The steps S201 to S204 shown in the embodiments of the present application realize adaptive fixture control and improve efficiency and fixture stability.

[0113] In step S201 of some embodiments, humanoid robot data can be obtained from a robot database. Humanoid robot data can also be obtained in other ways, without being limited thereto. The humanoid robot data can be uploaded to a PLM platform, where PLM represents Product Lifecycle Management (PLM).

[0114] In some embodiments, in step S202, a design semantic vector and a virtual production line scene are generated according to the humanoid robot data, which can include but is not limited to the following steps:

[0115] The humanoid robot joint CAD model is parsed and processed using a CAD parser to obtain basic elements and adjacency relationships, the basic elements including faces, edges, and vertices;

[0116] A geometric topology graph is generated according to the basic elements and the adjacency relationships;

[0117] A geometry vector is generated using a graph convolution network according to the geometric topology graph;

[0118] A first semantic vector is generated according to a bill of materials, an assembly sequence, and a torque curve;

[0119] The geometry vector and the first semantic vector are spliced to obtain a design semantic vector;

[0120] A virtual production line scene is generated according to the design semantic vector.

[0121] In some embodiments, in the design semantic parsing layer, the CAD model of the humanoid robot joint can be parsed by a CAD parser to obtain basic elements and adjacency relationships, wherein the basic elements include faces, edges and vertices. Then, according to the basic elements and the adjacency relationships, a geometric topology graph is generated, and the adjacency relationships can be saved in a graph structure. According to the geometric topology graph, a geometric vector is generated by using a graph convolution network. For example, the geometric topology graph is mapped to a 1024-dimensional vector as a geometric vector by a set of lightweight graph convolution networks (GCN), and each sub-section of the vector corresponds to an interpretable feature such as local curvature, hole diameter or tolerance bandwidth. Then, according to the bill of materials, assembly sequence and torque curve, a first semantic vector is generated. For example, the materials, heat treatment methods and stress limits in the BOM (Bill of Material) table are converted into a 256-dimensional semantic vector by a pre-trained text encoder (BERT), and combined with process data such as assembly sequence and torque curve to encode the first semantic vector. The geometric vector and the first semantic vector are spliced to obtain a design semantic vector. Through this "double-track coding", the system not only saves spatial information, but also carries engineering constraints, realizing the design semantics of "readable, calculable and traceable". Furthermore, the assembly sequence and torque curve can be written into a YAML process file as priori for subsequent reinforcement learning. The process file and the geometric topology graph Figure 1 Understandably, YAML is a high-readability format used to express data serialization. Finally, according to the design semantic vector, a virtual production line scene is generated, which can create a virtual production line scene corresponding to the real production line, providing a high-fidelity environment for subsequent offline simulation and strategy verification. Understandably, the design semantic parsing layer is responsible for disassembling the CAD model into a "geometric topology graph", and then encoding it into a high-dimensional vector together with the BOM tolerance and material information. In this way, "what to do" can be clearly passed to the subsequent algorithm. The goal is to reliably project the research and development intention into the algorithm world.

[0122] In some embodiments, in step S203, according to the design semantic vector, offline simulation and distillation are performed to generate a fixture control model and a programmable logic controller ladder diagram, which can include but is not limited to the following steps:

[0123] According to the design semantic vector, incoming material tolerance, temperature and humidity, and assembly posture, a model training is performed by using reinforcement learning to obtain a fixture control model;

[0124] The fixture control model is compressed by using a contrast distillation strategy to obtain lightweight network weights;

[0125] According to the lightweight network weights, a programmable logic controller ladder diagram is generated.

[0126] In some embodiments, in a virtual production line scenario, a high-robustness clamp control model can be obtained by using reinforcement learning to perform model training according to the design semantic vector, incoming tolerance, temperature and humidity, and assembly posture. After training, the clamp control model is compressed using a contrast distillation strategy to obtain light-weight network weights, and then a programmable logic controller ladder diagram (PLC ladder diagram) is generated according to the light-weight network weights. If the verification is passed, the clamp control model is pushed to the edge computing box and the on-site PLC with a version number.

[0127] In some embodiments, in step S204, adaptive clamping control is performed according to the clamp control model and the programmable logic controller ladder diagram to obtain a component assembly record, which can include but is not limited to the following steps:

[0128] After calibrating the multi-modal sensing module, the assembly parameters are collected;

[0129] The assembly parameters are aligned and fused using a cross-modal multi-head attention network to obtain an assembly semantic vector;

[0130] According to the assembly semantic vector, an instruction parameter is generated using a graph neural network and a clamp control model, and the instruction parameter includes a posture adjustment amount, a clamping force size, an insertion and assembly speed curve, and a motion trajectory;

[0131] According to the instruction parameter and the programmable logic controller ladder diagram, the adaptive clamping module is controlled to perform a component assembly operation to obtain a component assembly record.

[0132] In some embodiments, the assembly parameters can be collected after calibrating the multi-modal sensing module. For example, before the production line starts, an engineer performs a no-load calibration: the industrial-grade binocular RGB-D camera completes coordinate calibration on a target plate, the six-dimensional force-displacement sensor is zeroed on a standard test block, and the miniature inertial measurement unit and flexible temperature patch record the baseline values. The edge computing box runs a self-checking script, and if the prediction-execution error, clock drift, or network latency exceeds the limit, it will prevent the start. Then the collaborative robot tries to grab a few sample pieces, and the model compares the theoretical and measured force curves. If the deviation exceeds the threshold, the clamping parameters are adjusted online until the first piece is released. The multi-modal perception layer can be used to collect assembly parameters, including visual, force, temperature, and inertial signals. The visual encoder uses an improved Swin Transformer: a layer of deep separable convolution is added before the basic window attention structure, so that micro features such as edges and holes can be extracted at a low resolution stage. The force and temperature signals enter a one-dimensional convolutional network + gated recurrent unit (GRU) to capture the dynamic evolution of the three stages of “micro-jitter-pressing-steady state” during assembly. The inertial measurement unit (IMU) data is filtered by Kalman filter and then stacked into the same GRU to allow the model to perceive micron-level slip. Finally, these heterogeneous features are resampled to the same beat and spliced into a “perception vector” of length 768. To prevent perception noise from drifting over time, the system uses self-supervised contrastive learning: raw signals are collected at the fragment level every day, and the “same station, same time period” is used as the positive sample, and the “cross-station, cross-time period” is used as the negative sample, to continuously update the encoder weights and ensure the stability of the feature distribution.

[0133] Then the assembly parameters are aligned and fused using a cross-modal multi-head attention network to obtain an assembly semantic vector. For example, the object shape and pose can be extracted by the visual network, the force change during clamping can be analyzed by the force network, and the friction or slip signs can be monitored by the temperature and inertial network. A cross-modal Transformer (multi-head attention network) is used to align the above features with the design vector to generate a unified assembly semantic vector. It can be understood that cross-modal fusion is the core of the entire algorithm stack. The system uses a 12-layer multi-head attention Transformer, with 16 heads per layer, and each attention head receives both “design vectors” (size, material, tolerance) and “perception vectors” (shape, force-temperature, inertia). In order to align the time scales of different modalities, a “dynamic time delay compensation block” is inserted after each attention. This module will automatically align the visual frames and force sampling in the feature space by learning a set of differentiable shift parameters based on the correlation coefficient of the previous attention. After fusion, an assembly semantic vector of length 512 is output. This step does not directly give motion instructions, but compresses implicit information such as “how to grab, how much force, and where to insert” into a unified semantic space for downstream decision networks to consume.

[0134] According to the assembly semantic vector, the instruction parameters are generated by using the graph neural network and the fixture control model, and the instruction parameters include the pose adjustment amount, the clamping force size, the insertion speed curve and the motion trajectory. Exemplarily, the decision and motion planning layer can be used for calculation, and the decision layer simultaneously considers the "geometric contact relationship" and the "dynamic mechanical feedback". The system abstracts the mechanical fingertips of the fixture, the parts to be assembled and the reference hole positions into nodes of a graph, and the edge weights between the nodes are initialized by the assembly semantic vector. Subsequently, a set of three-layer graph neural networks (GNN) performs information propagation on the contact graph, and outputs three continuous action parameters including the end pose compensation, the clamping force setting and the insertion speed curve. The end pose compensation refers to the x, y and z displacements and the attitude angle which need to be fine-tuned by the robot arm. The clamping force setting refers to the target force values of the three channels of the fingertips, the vacuum and the magnetic force. The insertion speed curve refers to a speed-time table of the stages of increasing, uniform and decreasing. The action parameters are sent to a policy head based on reinforcement learning, and the policy head uses the proximal policy optimization algorithm (PPO). The reward function comprehensively considers the coaxiality error, the maximum stress, the assembly time and whether secondary alignment is needed, so as to ensure that the model pursues the beat while not excessively risks. Periodically, the cloud uses digital twinning to simulate rare extreme working conditions, and performs "risk rollback training" on the policy head, so that the model learns to actively slow down or evacuate when an abnormality occurs. Finally, the graph neural network (GNN) combines the reinforcement learning strategy to output three sets of instruction parameters, including the pose adjustment amount (position offset and angle offset), the clamping force size (Newton, which has considered different materials) and the motion trajectory (the path of the robot arm from the current point to the target point). The insertion speed curve can also be included.

[0135] Finally, according to the instruction parameters and the programmable logic controller ladder diagram, the adaptive gripping module is controlled to perform the component assembly operation, and a component assembly record is obtained. Exemplarily, the first piece assembly can be performed, or after entering mass production, the camera captures the shape and pose of the part, and the edge computing box fuses vision, force, temperature and design semantics within about 20 milliseconds to output the pose adjustment amount, the clamping force size and the insertion speed curve. The PLC drives the robot arm to execute at a frequency of 1 kHz, and the force-position-temperature-inertia signals are fed back in real time. If an abnormal stress or pose drift is detected, the model immediately calculates a retreat path and performs secondary alignment, so as to ensure that the assembly is compliant and damage-free.

[0136] In some embodiments, the method further comprises:

[0137] When online quality monitoring is performed, the fixture control model is fine-tuned according to the component assembly record, and the fixture control model is updated;

[0138] When full-day data aggregation is performed, the fixture control model is aggregated according to the component assembly record and the extreme working condition simulation result, and the fixture control model is updated;

[0139] When model change is performed, the fixture control model is updated according to the component assembly record and the small-scale simulation result.

[0140] In some embodiments, model training can be performed using an online self-learning layer. The cloud can simulate various assembly conditions in batches using digital twins, distill offline experience into lightweight model weights, and collect real-time "success / failure" labels at the edge. The model weights are sent to the PLC through container hot updates, and the entire process does not require downtime or manual re-teaching.

[0141] When online quality monitoring is performed, the fixture control model can be parameter-tuned according to the component assembly record, and the fixture control model is updated. For example, after completing each assembly, the system records indicators such as coaxiality, maximum force, temperature rise, and retry count and writes them into a quality database. The edge computing box aggregates "success / failure" samples on an hourly basis, performs small-step gradient updates, and only tunes the decision layer and the time delay compensation block, thereby quickly adapting to batch errors without damaging the main features.

[0142] When full-day data aggregation is performed, the fixture control model can be parameter-aggregated according to the component assembly record and the extreme condition simulation result, and the fixture control model is updated. For example, at night every day, the cloud compares the production log of the day with the digital twin scenario again, supplements the extreme condition simulation result, and generates a new version of incremental weights. The new weights are first released in gray to a single robot for testing for three hours; if the pass rate improves or remains the same and no alarm is triggered, the new weights are automatically promoted to the entire line; otherwise, the weights are rolled back to the previous stable version and the rollback reason is recorded. The entire release or rollback process takes about three minutes and does not require downtime.

[0143] When model change is performed, the fixture control model is updated according to the component assembly record and the small-scale simulation result. For example, when the research and development department uploads a new joint design file, the digital twin platform automatically parses and runs a small-scale simulation to generate a small-scale simulation result. The edge computing box completes the tuning within the first ten trial assemblies, and if the first assembly passes, it can directly enter mass production. In this way, "design change goes online" can be achieved on the same flexible production line without the need to replace fixtures or manually teach or stop for a long time.

[0144] The closed-loop cycle of this embodiment realizes no-stop-line quick change, sub-millimeter level alignment accuracy and millinewton level compliant force control, providing an intelligent solution for sustainable evolution of humanoid robot multi-model and high iteration assembly. More specifically, in real-time production, the model records the difference between "actual action" and "planned action" as a playback segment and labels it as "success / failure". The edge box aggregates the data of the last twenty-four hours every day and performs a small-scale gradient update: this step only adjusts the decision head and time delay compensation block, keeping the encoder and Transformer backbone stable to avoid catastrophic forgetting. The updated weights are time-stamped and work-station-numbered by the version control system, and then go through the gray release process: first, small batch trial operation on a single robot arm, if the qualified rate does not decrease but increases, then gradually promote the new model to the whole line. If a rollback threshold occurs during the gray period (for example, three consecutive assembly failures), the scheduler will immediately switch back to the previous stable version and record the rollback reason for subsequent offline analysis. The entire software stack runs in a real-time Linux container and relies on GPU virtualization technology to serve multiple production lines simultaneously. Each production line has an independent "model sandbox" to ensure that on-line fine-tuning does not interfere with other workstations; the cloud acts as a "knowledge hub" and continuously extracts the most valuable strategy increment from each sandbox to benefit the next round of offline training. Through this "cloud fusion-edge self-learning" two-way closed loop, the system can continuously evolve under no-stop-line conditions to meet the long-term needs of humanoid robot assembly "multi-model, high precision, and rapid iteration".

[0145] In some embodiments, to verify the feasibility of the method of the present embodiment in the assembly scene of the elbow joint module of a humanoid robot, an experiment can be conducted to build a single-station verification line in a trial production workshop. The experimental environment is as follows.

[0146] (1) Workstation and equipment configuration. Collaborative robot: 1x UR10e (6-DOF, working radius 1300mm, repeatability ±0.03mm), end-mounted self-developed three-finger-vacuum-electropermanent magnet integrated gripper; maximum opening of fingers 120mm, clamping force 0-200N continuously adjustable. Force-position sensor: ATI Mini45-E (resolution 0.01N / 0.5mN-m, sampling 1kHz), mounted between gripper and flange. Vision system: dual-eye Intel Realsense D435 camera (1280x720@120fps), one fixed on top of workstation looking down, one moving with the robot end. Industrial PLC: Beckhoff CX5140 embedded controller (Intel Atom E3845 1.9GHz, EtherCAT master), cycle 1ms. Edge computing box: NVIDIA Jetson AGX Orin (2048 CUDA core, 64 Tensor core, 22TOPS, Ubuntu 22.04 RT Patch), connected to EtherCAT clock domain through TSN dual-network-port. Temperature & IMU patch: flexible NTC array + Bosch BMI270 (1kHz sampling), adhered to the surface of the aluminum alloy shell of the elbow joint and the silicone protective layer.

[0147] (2) Network and time synchronization. Motion layer bus: EtherCAT 100Mb / s, PLC master cycle 1ms; robot and force sensor are slave stations. Data layer network: Cisco IE-4000 TSN switch, 1Gb / s, uplink to edge box; all nodes enable IEEE 1588 PTP, clock drift controlled within 0.8μs. Cloud connection: independent VPN channel for the laboratory, upload production logs and model weights; download bandwidth 200Mb / s.

[0148] (3) Software and algorithm stack.

[0149]

[0150] (4) Sample and evaluation index. Sample: 2025-A type elbow joint module, weight 1.2kg, outer shell magnesium-aluminum alloy + local silicone coating; paired shoulder reference hole diameter 25mm, tolerance IT7. Batch: a total of 500 pieces, randomly introduced ±2mm pose deviation and ±0.03mm hole diameter deviation. Core indicators: △ coaxiality ≤0.04mm (90% percentile); maximum clamping force peak ≤85N (silicone surface must not appear indentation); first-time insertion success rate ≥95%; workstation beat 11s (including alignment, insertion, clamping release, and reset).

[0151] (5) Environment control and safety. The experimental workshop is kept at a constant temperature of 23±2°C and humidity of 45±5% RH. The safety torque of the robot arm is limited to 100 N-m, and the overload and force-position anomaly of the clamp trigger the secondary EMS shutdown. All data are encrypted and stored in the local NAS, and are mirrored to the cloud object storage every day.

[0152] During the experiment, the clamp control operation process is as shown in Figure 3 , which can first call the Unity-ROS virtual production line in the cloud digital twin platform, and perform 20,000 off-line training on the 2025-A elbow joint under the conditions of ±2mm pose disturbance and ±0.03mm hole diameter tolerance. After the reinforcement learning converges, the distilled 22MB lightweight weight file (version number cfs-elbow-v5.2) is packaged with the matching PLC ladder diagram into a "strategy package". The edge computing box pulls the strategy package through VPN, and the container hot loads the new weight; the PLC calls the TwinCAT "Online Change" function to complete the ladder diagram hot replacement in the idle period, which takes about 45s. After loading, the model SHA-256 checksum and version number are recorded to ensure traceability in the future.

[0153] Then the unloaded system calibration is performed, including: vision calibration: fix the camera to align with the cross target board, and run the ROScamera_calibration node to obtain the internal and external parameter matrix; the end camera is aligned with the station coordinate system through four-point ARUCO code. Force-position zero point calibration: the robot moves to "zero attitude", and the clamp is not pressurized. The ATIMini45 automatically collects 2000 frames of unloaded data, calculates the average value and writes it into the zero bias compensation register. Temperature and inertia reference: the flexible NTC array and BMI270 collect 30s environmental baseline respectively; the edge box caches the baseline sequence in Redis, and subsequently judges the temperature rise and vibration anomaly. After calibration, the system executes the self-check script: if the vision re-projection error is >0.15px, the force sensor zero drift is >1N, or the network latency is >0.5ms, then prevent entering the next step.

[0154] Then the first piece trial assembly and model fine tuning are performed. The collaborative robot picks up 10 trial samples one by one for assembly, and after the assembly of each piece is completed, the edge box aligns the "assembly semantic vector" with the actual force-position-temperature curve, calculates the coaxiality error and the maximum clamping force. If the coaxiality error is ≤0.04mm and the maximum stress is ≤85N, then mark "success"; otherwise, trigger online gradient update, set the learning rate of the decision head to 1e-5, and after 200 iterations, try again. When 3 consecutive pieces are marked as successful, the First-Article-Pass report is generated, and the fine-tuned weight is saved as cfs-elbow-v5.2.1. The entire first piece stage takes 18 minutes, and there are 2 online updates.

[0155] Batch assembly and real-time closed-loop processing are performed, the edge box switches to production mode, and the MES system issues a 500-piece production task: when the parts to be assembled are in place, the overhead camera sends a / part_ready trigger signal; the end camera synchronously acquires the local shape. The multi-modal fusion model outputs within 22 ms: end pose compensation Δx, Δy, Δz, ΔRoll, ΔPitch, ΔYaw; clamping force setting 65N (for aluminum alloy surface) / 40N (for silicone surface); insertion speed curve: acceleration 0.5 m / s 2 → uniform speed 0.2 m / s → deceleration 0.6 m / s 2 . The PLC updates the servo every 1 ms: if the force-position sensor detects a transient force peak > 90N or torque direction reversal, trigger the "yield path", the robot retreats 5mm and repositions. After insertion is completed, the clamp is released, and the collaborative robot returns to the Home position; the system synchronously records the assembly log entries and pushes them to Kafka.

[0156] Online quality monitoring is performed during assembly, and the quality monitoring service subscribes to the Kafka stream and calculates in real time: coaxiality distribution: box plot output every 50 pieces, if the 90% percentile > 0.04mm triggers a yellow light warning. Clamping force peak trend: rolling window 100 pieces, linear regression slope > 2N / 100 pieces triggers maintenance prompt. Temperature rise curve: local temperature higher than baseline +3℃ for 10s triggers production suspension and prompts to check the silicone coating. The edge box uses success / failure tags to make fine adjustments with a 60min cycle, fine adjustments only update the last layer of fully connected and time delay compensation blocks, which takes <45s and does not affect the beat.

[0157] Finally, cloud distillation and gray release are performed, after production is completed, the edge box uploads complete logs and the latest weight v5.2.1-edge. The cloud supplements real data to the simulation environment in the night task, re-trains 2h for extreme pose and temperature rise scenarios, and generates an incremental version v5.3-candidate. The next day, the scheduler puts v5.3-candidate into gray release to this station: assemble 25 pieces according to a 5% sampling ratio; if the qualified rate ≥95% and the average coaxiality does not deteriorate, automatically release 100%; if not up to standard, roll back to v5.2.1 and generate a rollback report. In this experiment, v5.3-candidate successfully passed the gray release, the entire line switching time was 2min37s, and there was no need to stop the line.

[0158] In some embodiments, advantages of the present embodiments include:

[0159] (1) The changeover time is almost zero. Traditional production lines must replace special fixtures, recalibrate, and stop for 30-60 minutes; the present system relies on multi-modal algorithms and PLC ladder diagram hot replacement, and only container loading and coordinate self-checking are needed to switch to a new model, the entire process takes less than 3 minutes, and the impact on production line productivity is negligible.

[0160] (2) Assembly precision and compliance force control. In a batch of 500 experiments, the coaxiality of the 90% percentile is maintained at 0.037 mm, and the maximum clamping force peak is 82 N, which is better than the traditional rigid fixture (commonly 0.07-0.10 mm, >100 N). The silicone-coated surface does not appear to be indented or silicon-cracked, proving that multi-modal closed-loop can achieve millinewton force control and sub-millimeter positioning under hard-soft mixed materials.

[0161] (3) First-time assembly success rate significantly improved. Benefiting from the immediate "yielding-second alignment" strategy, the first-time insertion success rate is improved from 86-88% of manual demonstration line to 96.8%, and the rework rate is reduced to less than 2%, saving about 15% of labor and material costs per month.

[0162] (4) Continuous self-learning, yield climbing over time. Through the closed loop of "cloud distillation + edge fine-tuning + gray release", the model weight can be iterated without stopping the line; in the experiment, from v5.2 to v5.3, the yield is improved by 1.3 percentage points, and there is no manual demonstration or parameter tuning. This capability enables the production line to automatically "evolve" with design improvements and incoming material fluctuations, avoiding the need for repeated on-site adjustments by engineers in traditional solutions.

[0163] (5) Fixture and tooling assets halved. A single adaptive fixture combines mechanical fingertips, vacuum, and electro-permanent magnet modes, covering multiple materials and sizes; the same production line shares cameras, force sensors, and edge boxes, reducing fixture procurement and maintenance costs by about 40-50% compared to the "model-fixture one-to-one" mode.

[0164] (6) Data traceability and rapid attribution. Multi-modal data (design, vision, force-position, temperature, IMU) is archived at the piece level, and defects can be located to a certain clamping or trajectory within seconds; compared to traditional visual inspection and intermittent sampling, traceability efficiency is improved by an order of magnitude, laying the foundation for subsequent large-scale quality management and AI diagnosis.

[0165] (7) Universality and generalizability. The algorithm framework does not depend on specific brand robotic arms or PLCs, and can be transplanted to UR, ABB, KUKA, and EtherCAT, Profinet networks; multi-modal features and reward functions only need to adjust weights to adapt to shoulder joints, grippers, battery packs, and other different assembly units, achieving "the same technology stack, covering the entire machine".

[0166] In some embodiments, the embodiment implements a design-sensing bidirectional semantic mapping mechanism, which parses three-dimensional CAD models and BOM process files into geometric topology graphs according to face, edge, vertex and material attributes, and then extracts high-dimensional design semantic vectors by using graph convolution and text encoder; at the same time, multi-modal signals such as RGB-D, force-temperature, inertia and the like collected in real time on site are uniformly encoded into the same dimension of perception vectors. After alignment in the same feature space, the algorithm can inherit both engineering priori and instantaneous state, and provide accurate and traceable "semantic coordinate system" for subsequent decision-making.

[0167] The embodiment implements a cross-modal Transformer fusion network and dynamic time delay compensation. A multi-head attention Transformer performs layer-by-layer fusion on design and perception vectors, and a learnable delay compensation module is inserted after each layer to automatically adjust the time alignment of different modalities according to the previous attention weight, and eliminate the misalignment caused by the difference in visual and force sampling frequency. The network outputs a unified assembly semantic vector within twenty milliseconds, realizing high-precision and low-delay information fusion.

[0168] The embodiment implements an adaptive assembly decision-making framework based on graph neural network and proximal policy optimization. The algorithm abstracts the fixture fingertips, parts to be assembled and reference hole positions into nodes, and the edge weights between the nodes are derived from the semantic vectors, which are passed to a three-layer graph neural network to transfer contact topology information; then, proximal policy optimization (PPO) is used to generate end position compensation, clamping force setting and insertion speed curve. When the incoming material tolerance, material hardness or assembly posture changes, the decision-making framework can automatically adjust the grab-insert-release strategy to realize compliant and stable assembly.

[0169] The embodiment implements a three-layer self-learning closed loop of cloud distillation-edge fine-tuning-gray release. The system first performs large-scale reinforcement learning in the digital twin environment, and then distills the model into lightweight weights and releases it to the production line; the edge end fine-tunes a small number of parameters combined with real success / failure samples on an hourly basis; the updated model is released or quickly rolled back in a gray manner, completely without stopping the line or manual demonstration, thereby ensuring continuous ramping of the production line without interruption.

[0170] The embodiment implements online hot replacement of PLC ladder diagram and clock consistency strategy. The edge AI reasoning result is converted into a hot-loadable ladder diagram segment by safety constraint, and is replaced seamlessly by the TwinCAT online change interface; all nodes are kept synchronized at the microsecond level through TSN-PTP clock domains, which meets the requirements of industrial hard real-time and avoids the downtime risk of traditional version switching.

[0171] The embodiment realizes a mechanical fingertip-vacuum-electromagnetic three-mode fast switching self-adaptive clamp, which is internally integrated with a ball screw automatic width adjusting mechanism, a double-loop vacuum channel and a matrix type permanent magnet, can complete clamping mode switching within a single action instruction, and can complete clamp replacement within one minute through blind insertion quick change interface. The algorithm can continuously adjust the clamping force in the range of 0-200 Newton, and takes into account the rigid fixation of the metal shell and the flexible protection of the silica gel sheath.

[0172] The embodiment realizes a multi-modal safety monitoring and abnormal self-recovery strategy, which simultaneously monitors force-position, temperature rise and inertia signals, adopts redundant threshold judgment and trend analysis double insurance; if over force, slip or abnormal temperature rise is detected, the decision layer generates a retreat path in milliseconds and triggers secondary alignment to prevent pressure injury, biting and wire winding, and realizes active safety in the assembly process.

[0173] The embodiment realizes a piece-level assembly full data tracing framework, which archives design semantic vector, action parameter, actual multi-modal curve and qualification judgment for each part, realizes the whole process data chain of one thing one code. When defects occur, the corresponding clamping action or force peak abnormality can be located in seconds, providing high-value training data for engineering analysis, quality improvement and subsequent artificial intelligence diagnosis.

[0174] The beneficial effects of implementing the embodiment of the present application include that the embodiment of the present application first acquires humanoid robot data, then generates a design semantic vector and a virtual production line scene according to the humanoid robot data, then performs offline simulation and distillation according to the design semantic vector to generate a clamp control model and a programmable logic controller ladder diagram, and finally performs adaptive gripping control according to the clamp control model and the programmable logic controller ladder diagram to obtain a component assembly record, thereby realizing adaptive gripping control and further improving efficiency and gripping stability.

[0175] As shown in Figure 4 The embodiment of the present application also provides a computer device, which comprises:

[0176] at least one processor 901;

[0177] at least one memory 902 for storing at least one program;

[0178] When the at least one program is executed by the at least one processor, the at least one processor implements the method shown in Figure 2 .

[0179] The contents in the above method embodiments are all applicable to the device embodiments, the device embodiments specifically realize the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.

[0180] The embodiment of the present application also provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the method shown in the embodiment of the present application. Figure 2 The embodiment of the present application also provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the method shown in the embodiment of the present application.

[0181] The content in the method embodiments is applicable to the storage medium embodiments, the storage medium embodiments specifically realize the same functions as the method embodiments, and the beneficial effects achieved by the storage medium embodiments are also the same as the beneficial effects achieved by the method embodiments.

[0182] The preferred embodiments of the present application are described above with reference to the accompanying drawings, and the scope of the present application is not limited by the above description. Any modification, equivalent replacement and improvement made by those skilled in the art without departing from the scope and essence of the present application should be within the scope of the present application.

Claims

1. A control method for a multi-modal adaptive grasping system for humanoid robot key component assembly, characterized in that, The multi-modal adaptive gripping system comprises: a multi-modal sensing module for collecting assembly parameters; an edge computing module for generating control instructions using a multi-modal deep learning model according to the assembly parameters; a cloud service module for generating distilled data according to the assembly parameters; an adaptive gripping module for gripping and moving key components of a humanoid robot; a programmable logic control module for controlling the adaptive gripping module to perform component assembly operations according to the control instructions and the distilled data; The control method comprises the following steps: obtaining humanoid robot data, which includes a humanoid robot joint CAD model and a bill of materials; generating a design semantic vector and a virtual production line scene according to the humanoid robot data; in the virtual production line scene, performing offline simulation and distillation according to the design semantic vector to generate a gripper control model and a programmable logic controller ladder diagram; performing adaptive gripping control according to the gripper control model and the programmable logic controller ladder diagram to obtain component assembly records, which include coaxiality, maximum stress, temperature rise, and retry count.

2. The method of claim 1, wherein, The method further comprises: ​ ​ ​ ​ ​ ​ 3. The method of claim 1, wherein, ​ ​ ​ ​ 4. The method of claim 1, wherein, ​ ​ ​ ​ ​ 5. The method of claim 1, wherein, ​ When online quality monitoring is performed, the fixture control model is parameterized according to the component assembly record, the fixture control model is updated; When full-day data aggregation is performed, the fixture control model is parameterized according to the component assembly record and extreme working condition simulation results, and the fixture control model is updated; When model conversion is performed, the fixture control model is updated according to the component assembly record and small-scale simulation results.

6. The control method according to claim 1, characterized by, The assembly parameters include color images, depth information, three-directional forces, three-directional force moments, position slip amounts of fixture silica gel shells, and temperature change amounts of fixture silica gel shells, and the multi-modal sensing module includes: Industrial grade RGB D camera for capturing color images and depth information; Six-dimensional force Displacement sensors for collecting three-way force and three-way moment A miniature inertial measurement unit is used to collect the position slip amounts of the fixture silica gel shells. A flexible temperature patch is used to collect the temperature change amounts of the fixture silica gel shells.

7. The control method according to claim 1, characterized by, The cloud service module includes: A digital twin cluster is used to perform assembly simulation by using a digital twin framework according to the assembly parameters, and simulation data is obtained; A cloud server is used to perform model training to generate the distillation data according to the simulation data.

8. The control method according to claim 1, characterized by, The adaptive gripping module includes: An adaptive fixture is used to grip key components of a humanoid robot; A collaborative mechanical arm is used to move the adaptive fixture to a predetermined position.

9. The control method according to claim 8, characterized by, The adaptive fixture includes: A mechanical fingertip subcomponent composed of a parallel three-finger structure and a miniature ball screw is used to grip the key components of the humanoid robot by width adjustment; A vacuum adsorption subcomponent is used to adsorb the key components of the humanoid robot by atmospheric pressure; An electro-permanent magnetic subcomponent is used to adsorb key components of the humanoid robot with metal shells by magnetic force.

Citation Information

Patent Citations

  • Industrial robot control method and device based on multi-modal large model

    CN117656082A

  • Large model and small model collaborative robot operation action real-time control method and system

    CN119567267A