Complex spaceflight product man-machine cooperation assembly-oriented body-equipped agent packaging method and device

Through the embodied intelligent agent encapsulation method, combined with the multimodal perception and automatic code generation of a large language model, the problems of difficult system integration and insufficient task adaptability in the human-machine collaborative assembly of complex aerospace products were solved, and the adaptability and execution accuracy of the assembly process were improved.

CN120791810AActive Publication Date: 2025-10-17NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Application Number
CN202511309303.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2025-10-17
Estimated Expiration
2045-09-15

AI Technical Summary

Technical Problem

Existing technologies in the human-machine collaborative assembly of complex aerospace products have problems such as low system integration efficiency, insufficient task adaptability and poor dynamics of human-machine collaboration. Especially when faced with unstructured information and sudden working conditions, it is difficult to meet the needs of efficient and flexible assembly.

Method used

By adopting the embodied intelligent agent encapsulation method, by building a mapping relationship between the embodied intelligent agent and the physical robot, and combining it with a large language model to enhance perception, reasoning and execution capabilities, multimodal perception, scene graph analysis and automatic code generation are achieved, thereby improving system integration efficiency and task adaptability.

Benefits of technology

It has achieved improvements in the modularization of system integration, task adaptability and execution accuracy in the human-machine collaborative assembly of complex aerospace products, and solved the problems of high coupling between hardware equipment and software algorithms, deviation in task semantic understanding and low degree of automation in the assembly process in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120791810A_ABST
    Figure CN120791810A_ABST
Patent Text Reader

Abstract

The invention discloses a complex aerospace product man-machine cooperation assembly-oriented body-equipped agent packaging method and device. The method comprises the following steps: a sensing layer realizes multi-modal intention and environment sensing by using a large language model; the reasoning layer analyzes a structured assembly plan sequence through an assembly scene map; and the execution layer converts the target coordinate into a rotation angle to generate an action code. According to the method, the assembly plan sequence is generated according to the sensing result and is automatically converted into the robot execution code through the combination of the body agent packaging and the large language model, so that the self-adaptive capability of the man-machine cooperative assembly process is improved; therefore, the problems of high system integration difficulty, insufficient task adaptability and high-efficiency cooperation difficulty in man-machine cooperation assembly of complex aerospace products are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of human-robot collaborative assembly, and particularly relates to a body intelligent agent packaging method and device for human-robot collaborative assembly of complex aerospace products. BACKGROUND

[0002] The assembly process of complex aerospace products (such as launch vehicles, satellite payloads, etc.) has the significant characteristics of high precision requirements, multi-disciplinary intersection, complex component configuration, and frequent dynamic collaboration, and the assembly quality directly affects the reliability of the aerospace mission. With the increasing demand for intelligent manufacturing of aerospace products, the human-robot collaborative assembly mode has become the mainstream, which requires the robot system to not only accurately perform repetitive operations, but also dynamically collaborate with human operators in unstructured environments. Under this background, building an intelligent assembly system with environmental perception, task reasoning and precise execution capability has become a key technical direction to break through the efficiency and quality bottlenecks of aerospace product assembly.

[0003] Currently, the research on human-robot collaborative assembly mainly focuses on the following directions. In terms of robot perception, existing methods use RGB-D cameras and UWB positioning technologies to build an assembly scene model, but the understanding of personnel behavior intention is limited to action trajectory recognition, lacking semantic level analysis capability. In terms of robot reasoning, existing methods divide tasks by predefining human-robot collaboration rules, which heavily relies on static rule library. In the face of sudden working conditions (such as component position deviation exceeding tolerance, temporary tool replacement), the system reconfiguration time is as long as 30 minutes or more, which is difficult to meet the real-time requirements of aerospace assembly. In terms of robot execution, existing methods achieve basic assembly operations through pre-set control logic, which can only handle rigid assembly tasks at pre-set workstations, and has weak adaptability to flexible and dynamic assembly tasks of complex aerospace products.

[0004] Therefore, the existing technology exposes three major core defects in actual application: (1) Due to the lack of effective packaging method, the system integration efficiency is low, and each layer module adopts customized interface, resulting in high coupling degree of hardware devices and software algorithms, and new device access requires redeveloping adaptation programs, which is difficult to meet the assembly requirements of rapid switching of multiple models of aerospace products. (2) The task reasoning capability is lacking, and traditional systems rely on manual preset assembly logic (such as static assembly sequence based on CAD model), which cannot automatically analyze the semantics of tasks in the face of unstructured information in the assembly process (such as oral instructions of operators, temporary design changes). (3) The human-robot collaboration is poor in dynamics, and the robot executes tasks based on pre-set assembly programs, with poor self-generation and adaptability of code. SUMMARY

[0005] To solve the above problems, the application provides a body intelligent agent packaging method and device for complex aerospace product human-machine collaborative assembly.

[0006] Technical scheme: A body intelligent agent packaging method for complex aerospace product human-machine collaborative assembly, comprising the following steps:

[0007] A body intelligent agent packaging method for complex aerospace product human-machine collaborative assembly, characterized by comprising the following steps:

[0008] S100: A mapping relationship between the body intelligent agent and the physical robot is constructed, the body intelligent agent is characterized and packaged, and contains a perception layer, an inference layer and an execution layer; an embedded controller is used as an integrated interface, and a communication channel between the physical robot and the virtual body intelligent agent is constructed;

[0009] S200: In the perception layer, the perception ability of the body intelligent agent is enhanced based on a large language model, scene features are extracted through a ResNet-50 visual encoder, a language vector is fused to generate a multi-modal representation through a Hadamard product, the assembly environment semantics are analyzed, personnel assembly behavior intention perception and assembly environment part position perception are realized;

[0010] S300: In the inference layer, the assembly task is described by a scene graph, and is converted into a structured collaborative robot plan sequence through a large language model;

[0011] S400: In the execution layer, the plan sequence is converted into robot control code through a hand-eye calibration model of the large language model and an inverse kinematics algorithm, and finally the action of the collaborative robot is generated to realize the human-machine collaborative assembly task.

[0012] Preferably, in step S100, the body intelligent agent is packaged by analyzing the structure design of the physical collaborative robot, including a double mapping mechanism of physical entity mapping and internal function mapping; the physical collaborative robot is equipped with an embedded controller, which collects sensor data through a preset communication protocol; in the operating system of the embedded controller, an independent thread is created for each body intelligent agent as a basic unit for the operation of the body intelligent agent, an execution function API of joint control / grasp release is provided, the current state information of the robot is obtained based on the execution function API, sensor data is read and the running state of the robot is collected in real time, a dynamic link library encapsulates general functions for interaction between the embedded controller and hardware devices, data processing algorithm functions and communication protocol processing functions; when multiple body intelligent agents work cooperatively, the threads of the intelligent agents realize task allocation and coordination through the scheduling mechanism of the embedded controller, and realize cooperative scheduling through shared memory / message queues.

[0013] Preferably, in step S100, the dynamic link library adopts a standardized interface design, and the encapsulated general functions include device initialization, data reading and writing, and state monitoring functions; the data processing algorithm function set integrates filtering and noise reduction, coordinate conversion, and kinematics inverse solution algorithm; the communication protocol processing function supports Modbus, CAN, and EtherCAT industrial bus protocol analysis, providing a unified hardware interaction interface for embodied agents;

[0014] When the system function is expanded, by creating a new embodied agent thread on the embedded controller, or modifying the execution function API and dynamic link library in the existing thread: the new function module realizes hardware interaction by calling the basic functions of the dynamic link library, and the algorithm logic that needs to be expanded is realized by adding data processing functions or modifying the existing algorithm parameters.

[0015] Preferably, step S200 includes:

[0016] S210: A ResNet-50 visual encoder is used to extract hierarchical visual representations from scene images, and a deep residual architecture is used to capture spatial context details in a cluttered workspace, forming high-order visual features for subsequent semantic segmentation;

[0017] S220: The high-order visual features and language vectors are fused through a Hadamard product-based fusion layer to generate discriminative multi-modal representations with visual spatial specificity and language semantic accuracy;

[0018] S230: For occluded and blurred inputs, an error auxiliary guidance mechanism is constructed, the differences in the semantic segmentation results are identified through a confidence analysis system, and iterative cross-modal recalibration is performed based on the multi-modal representations to reduce semantic misunderstandings of the spatial positions of assembly components and the dynamics of human-machine interaction, and adaptive decision-making for the assembly environment is achieved.

[0019] Preferably, step S300 includes:

[0020] S310: A scene graph is used to describe human-robot collaborative assembly tasks, including the target of the task, the components involved, the assembly process, and the specific requirements of each link. Specifically, a graph structure is used to represent the nodes, including at least task target nodes, component nodes, assembly process nodes, and link requirement nodes; edges are used to represent the association between nodes, including the association between components and assembly processes, and the association between link requirements and assembly processes.

[0021] S320: processing the scene graph by using the large language model, analyzing, extracting and organizing the task information described by language in the graph, decomposing the task into a series of instructions with logical order and execution parameters according to the structured rules, and forming a plan sequence suitable for collaborative robot execution; wherein the structured rules include the format of the instructions, the determination rules of the logical order, and the assignment rules of the execution parameters; the plan sequence includes a plurality of instructions, each instruction contains an action type, an execution object, an execution parameter and a constraint condition.

[0022] As a preferred, in step S300, when it is necessary to expand or modify the human-robot collaborative assembly task, the following methods are used: if a new assembly task is added, a new task target node, a part node, an assembly process node and a link requirement node are added in the scene graph, and corresponding edge connections are established; if an existing task is modified, the node information or the association relationship of the edge in the scene graph is adjusted; the adjusted scene graph is again converted into a collaborative robot plan sequence by the large language model processing module.

[0023] As a preferred, in step S300, the determination method of the execution parameter includes: determining the accuracy range of the position parameter according to the accuracy requirement of the link requirement node in the scene graph; determining the reasonable value of the speed parameter and the force parameter according to the weight and material of the part combined with the action type of the assembly process node; the determination method of the constraint condition includes: determining the minimum safety distance in the safety constraint according to the safety standard of human-robot collaboration; determining the execution time window of each instruction in the time constraint according to the time requirement of the assembly task; determining the robot motion path restriction in the space constraint according to the layout of the assembly space.

[0024] As a preferred, step S400 includes:

[0025] S410: converting the pixel coordinates in the image coordinate system into accurate three-dimensional space coordinates in the base coordinate system of the collaborative robot by using the hand-eye calibration model built in the large language model; based on the three-dimensional space accurate coordinates, the rotation angle of each joint motor of the collaborative robot is calculated by using the inverse kinematics algorithm model;

[0026] S420: define the function of the collaborative robot execution code as a sub-function, including: joint space motion sub-function, supporting absolute position positioning and relative position motion; Cartesian space motion sub-function, supporting linear interpolation, circular interpolation trajectory planning; force control mode sub-function, supporting contact force feedback control and impedance control; calling the control code sub-function library, inputting the rotation angle, generating motion control code conforming to the hardware interface protocol of the collaborative robot, and driving the collaborative robot to execute the target action; the large language model automatically matches the sub-function type according to the assembly task type through a dynamic parameter adjustment mechanism, substitutes the calculated joint angle parameters and task-specific constraint conditions into the corresponding sub-function, and generates complete control code containing motion trajectory, speed curve and safety threshold.

[0027] As preferred, step S420 comprises:

[0028] S421: support dynamic parameter adjustment mechanism, when the force sensor detects abnormal contact force or the visual sensor identifies the workpiece position deviation, trigger real-time re-planning process;

[0029] S422: the large language model judges the deviation type through the task-level logic inference module, calls the online trajectory optimization algorithm to adjust the joint angle parameters and execution speed of the remaining motion segment, and generates a compensation amount;

[0030] S423: bring the compensation amount into the sub-function to generate a revised control code containing a new trajectory and a speed curve, realize dynamic compensation of the motion trajectory, and ensure safety and assembly accuracy in human-machine cooperation.

[0031] A device for somatic agent packaging for complex aerospace product human-machine collaborative assembly, which stores programs for human-machine collaborative assembly perception, reasoning and execution, and when the device for somatic agent packaging is started and run, the somatic agent packaging method for aerospace product human-machine collaborative assembly is realized.

[0032] Compared with the prior art, the present application has the beneficial effects of:

[0033] (1) Compared with the traditional fragmented system integration method, the somatic agent modular packaging technology (function function mapping, running logic association, interface plug and play) proposed in the present application can solve the problems of high system integration difficulty and high component coupling degree in complex aerospace product human-machine collaborative assembly.

[0034] (2) Compared with the traditional single-modal perception technology, the multi-modal fusion perception layer (text, image and video joint perception) based on the large language model in the present application can solve the problems of insufficient task adaptability of personnel behavior intention misjudgment and incomplete part position perception in the assembly site.

[0035] (3) Compared with the traditional rule engine task parsing method, the present invention integrates the large language model reasoning layer of the assembly scene graph, which can solve the problems of task semantic understanding deviation and low efficiency of assembly plan structured parsing in human-machine collaboration.

[0036] (4) Compared with the traditional manual programming execution mode, the present invention can solve the problem of insufficient execution accuracy caused by the reliance on manual debugging of robot action instructions and the low degree of automation in the assembly process through the execution layer technology (intelligent conversion of target coordinates and rotation angles) generated by large language models. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flow chart of a preferred embodiment of the embodied intelligent body packaging method for human-machine collaborative assembly of aerospace products in the present invention.

[0038] Figure 2 This is a schematic diagram of the embodied intelligent human-machine collaborative scene perception based on the large language model enhancement in the present invention.

[0039] Figure 3 This is a schematic diagram of the reasoning of embodied intelligent human-computer collaborative tasks based on scene graphs and large language models in the present invention.

[0040] Figure 4 This is a schematic diagram of the autonomous execution of the embodied intelligent agent enhanced based on the large language model in the present invention.

[0041] Figure 5 It is a block diagram of the device principle of the present invention. DETAILED DESCRIPTION

[0042] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.

[0043] The present invention discloses an embodied intelligent body encapsulation method and device for the human-machine collaborative assembly of complex aerospace products. In terms of the method, first, an embodied intelligent body for the human-machine collaborative assembly of aerospace products is constructed, and the encapsulation of the embodied intelligent body is realized through the function mapping, operation logic association and interface plug-and-play of the robot, which includes a perception layer, a reasoning layer and an execution layer; in the perception layer encapsulated by the embodied intelligent body, the powerful generalization ability of the large language model is used to construct the perception layer of the embodied intelligent body, and the multimodal perception means of text, image and video are used to realize the perception of the assembly behavior intention of the personnel and the position of the parts in the assembly environment; in the reasoning layer encapsulated by the embodied intelligent body, the reasoning ability of the large language model for the human-machine collaborative assembly task is improved through the assembly scene graph, and the human-machine collaborative assembly task can be parsed into an assembly plan sequence in a structured language, so that the collaborative robot can understand its own and the operator's assembly tasks; in the execution layer encapsulated by the embodied intelligent body, the target point coordinates of the task object to be executed are converted into the rotation angle of the collaborative robot through the large language model, thereby generating a complete collaborative robot action code. The present invention combines embodied intelligent agent encapsulation with a large language model, effectively improving the adaptability of the human-machine collaborative assembly process.

[0044] Example 1 like Figure 1 As shown, this embodiment discloses an embodied intelligent body packaging method for human-machine collaborative assembly of aerospace products, including the following steps:

[0045] S100: By analyzing the structure and functional composition of the physical robot, designing the physical entity mapping mechanism and the internal function mapping mechanism, constructing the mapping relationship between the embodied intelligent agent and the physical robot, and characterizing and encapsulating the embodied intelligent agent. Using the embedded controller as the integration interface, build a communication channel between the physical robot and the virtual embodied intelligent agent. Specifically,

[0046] S110: Embodied Agents encapsulate physical collaborative robots into embodied agents with task reasoning, multimodal cognition, and autonomous execution capabilities. By analyzing the structure and functional components of physical collaborative robots, and designing physical entity mapping mechanisms and intrinsic function mapping mechanisms, the embodied agent is mapped to the physical collaborative robot, thus completing the representation and encapsulation of the embodied agent.

[0047] S120: The physical robot is equipped with an embedded controller that serves as a core data processing and transmission unit, and is communicatively connected to various sensors, actuators, and other hardware modules of the physical robot. The sensors collect production-related data in real time during the operation of the physical robot, including but not limited to the position coordinates of the robot, the movement speed, the working state parameters (such as motor speed and torque), the size parameters of the processed parts, and the quality detection data. The embedded controller collects these production data through a pre-set communication protocol (such as TCP / IP, Modbus, CAN bus protocol, etc.), and transmits the production data through the embedded controller, ensuring the real-time, accuracy and stability of data transmission, and providing reliable data support for the monitoring, management and optimization of the production process.

[0048] S130: In the operating system of the embedded controller, an independent thread is created for each embodied agent, which serves as the basic unit for the operation of the embodied agent. The execution function API includes but is not limited to action functions corresponding to the physical robot, such as functions for controlling the joint movement of the robot and functions for implementing the grasping and releasing actions of the end effector, which are used to drive the physical robot to perform specific operation tasks.

[0049] In one embodiment, as shown in Figure 1 the functions include functions for obtaining the current state information of the robot and functions for reading sensor data, which are used to collect the running state parameters of the robot in real time. The dynamic link library encapsulates general functions for interacting between the embedded controller and hardware devices, data processing algorithm functions, and communication protocol processing functions, which are shared and called by multiple embodied agent threads to realize code reuse.

[0050] The dynamic link library adopts a standardized interface design, and its encapsulated general functions include device initialization, data reading and writing, and state monitoring functions; the data processing algorithm functions integrate filtering and noise reduction, coordinate conversion, and kinematics inverse solution algorithms; and the communication protocol processing functions support Modbus, CAN, and EtherCAT industrial bus protocol analysis, providing a unified hardware interaction interface for embodied agents.

[0051] When the system function is expanded, new embodied agent threads are created on the embedded controller, or the execution function API and the dynamic link library in the existing threads are modified: new function modules achieve hardware interaction by calling the basic functions of the dynamic link library, and the algorithm logic that needs to be expanded can be realized by adding data processing functions or modifying existing algorithm parameters, without the need for large-scale adjustment of the system hardware architecture and underlying software, thereby realizing the expandable control of the embodied agent system.

[0052] S140: When multiple embodied intelligent agents work together, the threads of each embodied intelligent agent realize task allocation and coordination through the scheduling mechanism of the embedded controller. The threads interact with each other and share information through communication methods such as shared memory, message queues, and sockets. According to the collaborative work strategies and algorithms (such as distributed collaborative control algorithms, multi-agent collaborative protocols, etc.), the embodied intelligent agents can cooperate with each other to jointly complete complex production tasks, such as multi-robot collaborative handling, assembly, and processing, which significantly improves the flexibility and efficiency of the production system.

[0053] S200: Enhances the perception capabilities of embodied agents based on a large language model, enabling semantic recognition and contextual understanding of multiple environmental entities, such as assembly parts, assembly tools, and human behavior. Specifically, by leveraging cross-modal matching between visual input and language descriptions, the large language model facilitates the integration of multimodal cognitive data into a unified representational space. By embedding language knowledge into the perception processing pipeline, this approach ensures that collaborative robots can not only perceive their environment but also understand the functional and relational context of assembly tasks, thereby achieving efficient, accurate, and safe human-robot collaboration.

[0054] In one embodiment, Figure 2 As shown, S200 includes:

[0055] S210: Extracts hierarchical visual representations from scene images using the core feature extraction module. The core feature extraction module uses the ResNet-50 visual encoder, which uses its deep residual architecture to capture spatial contextual details in the cluttered workspace and form high-level visual features for subsequent semantic segmentation.

[0056] S220: The high-level visual features are fused with language vectors through a Hadamard product-based fusion layer to generate a discriminative multimodal representation that combines visual spatial specificity with linguistic semantic accuracy, providing cross-modal foundational data for scene semantic parsing. The language vectors are derived from textual information in the multimodal perceptual input (in the form of voice commands or a preset semantic library). Embedding technology is used to convert natural language into high-dimensional vectors.

[0057] S230: Build an error-assisted guidance mechanism to address practical challenges such as occlusion and ambiguous input; identify differences in semantic segmentation results (performed during multimodal fusion) through a confidence analysis system, perform iterative cross-modal recalibration based on the multimodal representation, and gradually optimize scene understanding accuracy.

[0058] S240: For semantic misunderstanding phenomena such as visual occlusion, language ambiguity, cross-modal conflict, etc., in the iterative cross-modal recalibration process, through difference detection, recalibration and iterative optimization, the semantic misunderstanding of the spatial position of the assembly component and the dynamic of human-computer interaction is reduced, the robustness of the human-computer collaborative assembly system in complex scenes is enhanced, and adaptive decision-making for the assembly environment is realized.

[0059] The semantic analysis method is suitable for resource-constrained assembly scenes, and through accurate coordination of the visual encoder and the language reasoning module, the robot perception system can complete reliable semantic analysis under complex working conditions, and improve the efficiency of human-computer collaborative assembly.

[0060] S300: Embodied reasoning describes the task of human-computer collaborative assembly using a scene graph, and processes the scene graph using a large language model to convert it into a structured language written collaborative robot plan sequence.

[0061] In one embodiment, as shown in Figure 3 S300 includes:

[0062] S310: The scene graph is used to describe the task of human-computer collaborative assembly, and the scene graph covers the target of the task, the components involved, the process of assembly, and the specific requirements of each link, etc., providing the original task information basis for subsequent processing; the large language model processing module is used to process the scene graph, and convert it into a structured language written collaborative robot plan sequence. In the conversion process, the large language model processing module analyzes, extracts and organizes the task information in natural language, decomposes the task into a series of instructions with logical order and execution parameters according to the structured rules, and forms a plan sequence suitable for collaborative robot execution;

[0063] S320: The scene graph uses a graph structure to represent, including nodes and edges; wherein the nodes at least include a task target node, a component node, an assembly process node and a link requirement node. The task target node is used to describe the final target of the assembly task; the component node is used to describe all components involved in the assembly process, including the geometric size, material, connection mode and other attributes of the components; the assembly process node is used to describe the specific process of the assembly task, arranged in time or logical order; the link requirement node is used to describe the specific requirements of each assembly link, including precision requirements, force requirements, safety requirements, etc. The edges are used to represent the association between the nodes, including the association between the components and the assembly process, the association between the link requirements and the assembly process, etc.

[0064] S330: The processing process of the large language model processing module on the scene graph includes: first, analyzing the natural language description in the scene graph, identifying key information such as task target, parts, assembly process and requirements of each link; then, extracting the parsed information, including extracting attribute information of parts, sequence information of assembly process, specific parameters of link requirements, etc.; finally, according to the structured rules, the extracted information is organized into a plan sequence executable by the collaborative robot, and the structured rules include the format of the instruction, the determination rule of the logical sequence, the assignment rule of the execution parameter, etc.

[0065] S340: The collaborative robot plan sequence includes multiple instructions, each instruction contains action type, execution object, execution parameter and constraint condition; the action type includes grasping, moving, assembling, detecting, etc.; the execution object is the parts involved in the assembly process; the execution parameter is determined according to the specific requirements of the assembly link, including position parameter, speed parameter, force parameter, etc.; the constraint condition includes time constraint, space constraint, safety constraint, etc., to ensure the safety and accuracy of the collaborative robot when executing the task.

[0066] S350: When it is necessary to expand or modify the human-robot collaborative assembly task, the following methods can be used: if a new assembly task is added, new task target nodes, part nodes, assembly process nodes and link requirement nodes are added in the scene graph, and corresponding edge connections are established; if the existing task is modified, the node information or the association relationship of the edge in the scene graph is adjusted; the adjusted scene graph is converted into a collaborative robot plan sequence again through the large language model processing module, without the need for large-scale modification of the underlying architecture and core algorithm of the system, realizing the flexible expansion of the system.

[0067] S360: The determination method of the execution parameter includes: determining the accuracy range of the position parameter according to the accuracy requirement of the link requirement node in the scene graph; determining the reasonable value of the speed parameter and the force parameter according to the weight and material of the part, combined with the action type of the assembly process node. The determination method of the constraint condition includes: determining the minimum safety distance in the safety constraint according to the safety standard of human-robot collaboration; determining the execution time window of each instruction in the time constraint according to the time requirement of the assembly task; determining the robot motion path restriction in the space constraint according to the layout of the assembly space.

[0068] S400: The embodied execution can input the generated collaborative robot plan sequence into the subsequent module, and according to the instructions and parameters in the plan sequence, combined with the hardware characteristics of the collaborative robot, the working environment and other factors, further processing and calculation are carried out, and finally the action of the collaborative robot is generated to realize the human-robot collaborative assembly task.

[0069] In one embodiment, as Figure 4As shown, S400 includes:

[0070] S410: The large language model converts the pixel coordinates in the image coordinate system to precise three-dimensional space coordinates in the collaborative robot base coordinate system through the built-in hand-eye calibration model. The coordinate conversion calculation unit calculates the rotation angles of each joint motor of the collaborative robot based on the precise three-dimensional space coordinates using an inverse kinematics algorithm model. The control code generation unit substitutes the rotation angles into the control code sub-function library preset by the large language model to generate motion control codes that conform to the hardware interface protocol of the collaborative robot, thereby driving the collaborative robot to perform target actions such as grabbing, moving, and assembling.

[0071] S410 specifically includes:

[0072] S411: The functions of the execution code need to be defined as sub-functions. Based on the analysis results of the structured planning, these sub-functions are used to arrange the action sequence of the collaborative robot, thereby forming the main function code for the collaborative robot to perform the action.

[0073] S412: Based on the results of multi-modal perception, the joint rotation angles of the target object that needs to be grabbed or carried are calculated, and the calculation results are input as the calling parameters of the corresponding sub-function code. The specific action code of the robot can be automatically generated, the abstract task description is systematically converted into low-level executable instructions, thereby eliminating manual coding work and reducing human intervention in the robot programming process.

[0074] S420: The control code sub-function library includes: joint space motion sub-function, supporting absolute position positioning and relative position motion; Cartesian space motion sub-function, supporting linear interpolation, circular interpolation, and other trajectory planning; force control mode sub-function, supporting contact force feedback control and impedance control; The large language model adjusts the dynamic parameters according to the assembly task type (automatically matches the sub-function type, substitutes the calculated joint angle parameters and task-specific constraint conditions into the corresponding sub-function, and generates complete control codes containing motion trajectories, velocity curves, and safety thresholds.

[0075] S420 specifically includes:

[0076] S421: The module supports a dynamic parameter adjustment mechanism. When the force sensor detects abnormal contact force or the vision sensor identifies that the workpiece position deviates, a real-time re-planning process is triggered.

[0077] S422: The large language model determines the deviation type through the task-level logical inference module and calls an online trajectory optimization algorithm (such as the time-optimal quadratic programming algorithm) to adjust the joint angle parameters and execution speed of the remaining motion segment.

[0078] S423: The control code generation unit generates a corrected control code according to the updated parameters, realizes dynamic compensation of the motion trajectory, and ensures safety and assembly accuracy in the human-robot collaboration process.

[0079] The application improves the adaptive ability of the human-robot collaboration assembly process by combining the embodiment of the intelligent agent with the large language model, generating an assembly plan sequence according to the perception result and automatically converting it into robot execution code, thereby solving the problems of difficult system integration, insufficient task adaptability and difficult efficient collaboration in the human-robot collaboration assembly of complex aerospace products.

[0080] Embodiment 2

[0081] In one embodiment, as shown in Figure 5 An apparatus for embodiment of intelligent agent is provided, which stores programs for human-robot collaboration assembly perception, reasoning and execution. When the apparatus is started and run, the encapsulation method in embodiment 1 is realized.

[0082] Specifically, on the hardware side, it includes:

[0083] An embedded controller unit is used to create an independent thread for each embodiment of the intelligent agent; a dynamic link library is integrated to encapsulate device initialization, kinematics inverse calculation algorithm and Modbus / CAN / EtherCAT protocol analysis; and a TCP / IP port is used to communicate with the upper management system.

[0084] A multi-modal sensor group includes a visual sensor for real-time acquisition of assembly scene data, a force sensor for detection of robot end contact force and triggering of dynamic compensation, and a voice sensor for capturing voice instructions to generate a language vector.

[0085] A collaborative robot execution mechanism includes a multi-degree-of-freedom mechanical arm.

[0086] On the software side, it includes:

[0087] On the perception layer, a ResNet-50 visual encoder is included to extract spatial features; a language vector generation module is included to convert voice instructions into a multi-dimensional vector; and a multi-modal fusion module is included to fuse visual / language features based on Hadamard product.

[0088] On the reasoning layer, a scene graph database is included to store task targets, part attributes and assembly process nodes; and a large language model is included to analyze the graph and generate a structured instruction sequence.

[0089] On the execution layer, a control code sub-function library is included, which contains functions such as joint motion and linear interpolation; and a dynamic re-planning module is included to integrate a time-optimal quadratic programming algorithm.

[0090] The working process of the device is as follows: firstly, the environment is perceived: the assembly table image is captured through the visual sensor, and the language vector is output based on the multi-modal fusion module according to the voice instruction of the operator. Then task reasoning is performed, the scene graph node is updated, the graph is parsed based on the large language model, and the structured sequence is generated. Finally, the command is dynamically executed, the joint angle calculated based on the inverse kinematics is obtained, the mechanical arm moves to the target coordinate; whether the re-planning is triggered is judged based on the detection data of the force sensor, if the re-planning is triggered, the control code generation unit mobilizes the control sub-function, updates the trajectory parameters, and the assembly is completed.

[0091] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand; the technical solutions recorded in the foregoing embodiments can be modified, or some technical features can be replaced by equivalents; and these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the corresponding embodiments, and should be included in the protection scope of the present application.

Claims

1. A method for packaging embodied intelligent bodies for human-machine collaborative assembly of complex aerospace products, characterized by: The following steps are involved: S100: Construct a mapping relationship between the embodied intelligent agent and the physical robot, characterize and encapsulate the embodied intelligent agent, which includes a perception layer, a reasoning layer, and an execution layer; use the embedded controller as an integration interface to build a communication channel between the physical robot and the virtual embodied intelligent agent; S200: At the perception layer, the perception capability of the embodied intelligent agent is enhanced based on a large language model. Scene features are extracted through a ResNet-50 visual encoder and fused with language vectors via a Hadamard product to generate a multimodal representation. This semantics of the assembly environment is analyzed to achieve perception of the assembly behavior intentions of the personnel and the location of components in the assembly environment. S300: At the reasoning layer, the assembly task is described using a scene graph, which is converted into a structured collaborative robot plan sequence using a large language model; S400: At the execution layer, the plan sequence is used to generate robot control code through the hand-eye calibration model of the large language model and the inverse kinematics algorithm, and finally the action of the collaborative robot is generated to achieve the human-machine collaborative assembly task.

2. The method according to claim 1, characterized in that In step S100, the encapsulation of the embodied intelligent body is completed by analyzing the structural design of the physical collaborative robot, which includes a dual mapping mechanism of physical entity mapping and intrinsic function mapping; the physical collaborative robot is equipped with an embedded controller, which collects sensor data through a preset communication protocol; in the operating system of the embedded controller, an independent thread is created for each embodied intelligent body, which serves as the basic unit for the operation of the embodied intelligent body, and provides an execution function API for joint control / grasping and release. Based on the execution function API, the current state information of the robot is obtained, sensor data is read, and the operation status of the robot is collected in real time. The dynamic link library encapsulates the general functions, data processing algorithm functions and communication protocol processing functions for the interaction between the embedded controller and the hardware devices; when multiple embodied intelligent bodies work together, the threads of each intelligent body realize task allocation and coordination through the scheduling mechanism of the embedded controller, and realize collaborative scheduling through shared memory / message queues.

3. The method according to claim 2, characterized in that In step S100, the dynamic link library adopts a standardized interface design. The general functions it encapsulates include device initialization, data reading and writing, and status monitoring functions; the data processing algorithm function integrates filtering and noise reduction, coordinate transformation, and kinematic inverse solution algorithms; and the communication protocol processing function supports Modbus, CAN, and EtherCAT industrial bus protocol parsing, providing a unified hardware interaction interface for the embodied intelligent body. When expanding system functions, a new embodied intelligent body thread is created on the embedded controller, or the execution function API and dynamic link library in the existing thread are modified in a targeted manner: the new functional module implements hardware interaction by calling the basic functions of the dynamic link library, and the algorithm logic that needs to be expanded is implemented by adding data processing functions or modifying the existing algorithm parameters.

4. The method according to claim 1, wherein Step S200 includes: S210: Uses the ResNet-50 visual encoder to extract hierarchical visual representations from scene images, and uses a deep residual architecture to capture spatial context details in a cluttered workspace to form high-order visual features for subsequent semantic segmentation; S220: fusing the high-order visual features with the language vectors through a fusion layer based on the Hadamard product to generate a discriminative multimodal representation with both visual spatial specificity and language semantic accuracy; S230: For occluded and ambiguous inputs, an error-assisted guidance mechanism is constructed, and the differences in the semantic segmentation results are identified through a confidence analysis system. Based on the multimodal representation, iterative cross-modal recalibration is performed to reduce semantic misunderstandings of the spatial positions of assembly components and the dynamics of human-computer interaction, thereby achieving adaptive decision-making for the assembly environment.

5. The method according to any one of claims 1 to 4, characterized in that Step S300 includes: S310: Describe the human-machine collaborative assembly task using a scenario graph, including the task objective, the components involved, the assembly process, and the specific requirements of each link. Specifically, a graph structure containing nodes and edges is used to represent the task. The nodes include at least task objective nodes, component nodes, assembly process nodes, and link requirement nodes. Edges are used to represent the associations between nodes, including the association between components and assembly processes, and the association between link requirements and assembly processes. S320: Use a large language model to process the scene graph, parse, extract and organize the task information described in the language in the graph, and decompose the task into a series of instructions with logical order and execution parameters according to structured rules to form a plan sequence suitable for execution by collaborative robots; wherein the structured rules include the format of instructions, the rules for determining the logical order, and the rules for assigning execution parameters; the plan sequence includes multiple instructions, each instruction contains an action type, an execution object, execution parameters and constraints.

6. The method according to claim 5, characterized in that In step S300, when the human-machine collaborative assembly task needs to be expanded or modified, it is achieved in the following way: if a new assembly task is added, a new task target node, component node, assembly process node and link requirement node are added to the scene graph, and corresponding edge connections are established; if an existing task is modified, the corresponding node information or edge association relationship in the scene graph is adjusted; the adjusted scene graph is again converted into a collaborative robot plan sequence through the large language model processing module.

7. The method according to claim 5, characterized in that In step S300, the method for determining the execution parameters includes: determining the accuracy range of the position parameters according to the accuracy requirements of the nodes in the link requirements in the scene graph; determining the reasonable values ​​of the speed parameters and the force parameters according to the weight and material of the parts and the action type of the assembly process node; the method for determining the constraint conditions includes: determining the minimum safety distance in the safety constraint according to the safety standards of human-machine collaboration; determining the execution time window of each instruction in the time constraint according to the time requirements of the assembly task; and determining the robot motion path restriction in the space constraint according to the layout of the assembly space.

8. The method according to claim 1, characterized in that Step S400 includes: S410: Using the hand-eye calibration model built into the large language model, the pixel coordinates in the image coordinate system are converted into precise three-dimensional coordinates in the base coordinate system of the collaborative robot; based on the precise three-dimensional coordinates, the rotation angle corresponding to the motor of each joint of the collaborative robot is calculated using the inverse kinematics algorithm model; S420: The functions of the collaborative robot execution code are defined as sub-functions, including: joint space motion sub-function, which supports absolute position positioning and relative position motion; Cartesian space motion sub-function, which supports trajectory planning of linear interpolation and circular interpolation; force control mode sub-function, which supports contact force feedback control and impedance control; calling the control code sub-function library, inputting the rotation angle, generating motion control code that complies with the collaborative robot hardware interface protocol, and driving the collaborative robot to perform the target action; the large language model automatically matches the sub-function type according to the assembly task type through the dynamic parameter adjustment mechanism, and substitutes the calculated joint angle parameters and task-specific constraints into the corresponding sub-function to generate a complete control code including motion trajectory, speed curve, and safety threshold.

9. The method according to claim 8, characterized in that Step S420 includes: S421: Supports dynamic parameter adjustment mechanism, triggering real-time re-planning process when force sensor detects abnormal contact force or vision sensor identifies workpiece position deviation; S422: The large language model determines the deviation type through the task-level logical reasoning module and calls the online trajectory optimization algorithm to adjust the joint angle parameters and execution speed of the remaining motion segments to generate compensation. S423: The compensation amount is brought into the sub-function to generate a correction control code containing a new trajectory and speed curve to achieve dynamic compensation of the motion trajectory and ensure safety and assembly accuracy during the human-machine collaboration process.

10. A device for embodied intelligent body packaging for human-machine collaborative assembly of complex aerospace products, characterized by: The device stores programs for human-machine collaborative assembly perception, reasoning and execution, and when the device for embodied intelligent body encapsulation is started and run, the encapsulation method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Representation understanding method based on multi-scale cross-modal feature fusion

    CN115496991A

  • Target object grabbing method and system based on multi-source knowledge driving

    CN118038221A

  • Method for uniformly converting coordinates of robot into three-dimensional coordinate system and virtual platform

    CN118744432A

  • Gas turbine augmented reality auxiliary assembly method based on image recognition technology

    CN120122820A

  • Man-machine cooperation safe assembly method based on safe skin of industrial robot

    CN120287308A

Cited By

  • Autonomous decision-making method and system for intelligent robot with body

    CN121716084A

  • An autonomous decision-making method and system for embodied intelligent robots

    CN121716084B

  • Control method and system for mechanical arm of collaborative robot

    CN121756362A