Industrial robot autonomous assembly task planning method based on multi-agent large model

By analyzing assembly task instructions and drawings using a large multi-agent model, and combining multimodal perception and behavior tree planning, the autonomy and efficiency issues of industrial robots in complex assembly tasks are solved, enabling intelligent assembly of robots in dynamic environments.

CN120715913BActive Publication Date: 2025-11-18THE HONG KONG POLYTECHNIC UNIV SHENZHEN RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511222403.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-11-18
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing industrial robots rely on static programming and predefined processes, lack the ability to parse natural language instructions and drawing information, and are unable to handle complex assembly tasks and frequently changing production processes. The perception and decision-making modules are fragmented, and they lack multimodal information fusion and dynamic planning capabilities.

Method used

A multi-agent large model-based approach is adopted to parse assembly task instructions and drawings through a visual language model and construct task execution logic. A multimodal perception system is used to collect assembly environment images in real time, generate symbolic behavior trees for assembly task planning, and combine visual and sensor data for dynamic adjustment, execution monitoring, and feedback adjustment.

Benefits of technology

It achieves semantic parsing of unstructured tasks and dynamic adjustment of environmental states, improving the autonomous assembly efficiency and accuracy of the autonomous assembly task planning system of the assembly system, and enhancing the task completion rate and execution efficiency of robots in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120715913B_ABST
    Figure CN120715913B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of industrial robot autonomous assembly task planning method based on multi-agent large model, belong to the technical field of automated assembly.The method comprises: obtaining the assembly task data of assembly object, and the assembly task data includes the task instruction and assembly manual for assembling assembly object;Task instruction and assembly manual are parsed, and task execution logic is obtained;The assembly environment when assembling assembly object is collected in real time, and the target image is obtained;Multi-modal perception is carried out on the target image, and the perception result is obtained, and the perception result includes the position of at least one component of assembly object in assembly environment;Task execution logic and perception result are converted into planning domain definition language format, and assembly task planning is carried out through behavior tree, and the target assembly task is obtained.The embodiment of the application improves the assembly efficiency and precision of industrial robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automated assembly technology, and in particular to a method for planning autonomous assembly tasks for industrial robots based on a multi-agent large model. Background Technology

[0002] As Industry 4.0 evolves into Industry 5.0, manufacturing models are shifting from efficiency-centric full automation to human-machine collaborative systems centered on flexible cooperation, with particular emphasis on the autonomous task execution and collaborative capabilities of intelligent robots in dynamic and changing environments. In this context, industrial robots not only need basic operational functions but also semantic understanding, environmental perception, task reasoning, and strategy generation capabilities to adapt to complex, personalized, and unstructured production demands. This transformation places higher demands on the intelligence level of robot systems, especially in areas such as task understanding, multimodal information fusion, and long-term planning.

[0003] However, most industrial robots currently rely on static programming and predefined process-driven operations, lacking the ability to parse unstructured inputs such as natural language instructions, drawing information, and task specifications. This results in the understanding, decomposition, and execution of tasks being highly dependent on manual configuration, making them unsuitable for complex assembly tasks or frequently changing production processes. Summary of the Invention

[0004] The main objective of this application is to propose an autonomous assembly task planning method for industrial robots based on a multi-agent large model, which aims to solve the technical problem of existing industrial robot assembly relying on manual labor, thereby improving the assembly efficiency and accuracy of industrial robots.

[0005] To achieve the above objectives, a first aspect of this application proposes a method for autonomous assembly task planning of industrial robots based on a multi-agent large model, the method comprising:

[0006] Acquire assembly task data for the assembly object, wherein the assembly task data includes task instructions and assembly manuals for assembling the assembly object;

[0007] The task instructions and the assembly manual are parsed to obtain the task execution logic;

[0008] The assembly environment during the assembly of the assembly object is captured in real time to obtain the target image;

[0009] Multimodal perception is performed on the target image to obtain perception results, the perception results including the position of at least one component of the assembly object in the assembly environment;

[0010] The task execution logic and the perception results are converted into a planning domain definition language format, and the assembly task is planned using a behavior tree to obtain the target assembly task.

[0011] In some embodiments, parsing the task instructions and the assembly manual to obtain the task execution logic includes:

[0012] The task instructions and assembly manual are semantically parsed using a visual language model to obtain the category of each component, the functional role of each component, and the spatial structural relationship between each component.

[0013] The task execution logic is constructed based on the category of each component, the functional role of each component, and the spatial structural relationship between each component. The task execution logic includes multiple sub-tasks, each sub-task including installation description information of at least one component, and the execution order of the sub-tasks.

[0014] In some embodiments, performing multimodal perception on the target image to obtain a perception result includes:

[0015] The target image is segmented to obtain segmented images of each assembly object in the target image;

[0016] The segmented image of each assembly object is subjected to grasping pose perception to obtain the pose information of each assembly object, which includes the graspability score and grasping pose of each assembly object.

[0017] The perception results are obtained by performing semantic analysis on the segmented images and pose information of each assembly object using a visual language model. The perception results include the spatial relationships between each assembly object, the operational feasibility of each assembly object, and the attributes of each assembly object.

[0018] In some embodiments, the step of performing image segmentation on the target image to obtain segmented images of each assembly object in the target image includes:

[0019] Target detection is performed on the target image to obtain the bounding boxes and preliminary category labels of each assembly object in the target image;

[0020] The target image is segmented based on the bounding box and the preliminary category label of each assembly object to obtain a segmented image of each assembly object. The segmented image includes the outline boundary and target category label of each assembly object.

[0021] In some embodiments, the step of converting the task execution logic and the perception results into a planning domain definition language format, and performing assembly task planning through a behavior tree to obtain the target assembly task, includes:

[0022] The task execution logic and the perception results are converted into data formats to obtain new task execution logic and new perception results in the planning domain definition language format.

[0023] Based on the new task execution logic and the new perception results, a behavior tree task architecture is generated. The behavior tree task architecture includes multiple task nodes, including sequence nodes, condition nodes, and backoff nodes. The sequence nodes are used to determine the execution order of subtasks. The condition nodes are used to determine whether the current state meets the conditions for executing the current subtask. The backoff nodes are used to re-execute the current subtask or re-plan the assembly task when the current subtask fails to execute.

[0024] The assembly task is planned based on the behavior tree task architecture to obtain the target assembly task.

[0025] In some embodiments, after converting the task execution logic and the perception result into a planning domain definition language format, and performing assembly task planning through a behavior tree to obtain the target assembly task, the method further includes:

[0026] The assembly environment during the assembly of the assembled object was filmed to obtain a video.

[0027] Image analysis is performed on the captured video. If it is determined that the target assembly task has failed actions or the action deviation exceeds a preset threshold during execution, the current subtask is re-executed or the process is redirected to convert the task execution logic and the perception result into a planning domain definition language format. Then, the assembly task is planned using a behavior tree to obtain the steps of the target assembly task.

[0028] In some embodiments, the image analysis of the captured video, if it is determined that the target assembly task has failed actions or deviated from a preset threshold during execution, then the current subtask is re-executed or the process jumps to the step execution of converting the task execution logic and the perception results into a planning domain definition language format, and planning the assembly task through a behavior tree to obtain the steps of the target assembly task, includes:

[0029] The captured video is analyzed using a visual language model to obtain the execution result of the current subtask.

[0030] Acquire physical sensor data of the assembly equipment that assembles the assembly object to obtain verification data, which includes gripper status and force feedback data.

[0031] The execution result of the subtask is verified using the verification data to obtain the verification result;

[0032] If the verification result indicates that the target assembly task has failed an action or the action deviation exceeds a preset threshold during execution, then the current subtask is re-executed or the process jumps to the step execution of the target assembly task by converting the task execution logic and the perception result into a planning domain definition language format and planning the assembly task through a behavior tree.

[0033] To achieve the above objectives, a second aspect of this application proposes an industrial robot autonomous assembly task planning device based on a multi-agent large model, the device comprising:

[0034] The acquisition module is used to acquire assembly task data of the assembly object, the assembly task data including task instructions and assembly manuals for assembling the assembly object.

[0035] The parsing module is used to parse the task instructions and the assembly manual to obtain the task execution logic;

[0036] The acquisition module is used to acquire the assembly environment in real time when the assembly object is assembled, and obtain the target image.

[0037] A perception module is used to perform multimodal perception on the target image and obtain a perception result, the perception result including the position of at least one component of the assembly object in the assembly environment;

[0038] The planning module is used to convert the task execution logic and the perception results into a planning domain definition language format, and to plan the assembly task through a behavior tree to obtain the target assembly task.

[0039] To achieve the above objectives, a third aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described in the first aspect.

[0040] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0041] This application proposes an autonomous assembly task planning method for industrial robots based on a multi-agent large model. It parses assembly task instructions and operation manuals using a visual language model to extract task objects, roles, and spatial relationships, constructing task execution logic. A multimodal perception system aggregates perceive the 3D image of the assembly environment, ensuring the robot has real-time awareness of the current scene. A large language model generates a symbolic behavior tree with execution conditions and backoff strategies, where each leaf node corresponds to a skill defined in the robot's skill library, and dynamic scheduling is performed based on the node's return status. During execution, visual and sensor data are integrated for multimodal verification, and state judgment is achieved through visual question answering and physical sensor data, enabling execution monitoring and feedback adjustments. By constructing a task hierarchy graph structure, multimodal perception and symbolic mapping, dynamic behavior tree planning, and a multimodal verification mechanism, this application achieves semantic parsing of unstructured task instructions, symbolic representation of environmental states, and dynamic task adjustment capabilities, enhancing autonomous decision-making capabilities, realizing a closed-loop multimodal information fusion, and strengthening the dynamic adjustment capabilities for complex assembly tasks. Attached Figure Description

[0042] Figure 1 This is a flowchart illustrating the autonomous assembly task planning method for industrial robots based on a multi-agent large model provided in this application embodiment;

[0043] Figure 2 This is a logical schematic diagram of the industrial robot autonomous assembly task planning method based on a multi-agent large model provided in the embodiments of this application;

[0044] Figure 3 This is a schematic diagram of the structure of the industrial robot autonomous assembly task planning device based on a multi-agent large model provided in this application embodiment;

[0045] Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0047] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0049] As Industry 4.0 evolves into Industry 5.0, manufacturing models are shifting from efficiency-centric full automation to human-machine collaborative systems centered on flexible cooperation, with particular emphasis on the autonomous task execution and collaborative capabilities of intelligent robots in dynamic and changing environments. In this context, industrial robots not only need basic operational functions but also semantic understanding, environmental perception, task reasoning, and strategy generation capabilities to adapt to complex, personalized, and unstructured production demands. This transformation places higher demands on the intelligence level of robot systems, especially in areas such as task understanding, multimodal information fusion, and long-term planning.

[0050] However, most industrial robots currently rely on static programming and predefined process-driven operations, lacking the ability to parse unstructured inputs such as natural language instructions, drawing information, and task specifications. This results in the understanding, decomposition, and execution of tasks being highly dependent on manual configuration, making them unsuitable for complex assembly or frequently changing production tasks.

[0051] Furthermore, the existing system's perception and decision-making modules are disconnected, making it difficult to convert multi-source sensor information such as vision and force into usable symbolic representations for building task models or generating dynamic programming strategies. This severely restricts the robot's adaptability and robustness in open environments.

[0052] Especially when dealing with long-term tasks with multiple steps and strong dependencies, traditional methods show significant shortcomings in task structure modeling, execution interruption recovery, and multi-task switching, and lack the ability to dynamically schedule based on task semantics and environmental state.

[0053] Based on this, the embodiments of this application provide an autonomous assembly task planning method for industrial robots based on a multi-agent large model. The aim is to achieve structured planning and robust execution of complex industrial assembly tasks through multimodal perception, symbolic task modeling and behavior tree execution mechanism, thereby improving the task completion rate and execution efficiency of robots in complex dynamic environments.

[0054] The autonomous assembly task planning method for industrial robots based on a multi-agent large model provided in this application is specifically illustrated through the following embodiments. First, the autonomous assembly task planning method for industrial robots based on a multi-agent large model in this application is described.

[0055] The autonomous assembly task planning method for industrial robots based on a multi-agent large model provided in this application relates to the field of automated assembly technology. This method can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the autonomous assembly task planning method for industrial robots based on a multi-agent large model, but is not limited to the above forms.

[0056] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0057] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0058] Figure 1This is an optional flowchart of the industrial robot autonomous assembly task planning method based on a multi-agent large model provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S100 to S500.

[0059] Step S100: Obtain assembly task data of the assembly object, wherein the assembly task data includes task instructions and assembly manual for assembling the assembly object.

[0060] In this embodiment, the task instruction can specify which parts need to be assembled, the approximate assembly sequence, or key steps. The assembly manual may contain more detailed assembly procedures, technical requirements, precautions, part information, etc. Specifically, the system can first obtain assembly task data of the product to be assembled from a database, local files, or through a network interface. The assembly task data can include task instructions, assembly manuals, etc. The task instruction can be a simple text instruction or structured data, such as: install part A onto base B, and then install part C onto part A; the assembly manual can be an electronic document or structured database containing detailed steps, part drawings, and technical requirements, such as: part A has three mounting holes that need to be aligned with the corresponding threaded holes on base B; when installing part C, it is necessary to ensure that the groove on it fully engages with the protrusion on part A.

[0061] Step S200: Parse the task instructions and the assembly manual to obtain the task execution logic.

[0062] In this embodiment, the system can utilize Natural Language Processing (NLP) technology or Visual Language Modeling (VLM) to analyze task instructions and assembly manual text. For example, it can identify action keywords such as "installation," "alignment," and "engagement," extract the names of the involved parts (A, B, C), and parse out the assembly relationships and sequence constraints between them. For structured manual data, the corresponding rules and parameters can be directly extracted. The parsed result, i.e., the task execution logic, is generated by identifying object categories, functional roles, and spatial structural relationships. It outputs a component list and its relational attributes, and constructs the task execution logic, which can be represented by a task hierarchy diagram structure.

[0063] Specifically, a visual language model is used to perform semantic understanding of the input assembly drawings and operation manuals. A pre-trained visual language model can be used to extract object categories, functional roles, and spatial relationships from the images, thereby constructing a task hierarchy graph structure. By transforming unstructured visual information into a structured task graph, the problem of traditional methods being unable to parse drawings and natural language instructions is solved.

[0064] Step S300: The assembly environment during the assembly of the assembly object is collected in real time to obtain the target image.

[0065] In this embodiment, a 3D camera mounted on the robot acquires real-time images of the work scene, including color images and depth information, to support subsequent target detection, segmentation, and pose estimation. Multiple sensors, such as cameras mounted on the robot arm, fixed above or to the side of the workstation, and depth cameras, can be used to acquire real-time visual information of the assembly area. Specifically, multiple sensors, including at least one depth camera, are installed in the assembly work area. During the robot's assembly task, these sensors continuously operate, acquiring real-time image data of the assembly area, including color images and depth maps. For example, when the robot prepares to grasp part A, the sensors capture the actual position and pose of part A on the current worktable, as well as information about the surrounding environment (such as other parts, tools, and fixtures).

[0066] Step S400: Perform multimodal perception on the target image to obtain a perception result, the perception result including the position of at least one component of the assembly object in the assembly environment.

[0067] In this embodiment, the system processes the received target images (RGB images and depth images). First, target detection and recognition algorithms are used to identify the various components of the assembly object (parts A, B, C, etc.) in the RGB image. Then, combined with the depth image information, the three-dimensional position and pose of these components in the robot coordinate system are accurately calculated. The perception results include not only the position coordinates and orientation of part A, but may also include the component's state (such as whether it has been installed, whether it is correctly placed, whether it is in a graspable state, whether it is occluded, etc.), and its relative relationship with other components.

[0068] Step S500: The task execution logic and the perception result are converted into a planning domain definition language format, and the assembly task is planned through a behavior tree to obtain the target assembly task.

[0069] In this embodiment, the system integrates task execution logic with perception results. Specifically, natural language instructions, task execution logic, and perception results can be used as inputs to generate a symbolic behavior tree with execution conditions and fallback strategies using a large language model. Behavior tree task planning refers to using a large language model to generate a symbolic behavior tree from natural language instructions, task hierarchy diagrams, and perception states. Specifically, a large language model can be used to generate a behavior tree containing execution conditions and fallback strategies, with its leaf nodes bound to atomic skills defined by the Planning Domain Definition Language (PDDL). Each leaf node in the behavior tree corresponds to a robot skill (such as grasping, insertion, and transportation) defined in the robot skill library, and is dynamically scheduled during execution based on the Success, Failure, or Running status returned by the node.

[0070] Specifically, the system matches and integrates the task execution logic with the current perception results. If the task logic requires the installation of part B, but the perception results indicate that part B is not in the expected position or its orientation is incorrect, the system can dynamically plan a new assembly path or strategy adapted to the current environment based on the currently perceived actual position and orientation of part B, as well as the requirements regarding installation sequence and constraints in the task logic. The target assembly task can be a series of specific action instructions, such as "move to the X coordinate, grab part B in the Y orientation, and then move to the Z coordinate for installation."

[0071] The logic diagram of this embodiment is as follows: Figure 2As shown, firstly, a visual language model is used to parse visual information such as assembly drawings and operation manuals, extracting task objects, roles, and spatial relationships to construct a task hierarchy graph structure. This step realizes the transformation from unstructured input to a structured task model. Secondly, object recognition and pose estimation are performed through a multimodal perception system. Combined with the visual language model, object categories, spatial relationships, and task states are extracted, generating symbolic predicates and mapping them to initial states and action premises in PDDL format. This step solves the semantic gap problem between the multimodal perception and symbolic planning modules. Thirdly, a large language model is used to generate a symbolic behavior tree with execution conditions and backoff strategies. Each leaf node corresponds to a skill defined in the robot's skill library, and dynamic scheduling is performed based on the node's returned state. This step realizes the adjustment of execution logic based on real-time perceived state. Finally, during execution, visual and sensor data are integrated for multimodal verification. State judgment is performed through visual question answering and physical sensor data to achieve execution monitoring and feedback adjustment. This step ensures the accuracy of state judgment and supports backoff or replanning in case of failure. The collaborative work among the steps is reflected in the following: the task structure extracted by the visual language model provides a framework for behavior tree generation; the multimodal perception results are mapped into PDDL format, providing a basis for judging the execution conditions of the behavior tree; and the dynamic scheduling mechanism of the behavior tree, combined with multimodal verification, enables real-time monitoring and adjustment of task execution. This collaborative approach allows the system to adapt to complex and dynamic assembly environments, improving the flexibility and robustness of task execution.

[0072] This embodiment, by acquiring assembly environment images in real time and performing multimodal perception, can promptly detect changes in the assembly environment and dynamically adjust the assembly plan according to the task execution logic, thereby improving the assembly system's adaptability to complex and dynamic environments. It can optimally plan assembly paths and strategies based on the actual environment, avoiding unnecessary movements or adjustments, which may shorten assembly time and improve production efficiency. Combined with a logical understanding of the task itself, the planned task is not merely a simple reaction to the environment, but an intelligent decision based on assembly goals and rules, reducing assembly failures caused by minor environmental disturbances.

[0073] In some embodiments, step S200 may include, but is not limited to, steps S210 to S220:

[0074] Step S210: Semantic parsing of the task instructions and the assembly manual is performed using a visual language model to obtain the category of each component, the functional role of each component, and the spatial structural relationship between each component.

[0075] Step S220: Construct the task execution logic according to the category of each component, the functional role of each component, and the spatial structural relationship between each component. The task execution logic includes multiple sub-tasks, and each sub-task includes installation description information of at least one component and the execution order of the sub-tasks.

[0076] In this embodiment, a Vision-Language Model (VLM) is used to semantically parse the task instructions and assembly manual. This allows for the understanding of the semantic relationships between natural language text and images, thereby extracting the categories, functional roles, and spatial structural relationships of each component. This leads to the construction of the task execution logic, i.e., the task hierarchy graph structure G=(O,H,R). Here, O is the object set containing all identified parts. H represents the directed composition relationship of subtasks, indicating the assembly order and hierarchy. R represents the equivalence relationship of replaceable components, identifying interchangeable parts with the same function. Finally, a task problem definition in PDDL format is automatically generated through structural mapping. PDDL is a standard representation language widely used in the field of artificial intelligence planning; this conversion enables subsequent planning tools to understand and process it. Objects are mapped to PDDL objects, relationships to predicates, and assembly steps to actions. The generated PDDL file contains: domain, objects, init, and goal sections, corresponding to the problem domain, object set, initial state, and target state, respectively.

[0077] Specifically, the visual language model generates a structured component list through multimodal input parsing. The object set O contains the physical entities involved in the assembly task and their category labels, such as screws, bearings, and bases. The directed composition relations H of subtasks are generated by parsing the step sequence and conditional dependencies in the operation manual; for example, "installing the base" must be performed before "fixing the screws." The equivalence relations R of replaceable components are established by identifying functionally equivalent parts; for example, two types of screws of the same specification are interchangeable. The task problem definition in PDDL format is automatically generated through structural mapping, providing structured input for subsequent symbolic planning and behavior tree generation.

[0078] Specifically, the visual language model first jointly parses the graphic symbols and text annotations in the assembly drawings to extract object categories and spatial relationships, such as recognizing the annotation "Part A needs to be inserted into the slot of Part B" in the drawings. Then, the model constructs directed combination relationships between subtasks based on the step descriptions in the operation manual, such as generating the sequence relationship in H from "Step 1: Place the base; Step 2: Install the bracket" in the manual. For replaceable components, the model establishes equivalence relationships by comparing the part parameter table with the functional description; for example, two nuts with the same parameters are marked as interchangeable in R. After completing the graph structure construction, the system maps objects in O to object instances in PDDL, converts the sequence relationships in H into preconditions and postconditions for PDDL actions, and generates alternative action options from the equivalence relationships in R. Through this process, the originally unstructured task input is transformed into standardized symbolic descriptions, allowing the subsequent planning module to directly call the initial state and target constraints defined in PDDL, avoiding semantic bias and efficiency loss caused by manual configuration.

[0079] In one implementation of this embodiment, the visual language model can analyze images in the assembly manual or referenced in task instructions to identify the various components contained therein. Combined with text descriptions, it determines the category of each component, such as "bolt," "nut," "washer," "gear," and "bearing." The visual language model can infer the functional role of each component in the assembly by combining text descriptions (e.g., "install the motor onto the bracket," "use a sealing ring to prevent oil leakage") with the component positions in the images and their connection methods with other components. Examples of functional roles include "installed part," "mounting base," "fastener," "seal," and "transmission component." The visual language model can also analyze schematic diagrams or 3D model screenshots in the assembly manual to understand the spatial layout and connection relationships between components. For example, it can identify spatial structural relationships such as "component A is above component B," "component C is connected to component D via threads," and "component E surrounds component F." The visual language model can not only identify surface information but also understand the functional roles and spatial relationships of components. This makes the generated task execution logic more consistent with actual assembly requirements and reduces errors caused by misunderstandings.

[0080] Then, the system integrates the component categories, functional roles, and spatial structural relationships obtained from the previous step. For example, it knows that a "bolt" is a fastener, a "bracket" is a mounting base, and a "motor" is the component being installed. It also knows that the motor needs to be installed on the bracket, and the bolt is used to secure the motor and the bracket. Based on this integrated information, the system decomposes the entire assembly process into multiple logically independent subtasks according to basic assembly principles (e.g., installing the basic structure first, then the upper-level components; installing large components first, then small components; installing internal components first, then external components, etc.) and the analyzed spatial structural relationships. Each subtask focuses on completing a specific installation action. The system determines the reasonable execution order of these subtasks based on the spatial structural relationships between components and the physical constraints of the assembly. Ultimately, the constructed task execution logic is an ordered list containing multiple subtasks. Each subtask explicitly includes installation description information for at least one component (e.g., which component needs to be installed, where it is installed, and which other components it interacts with), as well as the execution order of that subtask within the entire sequence.

[0081] This embodiment introduces a visual language model to automatically parse complex graphic assembly information, greatly reducing the workload of manually understanding, annotating, and inputting assembly steps, thus improving efficiency. The visual language model can not only recognize surface information but also understand the functional roles and spatial relationships of components. This makes the generated task execution logic more in line with actual assembly needs and reduces errors caused by misunderstandings. The generated PDDL definition provides structured input for subsequent planning, improving the overall system's intelligence level and task adaptability.

[0082] In some embodiments, step S400 may include, but is not limited to, steps S410 to S430:

[0083] Step S410: Perform image segmentation on the target image to obtain segmented images of each assembly object in the target image;

[0084] Step S420: Perform grasping pose perception on the segmented image of each assembly object to obtain the pose information of each assembly object. The pose information includes the graspability score and grasping pose of each assembly object.

[0085] Step S430: Semantic analysis is performed on the segmented images and pose information of each assembly object using a visual language model to obtain the perception results. The perception results include the spatial relationships between each assembly object, the operational feasibility of each assembly object, and the attributes of each assembly object.

[0086] In this embodiment, an image segmentation algorithm (such as the Grounded SAM model) can be used to process the real-time acquired target image. The algorithm can identify and distinguish different assembly objects in the image (such as parts to be installed, installed components, tools, etc.), and generate a precise bounding box or pixel-level mask for each object, thereby obtaining segmented images of each assembly object. For example, if the target image contains a bolt and a nut, the segmentation algorithm will generate two segmented images representing the bolt and nut respectively, clearly defining their outlines.

[0087] In this embodiment, for each assembly object defined by the segmented image, grasping pose awareness is further performed. This is typically achieved through specialized pose estimation or grasping detection algorithms. These algorithms analyze the shape, edges, texture, and other features of the object in the segmented image to estimate the object's pose (i.e., position and orientation) in three-dimensional space, and evaluate the feasibility and stability of grasping the object from different angles. Pose information typically includes a graspability score and a grasping pose. The graspability score is a quantitative indicator representing the ease or success rate of grasping the object in the current state; the grasping pose is one or more suggested grasping points and their corresponding poses of the robotic arm's end effector (e.g., the center position and opening / closing direction of the gripper), which are considered the ideal position and angle for grasping the object.

[0088] In this embodiment, a Visual Language Model (VLM) is used, taking the segmented image (visual information) of each assembly object and its corresponding pose information as input. The VLM can understand the image content and, combined with the pose description, perform deeper semantic analysis. For example, the VLM (such as the GPT-4V model) can understand what the objects in the segmented image are (requiring association with component information in the assembly manual), analyze the spatial relationships between these objects (such as "the bolt is to the left of the nut" and "the washer is below the bolt"), determine whether the operation on the object is feasible under the current pose (such as "the bolt can be stably gripped by the gripper under the current pose" and "the nut is temporarily inoperable due to obstruction"), and extract or infer certain attributes of the object (which may include material, shape, size, and functional role, such as "the object is made of metal" and "the object needs to be tightened"). The final perception result is a rich set integrating visual, spatial, operational feasibility, and attribute information.

[0089] This information is converted into a standard symbol format using predefined predicate templates; for example, grabbability is mapped to `graspable(obj)`, and spatial location is mapped to `at(obj,loc)`. The perception-predicate mapping function encodes the above symbol set into an initial state for use by the task planning module. This process, through multi-model collaborative processing, transforms raw perception data into a structured symbol description, solving the problem of insufficient accuracy in perception-symbol mapping in traditional methods and providing reliable environmental state input for subsequent planning.

[0090] The specific implementation of this plan is as follows:

[0091] First, the robot's 3D camera acquires real-time images of the work scene, including color images and depth information, to support subsequent target detection, segmentation, and pose estimation.

[0092] Next, the Grounding DINO model is used to perform object detection on the image, obtaining bounding boxes and preliminary category labels for all candidate objects. Then, Grounded SAM is used to perform fine image segmentation on the regions within the detection boxes, obtaining semantic masks and contour boundaries for each object. The above segmentation results are input into the AnyGrasp model to calculate the graspability score and grasping pose for each object.

[0093] Furthermore, semantic information extraction is performed: the segmented image, object candidate boxes, and pose results are input into the visual language model GPT-4V, and semantic analysis guided by language is performed in combination with the scene context to extract spatial relationships between objects (such as "near", "above", "partially inserted"), operational feasibility (such as "can be grasped", "can be inserted"), and task-related attributes (such as "part A belongs to component X").

[0094] Finally, based on the predefined PDDL predicate template, the above perception results are uniformly mapped to a symbolic predicate set Lx={x1,x2,…,xM}, including formats such as graspable(obj), at(obj,loc), inserted(obj1,obj2), and type(obj,class), to describe the current environment state. The symbolic perception results are then encoded into a standard PDDL initial state format using the perception-predicate mapping function fpred(K(I)), which is then used by the behavior tree task planning module in subsequent steps.

[0095] This embodiment uses image segmentation and pose perception to accurately identify various objects and their states in the assembly environment. By combining pose information and semantic analysis, it can more intelligently determine whether an object can be operated by the current tool (such as a robotic arm gripper), avoiding invalid or dangerous attempts. Combining visual information (images, poses) with language description information provides a more comprehensive and intelligent perception of the assembly environment, providing high-quality data input for subsequent dynamic planning based on task logic and real-time perception results. By mapping the perception results to the standard PDDL format, seamless integration between the perception module and the planning module is achieved, enhancing the modularity and scalability of the system.

[0096] In some embodiments, step S410 may include, but is not limited to, steps S411 to S412:

[0097] Step S411: Perform target detection on the target image to obtain the bounding boxes and preliminary category labels of each assembly object in the target image;

[0098] Step S412: Perform image segmentation on the target image based on the bounding box and the preliminary category label of each assembly object to obtain segmented images of each assembly object. The segmented images include the outline boundary and target category label of each assembly object.

[0099] In this embodiment, to achieve high-precision identification and localization of assembly objects in the target image, a technical solution combining target detection and guided segmentation is adopted. First, the Grounding DINO model is used, combined with preset text prompts, to perform target detection on the target image. This model can output the bounding boxes of all candidate assembly objects in the image and their corresponding preliminary category labels.

[0100] Specifically, the Grounding DINO model is an object detection model that incorporates natural language guidance. Unlike traditional object detectors that rely solely on visual features, Grounding DINO can use textual cues (e.g., words like "screw," "nut," and "bracket" that might be involved in an assembly task) to guide the detection process. The Grounding DINO model scans the entire target image, identifying all potential objects related to the given text cues. For each identified candidate object in the image, the model outputs a bounding box that defines the object's approximate location within the image. Simultaneously, the model provides an initial category label, typically corresponding to the category name of the input text cue (e.g., "screw").

[0101] Subsequently, the bounding boxes are used as input to guide the Grounded SAM model to perform fine image segmentation on the regions covered by the bounding boxes. The Grounded SAM model generates high-precision semantic masks based on the visual features within the bounding boxes, thereby accurately obtaining the contour boundaries of each assembly object.

[0102] Specifically, the Grounded SAM model is a variant of the Segment Anything Model (SAM) specifically designed for accurate image segmentation based on given "grounding" information (i.e., bounding boxes and class labels provided by Grounding DINO). Taking bounding boxes as input, Grounded SAM utilizes visual features within the bounding boxes, combined with possible class information (or relying solely on the bounding box region), to accurately delineate the object's outline. For each input bounding box region, Grounded SAM generates a semantic mask. The semantic mask is a binary image (or an image with confidence levels), where pixels belonging to the object are labeled 1 (or high confidence), and pixels not belonging to the object are labeled 0 (or low confidence). The semantic mask accurately depicts the object's contour boundaries.

[0103] This embodiment effectively combines Grounding DINO's rapid localization and language guidance capabilities with Grounded SAM's high-precision segmentation capabilities, ensuring efficiency while providing accurate object shape and position information for subsequent pose perception and semantic analysis.

[0104] In some embodiments, step S500 may include, but is not limited to, steps S510 to S530:

[0105] Step S510: Convert the data format of the task execution logic and the perception result respectively to obtain a new task execution logic and a new perception result in the planning domain definition language format.

[0106] Step S520: Based on the new task execution logic and the new perception results, generate a behavior tree task architecture. The behavior tree task architecture includes multiple task nodes, including sequence nodes, condition nodes, and backoff nodes. The sequence nodes are used to determine the execution order of subtasks. The condition nodes are used to determine whether the current state meets the conditions for executing the current subtask. The backoff nodes are used to re-execute the current subtask or re-plan the assembly task when the current subtask fails to execute.

[0107] Step S530: Plan the assembly task according to the behavior tree task architecture to obtain the target assembly task.

[0108] In this embodiment, the task execution logic (task hierarchy graph structure) is converted into a task problem definition in PDDL format through structure mapping, providing structured input for subsequent symbolic planning and behavior tree generation. The structure mapping process converts nodes and edges in the graph structure into objects, initial states, and target states in PDDL; for example, "the base is located at the assembly station" is mapped to the PDDL predicate (at base station). Specifically, after completing the graph structure construction, the system maps objects in O to object instances in PDDL, sequence relations in H to preconditions and postconditions for PDDL actions, and equivalence relations in R to generate alternative action options. Through this process, the originally unstructured task input is transformed into a standardized symbolic description, allowing the subsequent planning module to directly call the initial state and target constraints defined in PDDL, avoiding semantic bias and efficiency loss caused by manual configuration.

[0109] In this embodiment, for information about the assembly environment (i.e., perception results) acquired from real-time sensing (such as cameras and sensors), the system uses a predefined PDDL predicate template to uniformly map these perception data into a set of symbolic predicates. These predicates formally describe various states of the current environment, such as (above), (next to), (grabbing), (idle), etc. This symbolic set of perception results is encoded into a standard PDDL initial state format through a "perception-predicate mapping function".

[0110] In this embodiment, by converting the data format of the task execution logic and the perception results, it is ensured that the high-level planning instructions (task execution logic) at the task level and the low-level environmental information (perception results) at the perception level can be used by the subsequent behavior tree task architecture generation module in a compatible manner.

[0111] In this embodiment, based on the transformed new task execution logic and the new perception results, the system generates a task architecture, namely a behavior tree. Specifically, the system generates a behavior tree structure by combining a large language model with a task skill library, based on the natural language task instructions, PDDL knowledge domain, and perception state.

[0112] Specifically, natural language task instructions are assembly task requirements given by the user or system in natural language form (e.g., "Please install the blue nut onto the red bracket"), serving as high-level semantic input to the task objective; the PDDL knowledge domain defines the object types, predicates, and available skill sets within the task domain; and the perceived state describes the current environmental state through a set of symbolic predicates.

[0113] Specifically, the behavior tree includes the following types of nodes: sequence nodes, condition nodes, and backoff nodes. Sequence nodes represent the pre-defined execution order of subtasks within the task execution logic. A sequence node contains multiple sub-nodes (usually other task nodes or specific execution actions), which are executed sequentially according to a predetermined order. For example, if the task logic specifies installing component A before component B, the sequence node ensures that the installation of component A is performed before component B. Condition nodes perform status checks before executing a specific subtask. Based on the information provided in the current perception results, they determine whether the preconditions for executing the subtask are met. For example, before attempting to install a nut, the condition node checks whether the corresponding bolt in the perception results is in place and ready for installation. If the conditions are met, the subsequent subtask is executed; otherwise, the subtask may be skipped or other processing logic (such as waiting or adjustment) may be triggered. This reflects the dynamic constraints and guidance of perception results on task execution. Backoff nodes handle potential failures during task execution, enhancing the robustness of the planning method. When a subtask fails (e.g., a grasping failure, incorrect installation location, etc., which may be detected through sensor feedback or execution results), a fallback node is activated. It can perform pre-defined recovery actions, such as retrying the current subtask (which may require adjusting pose or clearing interference first), performing cleanup operations, or, if necessary, triggering a more complex replanning process, invoking a higher-level planner to re-plan subsequent task steps based on the latest perception results. This ensures that the assembly process can continue or be safely terminated even in non-ideal environments.

[0114] In this embodiment, based on the generated behavior tree task architecture, the system plans specific assembly tasks and outputs the final target assembly task. For each specific subtask determined in the behavior tree task architecture (e.g., grasping a component, moving it to a designated location, or performing an assembly action), the system combines the latest perception results (the precise pose and reachability of the component) to calculate the specific path, end effector posture, force, and other control parameters required for the robot to perform the action. The system executes sequentially along the nodes of the behavior tree task architecture, making real-time judgments when encountering conditional nodes and handling exceptions when encountering backtracking nodes. The entire process is dynamic and iteratively optimized, continuously feeding the latest perception information back into the planning. Finally, the planning result is transformed into a specific instruction sequence that the robot controller can understand, i.e., the target assembly task. This task instruction sequence not only includes the actions to be performed but also the timing of these actions, condition judgment logic, and exception handling contingency plans.

[0115] The specific implementation of this plan is as follows:

[0116] First, input the natural language task instruction i, the PDDL knowledge domain D, and the perception state P. For example, the task instruction i is "insert part A into part B", the PDDL knowledge domain D defines the set of available robot skills, and the perception state P contains the position information of parts A and B in the current scene.

[0117] Secondly, a behavior tree structure B is generated by combining a large language model with a task skill library Lψ. Specifically, the task instruction i, knowledge domain D, and state P are taken as input and parsed and reasoned through a large language model such as GPT-4, outputting a behavior tree structure B containing sequential execution, conditional judgment, and backtracking nodes.

[0118] Then, each leaf node of the behavior tree is bound to an atomic skill ψi defined in PDDL. The format of the atomic skill ψi is (name, params, ψp, ψe, φ), where name is the skill name, params is the parameter list, ψp represents the preconditions, ψe is the skill effect, and φ is the underlying control interface. For example, the "grasp" skill can be defined as (grasp, [obj], {at(robot,obj_loc), graspable(obj)}, {holding(obj)}, grasp_control).

[0119] Finally, during behavior tree execution, the robot's task execution and failure recovery mechanisms are dynamically driven based on the node status return values. Node states include three types: Success, Failure, and Running. When a child node returns Failure, the parent node's rollback strategy is triggered; when all child nodes return Success, the task is completed.

[0120] This embodiment achieves information fusion through data format conversion and introduces logic and fault tolerance mechanisms through a behavior tree task architecture, ultimately generating specific and feasible target assembly tasks. This improves the intelligence level and adaptability of autonomous assembly task planning for industrial robots based on a multi-agent large model. It can not only follow the preset assembly process, but also respond to environmental changes in real time through perception results. By using condition nodes and backoff nodes in the behavior tree task architecture for condition judgment and exception handling, the probability of successful assembly is increased and the need for manual intervention is reduced. It is especially suitable for automated assembly scenarios with complex environments and changing object states.

[0121] In some embodiments, steps S600 to S700 may be included after step S500:

[0122] Step S600: Capture the assembly environment during the assembly of the assembly object to obtain a video.

[0123] Step S700: Perform image analysis on the captured video. If it is determined that the target assembly task has failed actions or the action deviation exceeds a preset threshold during execution, then re-execute the current sub-task or jump to the step of converting the task execution logic and the perception result into a planning domain definition language format, and perform assembly task planning through a behavior tree to obtain the step execution of the target assembly task.

[0124] In this embodiment, the system monitors the assembly site during the assembly task execution. Using cameras or other visual sensors in the assembly environment, it captures images of the assembly object being assembled or about to be assembled, along with its surroundings, thereby acquiring one or more video clips. After acquiring the videos, image analysis is performed to determine whether the target assembly task (or the currently executing sub-task) is proceeding as expected. Specifically, the analysis system checks whether the movements of the assembly equipment (such as robotic arms, grippers, etc.) in the video conform to the planned path, whether the components on the assembly object are correctly grasped, moved, or installed, and whether there are any unexpected interferences or errors (e.g., component position offset, grasping failure, collision with other objects, etc.).

[0125] In this embodiment, the system compares the analysis results with preset "success" criteria or "acceptable deviation" ranges. If it is determined that the target assembly task has failed in execution or the deviation exceeds a preset threshold, the system will proceed. A failure can mean that the assembly equipment fails to complete the intended action (e.g., gripping failure, inaccurate placement), or that the state of the assembly object does not meet expectations (e.g., screws are not tightened, parts are not aligned). A deviation exceeding the preset threshold means that the difference between the actual movement trajectory, speed, force, etc., of the assembly equipment and the planned or ideal value exceeds the system's allowable range. This threshold is preset based on the accuracy requirements of the assembly task and the physical characteristics of the assembly object. Once the image analysis results indicate that there is indeed a "failure" or "deviation exceeding the threshold," the system will not continue to execute subsequent tasks as originally planned, but will trigger corresponding corrective measures. If it is determined that the current subtask only has a correctable, non-fatal error or deviation during execution, the system can choose to stop the current action of the assembly equipment and then attempt to re-execute the same subtask. Before re-execution, the system may make minor parameter adjustments or path corrections based on the cause of the failure. If image analysis or subsequent judgment (possibly combined with other sensor information) deems the current problem serious, or if simple retries may be ineffective or even lead to worse results, the system will choose to interrupt the execution of the current subtask and "jump to the step execution of converting the task execution logic and the perception results into a planning domain definition language format, and performing assembly task planning through a behavior tree to obtain the target assembly task." This means that the system will return to the starting point of task planning or a node before the current subtask, utilize the latest perception results (possibly including new information obtained from failures) and task logic, and re-perform task planning one or more times to generate a revised target assembly task that may contain different strategies or paths, and then attempt to execute it again.

[0126] This embodiment, by introducing a real-time monitoring and dynamic adjustment process based on video image analysis, can significantly improve the assembly system's ability to respond to execution errors, avoid error accumulation, and ensure that the assembly task can be completed as expected as possible even under non-ideal conditions, thereby improving the reliability, robustness, and automation level of the entire assembly process.

[0127] In some embodiments, step S700 may include, but is not limited to, steps S710 to S740:

[0128] Step S710: Perform image analysis on the captured video using a visual language model to obtain the execution result of the current subtask;

[0129] Step S720: Obtain physical sensor data of the assembly equipment that assembles the assembly object to obtain verification data, which includes gripper status and force feedback data.

[0130] Step S730: Verify the execution result of the subtask using the verification data to obtain the verification result;

[0131] Step S740: If the verification result indicates that the target assembly task has failed an action or the action deviation exceeds a preset threshold during execution, then the current subtask is re-executed or the process is redirected to convert the task execution logic and the perception result into a planning domain definition language format, and the assembly task is planned through a behavior tree to obtain the step execution of the target assembly task.

[0132] In this embodiment, the process of performing image analysis on the captured video to determine the current subtask execution status is specifically accomplished through a visual language model. A visual language model is an advanced artificial intelligence model capable of simultaneously understanding and processing image and text information. The system inputs the captured video clips (or keyframes) into a pre-trained visual language model. This model can not only identify visual information such as objects, positions, and postures in the video frame, but also understand the relationship between this information and textual knowledge such as assembly task instructions and assembly manuals. Through this cross-modal understanding capability, the visual language model can generate descriptive judgments about the current subtask execution result, i.e., the "current subtask execution result." For example, the model might determine that "the screw has been successfully screwed in but not fully seated," "part A has not been correctly aligned with the slot of part B," or "the gripper has grasped the target part."

[0133] After obtaining the current subtask execution result output by the visual language model, the system also simultaneously acquires data collected by physical sensors installed on the assembly equipment (such as industrial robots, robotic arms, etc.). This sensor data constitutes important "verification data," specifically including but not limited to the gripper's state information (such as gripping / releasing state, current gripping force, gripping position coordinates, etc.) and force feedback data (such as the magnitude, direction, and distribution of the force generated when the gripper contacts the object). This physical sensor data directly reflects the physical interaction between the assembly equipment and the assembly object, and is used to verify whether the task execution was successful.

[0134] Finally, based on the "verification result" that integrates visual language model judgment and physical sensor data verification, the system makes a final decision. If the verification result clearly indicates that the target assembly task has indeed experienced action failures (such as parts not being picked up, incorrect placement, etc.) or action deviations (such as slight offset of part position, insufficient tightening force, etc.) exceeding the preset safety or accuracy threshold during execution, the system will trigger the corresponding correction mechanism, instructing the assembly equipment to re-execute the currently unsuccessful sub-task, hoping to correct the error in the retry. If the error is more serious or the retry may be ineffective, the system jumps back to the task planning stage, that is, re-executes the step of "converting the task execution logic and the perception result into a planning domain definition language format, and performing assembly task planning through a behavior tree to obtain the target assembly task". During replanning, the system can use the latest perception results (including information obtained from failed attempts) and task logic to adjust strategies, correct paths or parameters, generate a corrected task plan, and then execute it.

[0135] Specifically, the system can determine whether an operation has achieved the expected result through visual question answering, and perform cross-validation by combining data from low-level sensors such as gripper status and force feedback. If an action failure or deviation exceeding a threshold is detected, the system will automatically trigger a backtracking strategy in the behavior tree or replan the path. The visual question answering method can perform semantic judgment on the current scene image through a visual language model to extract the matching degree between the operation result and the expected target; the physical verification method can collect clamping force and contact force data through a six-dimensional force sensor and compare them with preset safety thresholds; the cross-validation mechanism comprehensively evaluates the visual semantic judgment results and physical sensor data through a weighted fusion algorithm.

[0136] The specific implementation of this plan is as follows:

[0137] First, the robot acquires images of the current scene using a high-resolution camera. Then, the image is input into a pre-trained visual language model, GPT-4V, which performs semantic judgment on the operation results through visual question answering. For example, the system can ask questions such as "Has part A been fully inserted into part B?" or "Is the orientation of component C correct?", and the model provides answers based on the image content.

[0138] Simultaneously, the system also incorporates data from the torque sensor and position encoder of the robot's end effector to quantitatively analyze force feedback and positional deviations during operation. For example, during insertion operations, if abnormal resistance or positional deviation is detected, the system will consider it a potential execution failure.

[0139] Furthermore, the system cross-validates the visual semantic judgment results with physical sensing data. If there is a significant inconsistency between the two, or if the display operation on either side fails to achieve the expected results, the system will automatically trigger a failure handling mechanism.

[0140] Therefore, if an action execution failure or deviation exceeding a preset threshold is detected, the system will first attempt to activate a predefined rollback strategy in the behavior tree. For example, if part insertion fails, the system may execute a rollback sequence of "extract - adjust posture - re-insert". If multiple rollbacks still fail to resolve the issue, the system will re-invoke the task planning module to generate a new execution path based on the current state.

[0141] Throughout the execution process, the system continuously records key metrics such as the success rate, time consumption, and accuracy of each operation. This data will be used for subsequent task optimization, such as adjusting operation parameters and optimizing planning strategies, thereby continuously improving the overall performance of the system.

[0142] This embodiment significantly improves the system's adaptability to complex environmental changes by integrating visual semantic understanding and physical sensing data; through multimodal verification and real-time feedback mechanisms, the system can promptly detect and correct execution deviations, effectively reducing the task failure rate; the introduction of dynamic replanning and rollback strategies enhances the system's fault tolerance and execution stability, enabling it to cope with various unexpected situations.

[0143] This application embodiment uses a Visual Language Model (VLM) to parse assembly drawings or natural language instructions, extracting task objects, semantic roles, and spatial structural relationships to construct a task hierarchy graph. Combining 3D images, segmentation algorithms, and grasping estimators, it extracts object categories, poses, and operational attributes, generating a PDDL-formatted predicate set to describe the environmental state. Using a Large Language Model (LLM) to combine instructions, task graphs, and environmental state information, it automatically generates a behavior tree structure to organize the task execution sequence. The behavior tree supports conditional control, backoff mechanisms, and modular execution. During task execution, the system integrates visual verification and sensor feedback to automatically determine the success or failure of actions. In case of failure, it triggers alternative paths or replanning to ensure the stability and accuracy of task completion. This application embodiment, by integrating a Visual Language Model and a Large Language Model, constructs a robot long-term task planning system that supports complex semantic parsing, multimodal perception, and behavior tree-driven execution. This system possesses high intelligence, adaptability, and modularity, significantly improving the execution stability and operational efficiency of industrial robots in multi-task, multi-step, and dynamic environments. Leveraging the general semantic understanding capabilities of a pre-trained large-scale model, the system can parse new natural language instructions or assembly drawings and automatically generate executable task plans without requiring specific training for each task. This significantly expands the robot's adaptability in frequently changing or unstructured production lines. By extracting task objects, spatial relationships, and operation sequences through a visual language model, the system can automatically construct a task graph structure, identify hierarchical dependencies and concurrent relationships of tasks, and is suitable for assembly processes with multiple parts, multiple processes, and high precision requirements, meeting the complex needs of task modeling and planning in flexible and customized production scenarios. The system uses behavior trees as the task representation carrier and supports control logic such as sequential execution, conditional judgment, and failure rollback. During task execution, if execution fails due to environmental changes or perception errors, backup strategies or replanning can be automatically triggered to ensure uninterrupted task completion, effectively improving the system's fault tolerance and execution stability. The system significantly outperforms traditional state machine-based or template-driven robot methods in key indicators such as task completion rate, execution time, and repeatability. It exhibits stronger flexibility and robustness, especially in complex, multi-step tasks, and has good prospects for engineering application.

[0144] Please see Figure 3 This application also provides an industrial robot autonomous assembly task planning device 800 based on a multi-agent large model, which can implement the above-mentioned industrial robot autonomous assembly task planning method based on a multi-agent large model. The industrial robot autonomous assembly task planning device 800 based on a multi-agent large model includes:

[0145] The acquisition module 10 is used to acquire assembly task data of the assembly object, the assembly task data including task instructions and assembly manuals for assembling the assembly object.

[0146] The parsing module 20 is used to parse the task instructions and the assembly manual to obtain the task execution logic;

[0147] The acquisition module 30 is used to acquire the assembly environment in real time when the assembly object is assembled, and obtain the target image;

[0148] The perception module 40 is used to perform multimodal perception on the target image and obtain a perception result, the perception result including the position of at least one component of the assembly object in the assembly environment;

[0149] The planning module 50 is used to convert the task execution logic and the perception results into a planning domain definition language format, and to plan the assembly task through a behavior tree to obtain the target assembly task.

[0150] In some implementations, the parsing module 20 may include:

[0151] The parsing submodule is used to perform semantic parsing of the task instructions and the assembly manual through a visual language model to obtain the category of each component, the functional role of each component, and the spatial structural relationship between each component.

[0152] A submodule is constructed to build the task execution logic based on the category of each component, the functional role of each component, and the spatial structural relationship between each component. The task execution logic includes multiple subtasks, and each subtask includes installation description information of at least one component and the execution order of the subtasks.

[0153] In some implementations, the sensing module 40 may include:

[0154] The image segmentation submodule is used to segment the target image to obtain segmented images of each assembly object in the target image;

[0155] The pose perception submodule is used to perform grasping pose perception on the segmented image of each assembly object to obtain the pose information of each assembly object. The pose information includes the graspability score and grasping pose of each assembly object.

[0156] The semantic analysis submodule is used to perform semantic analysis on the segmented images and pose information of each assembly object through a visual language model to obtain the perception results. The perception results include the spatial relationships between each assembly object, the operational feasibility of each assembly object, and the attributes of each assembly object.

[0157] In some implementations, the image segmentation submodule may include:

[0158] The target detection unit is used to perform target detection on the target image to obtain the bounding boxes and preliminary category labels of each assembly object in the target image;

[0159] An image segmentation unit is used to segment the target image based on the bounding box and the preliminary category label of each assembly object to obtain a segmented image of each assembly object, wherein the segmented image includes the outline boundary and target category label of each assembly object.

[0160] In some implementations, the planning module 50 may include:

[0161] The conversion submodule is used to convert the data format of the task execution logic and the perception result respectively, so as to obtain a new task execution logic and a new perception result in the planning domain definition language format.

[0162] A generation submodule is used to generate a behavior tree task architecture based on new task execution logic and new perception results. The behavior tree task architecture includes multiple task nodes, including sequence nodes, condition nodes, and backoff nodes. The sequence nodes are used to determine the execution order of subtasks. The condition nodes are used to determine whether the current state meets the conditions for executing the current subtask. The backoff nodes are used to re-execute the current subtask or re-plan the assembly task when the current subtask fails to execute.

[0163] The planning submodule is used to plan the assembly task based on the behavior tree task architecture to obtain the target assembly task.

[0164] In some embodiments, the device further includes:

[0165] The shooting module is used to shoot the assembly environment when the assembly object is assembled, and obtain the shooting video;

[0166] The analysis module is used to perform image analysis on the captured video. If it is determined that the target assembly task has failed actions or the action deviation exceeds a preset threshold during execution, the current subtask is re-executed or the process is redirected to convert the task execution logic and the perception results into a planning domain definition language format, and the assembly task is planned through a behavior tree to obtain the steps of the target assembly task.

[0167] In some implementations, the analysis module may include:

[0168] The image analysis submodule is used to perform image analysis on the captured video using a visual language model to obtain the execution result of the current subtask.

[0169] The acquisition submodule is used to acquire physical sensor data of the assembly equipment that assembles the assembly object, and obtain verification data, which includes gripper status and force feedback data.

[0170] The verification submodule is used to verify the execution result of the subtask using the verification data, and obtain the verification result;

[0171] The execution submodule is used to re-execute the current sub-task or jump to the step execution of the target assembly task if the verification result indicates that the target assembly task has failed an action or the action deviation exceeds a preset threshold during execution. The task execution logic and the perception result are converted into a planning domain definition language format, and the assembly task is planned through a behavior tree to obtain the step execution of the target assembly task.

[0172] The specific implementation of the industrial robot autonomous assembly task planning device based on the multi-agent large model is basically the same as the specific implementation of the industrial robot autonomous assembly task planning method based on the multi-agent large model described above, and will not be repeated here.

[0173] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method for autonomous assembly task planning of industrial robots based on a multi-agent large model. This electronic device can be any intelligent terminal, including tablet computers, in-vehicle computers, etc.

[0174] Please see Figure 4 , Figure 4 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes:

[0175] The processor 801 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0176] The memory 802 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 802 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 802 and is called and executed by the processor 801 to implement the autonomous assembly task planning method for industrial robots based on a multi-agent large model according to the embodiments of this application.

[0177] The 803 input / output interface is used to implement information input and output.

[0178] The communication interface 804 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0179] Bus 805 transmits information between various components of the device (e.g., processor 801, memory 802, input / output interface 803, and communication interface 804);

[0180] The processor 801, memory 802, input / output interface 803, and communication interface 804 are connected to each other within the device via bus 805.

[0181] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for autonomous assembly task planning of industrial robots based on a multi-agent large model.

[0182] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0183] The present application provides an industrial robot autonomous assembly task planning method, device, electronic device, and storage medium based on a multi-agent large model. It parses assembly task instructions and operation manuals using a visual language model to extract task objects, roles, and spatial relationships, constructing task execution logic. It uses a multimodal perception system to perceive the 3D image of the assembly environment, ensuring the robot has real-time awareness of the current scene. It generates a symbolic behavior tree with execution conditions and backoff strategies using a large language model, with each leaf node corresponding to a skill defined in the robot's skill library, and dynamically schedules tasks based on the node's return status. During execution, it integrates visual and sensor data for multimodal verification, using visual question answering and physical sensor data for state judgment, achieving execution monitoring and feedback adjustment. By constructing a task hierarchy graph structure, multimodal perception and symbolic mapping, dynamic behavior tree planning, and a multimodal verification mechanism, this application achieves semantic parsing of unstructured task instructions, symbolic representation of environmental states, and dynamic task adjustment capabilities, improving autonomous decision-making capabilities, realizing a closed-loop multimodal information fusion, and enhancing the dynamic adjustment capabilities for complex assembly tasks.

[0184] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0185] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0186] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0187] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0188] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0189] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0190] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0191] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0192] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0193] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0194] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for autonomous assembly task planning of industrial robots based on a multi-agent large model, characterized in that, The method includes: Acquire assembly task data for the assembly object, wherein the assembly task data includes task instructions and assembly manuals for assembling the assembly object; The task instructions and the assembly manual are parsed to obtain the task execution logic; The assembly environment during the assembly of the assembly object is captured in real time to obtain the target image; Multimodal perception is performed on the target image to obtain perception results, the perception results including the position of at least one component of the assembly object in the assembly environment; The task execution logic and the perception results are converted into a planning domain definition language format, and the assembly task is planned through behavior tree to obtain the target assembly task. The process of converting the task execution logic and the perception results into a planning domain definition language format, and then using a behavior tree to plan the assembly task to obtain the target assembly task includes: The task execution logic and the perception results are converted into data formats to obtain new task execution logic and new perception results in the planning domain definition language format. Based on the new task execution logic and the new perception results, a behavior tree task architecture is generated. The behavior tree task architecture includes multiple task nodes, including sequence nodes, condition nodes, and backoff nodes. The sequence nodes are used to determine the execution order of subtasks. The condition nodes are used to determine whether the current state meets the conditions for executing the current subtask. The backoff nodes are used to re-execute the current subtask or re-plan the assembly task when the current subtask fails to execute. The assembly task is planned based on the behavior tree task architecture to obtain the target assembly task.

2. The method according to claim 1, characterized in that, The process of parsing the task instructions and the assembly manual to obtain the task execution logic includes: The task instructions and assembly manual are semantically parsed using a visual language model to obtain the category of each component, the functional role of each component, and the spatial structural relationship between each component. The task execution logic is constructed based on the category of each component, the functional role of each component, and the spatial structural relationship between each component. The task execution logic includes multiple sub-tasks, and each sub-task includes installation description information of at least one component and the execution order of the sub-tasks.

3. The method according to claim 1, characterized in that, The process of performing multimodal perception on the target image to obtain the perception result includes: The target image is segmented to obtain segmented images of each assembly object in the target image; The segmented image of each assembly object is subjected to grasping pose perception to obtain the pose information of each assembly object, which includes the graspability score and grasping pose of each assembly object. The perception results are obtained by performing semantic analysis on the segmented images and pose information of each assembly object using a visual language model. The perception results include the spatial relationships between each assembly object, the operational feasibility of each assembly object, and the attributes of each assembly object.

4. The method according to claim 3, characterized in that, The step of segmenting the target image to obtain segmented images of each assembly object in the target image includes: Target detection is performed on the target image to obtain the bounding boxes and preliminary category labels of each assembly object in the target image; The target image is segmented based on the bounding box and the preliminary category label of each assembly object to obtain a segmented image of each assembly object. The segmented image includes the outline boundary and target category label of each assembly object.

5. The method according to claim 1, characterized in that, After converting the task execution logic and the perception result into a planning domain definition language format, and performing assembly task planning through behavior trees to obtain the target assembly task, the method further includes: The assembly environment during the assembly of the assembled object was filmed to obtain a video. Image analysis is performed on the captured video. If it is determined that the target assembly task has failed actions or the action deviation exceeds a preset threshold during execution, the current subtask is re-executed or the process is redirected to convert the task execution logic and the perception result into a planning domain definition language format. Then, the assembly task is planned using a behavior tree to obtain the steps of the target assembly task.

6. The method according to claim 5, characterized in that, The step involves image analysis of the captured video. If it is determined that the target assembly task has failed actions or deviated from a preset threshold during execution, the current subtask is re-executed or the process jumps to the step of converting the task execution logic and the perception results into a planning domain definition language format, and then using a behavior tree to plan the assembly task to obtain the steps for executing the target assembly task, including: The captured video is analyzed using a visual language model to obtain the execution result of the current subtask. Acquire physical sensor data of the assembly equipment that assembles the assembly object to obtain verification data, which includes gripper status and force feedback data. The execution result of the subtask is verified using the verification data to obtain the verification result; If the verification result indicates that the target assembly task has failed an action or the action deviation exceeds a preset threshold during execution, then the current subtask is re-executed or the process jumps to the step execution of the target assembly task by converting the task execution logic and the perception result into a planning domain definition language format and planning the assembly task through a behavior tree.

7. An autonomous assembly task planning device for industrial robots based on a multi-agent large model, characterized in that, The device includes: The acquisition module is used to acquire assembly task data of the assembly object, the assembly task data including task instructions and assembly manuals for assembling the assembly object. The parsing module is used to parse the task instructions and the assembly manual to obtain the task execution logic; The acquisition module is used to acquire the assembly environment in real time when the assembly object is assembled, and obtain the target image. A perception module is used to perform multimodal perception on the target image and obtain a perception result, the perception result including the position of at least one component of the assembly object in the assembly environment; The planning module is used to convert the task execution logic and the perception results into a planning domain-defined language format, and to plan the assembly task using a behavior tree to obtain the target assembly task. It then performs data format conversion on the task execution logic and the perception results to obtain new task execution logic and new perception results in planning domain-defined language format. Based on the new task execution logic and the new perception results, it generates a behavior tree task architecture, which includes multiple task nodes, including sequence nodes, condition nodes, and backoff nodes. The sequence nodes determine the execution order of subtasks, the condition nodes determine whether the current state meets the conditions for executing the current subtask, and the backoff nodes re-execute the current subtask or replan the assembly task if the current subtask fails. Finally, it plans the assembly task according to the behavior tree task architecture to obtain the target assembly task.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the autonomous assembly task planning method for industrial robots based on a multi-agent large model as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the autonomous assembly task planning method for industrial robots based on a multi-agent large model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Robot assembly task planning method and system based on large language model

    CN118348983A

  • Man-machine interaction assembly method and system based on multi-modal large model and reinforcement learning

    CN118744426A