Assembly mechanical arm task planning method and device

The assembly robot task planning method leverages multi-modal data and large language models to enhance task understanding, feasibility, and adaptability by generating optimal action sequences, addressing limitations in traditional methods.

CN120307305AActive Publication Date: 2025-07-15YANGTZE DEITA GRADUATE SCHOOI OF BEIJING INST OF TECH (JIAXING) +1

Patent Information

Application Number
CN202510803738.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Traditional assembly robotic arm task planning has problems such as insufficient task understanding and decomposition capabilities, low feasibility and weak dynamic adaptability. It is difficult to analyze multimodal task instructions, separate planning and execution, and the strategy cannot be adjusted in real time.

Method used

Multimodal data processing and large language model are used to combine pre-trained action databases. By obtaining the image and text description of the assembly robot arm, joint feature encoding is determined, assembly sub-task sequences are generated using the assembly knowledge base and large language model, action parameters are optimized and optimal action sequences are generated, and action parameters are adjusted in real time to adapt to changes in the assembly environment.

Benefits of technology

The task understanding and decomposition capabilities of assembly robotic arm task planning are improved, feasibility and adaptability are enhanced, and efficient automation and dynamic adjustment of assembly tasks are achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120307305A_ABST
    Figure CN120307305A_ABST
Patent Text Reader

Abstract

The invention discloses an assembly mechanical arm task planning method and device, and relates to the technical field of assembly task planning, and the method comprises the steps: obtaining multi-modal data of an assembly mechanical arm, and determining a joint feature code, thereby determining an assembly retrieval result; determining a plurality of assembly subtask sequences based on the multi-modal data, the assembly retrieval result and a large language model; based on a pre-training action library, determining a plurality of feasible action sequences of each assembly subtask; based on the similarity scores and the optimal feasibility scores corresponding to all the actions in each feasible action sequence, determining the total score of the corresponding feasible action sequence so as to obtain an optimal action sequence; determining an optimal action parameter sequence based on a pre-training action library and the optimal action sequence; and based on the pre-training action library, the optimal action parameter sequence is converted into the movement of the assembly mechanical arm, and an assembly task is carried out. The task understanding and decomposition capability, feasibility and adaptability of task planning of the assembly mechanical arm are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of assembly task planning, and particularly to an assembly robotic arm task planning method and device. Background Art

[0002] Traditional assembly robotic arm task planning mainly relies on manually predefined rules or traditional optimization algorithms, and has three major limitations: (1) Insufficient task understanding and decomposition ability: It is difficult to parse multi-modal task instructions (such as text, images), resulting in a mismatch between the subtask sequence and the physical scenario, and manual adjustment is required repeatedly. (2) Disconnection between planning and execution, low feasibility: The abstract instructions generated by traditional planning require additional parameter calibration, and the lack of automatic optimization of action parameters (such as grasping poses, force control parameters) has a high failure rate, resulting in low feasibility of assembly robotic arm task planning. (3) Weak dynamic adaptability: In the face of changes in the assembly environment (such as part offset, fixture loosening), traditional methods cannot adjust the planning strategy in real time, and task interruption is likely to occur. Summary of the Invention

[0003] The purpose of this application is to provide an assembly robotic arm task planning method and device to solve the problems of insufficient task understanding and decomposition ability, low feasibility, and weak adaptability.

[0004] To achieve the above purpose, this application provides the following solutions: In the first aspect, this application provides an assembly robotic arm task planning method, including: Obtain multi-modal data of the assembly robotic arm; the multi-modal data includes: images of the assembly site and text descriptions of the assembly task; Based on the image and the text description, determine a joint feature encoding; Based on the joint feature encoding and the assembly knowledge base, determine an assembly retrieval result; Based on the multi-modal data, the assembly retrieval result, and a large language model, determine multiple assembly subtask sequences of the assembly task; Based on a pre-trained action library, respectively determine multiple executable action sequences for each assembly subtask in each assembly subtask sequence; Respectively determine the similarity score and the optimal feasibility score corresponding to each action in each executable action sequence; Based on the similarity scores and the optimal feasibility scores corresponding to all actions in each executable action sequence, respectively determine the total score of the corresponding executable action sequence; Determine the executable action sequence with the maximum total score as the optimal action sequence of the assembly robotic arm; Based on the pre-trained action library and the optimal action sequence, determine an optimal action parameter sequence; Based on the pre-trained action library, the optimal action parameter sequence is transformed into the motions of the joints of the assembly robot arm to perform the assembly task.

[0005] In one embodiment, before obtaining the multi-modal data of the assembly robot arm, it further includes: defining a pre-trained action library and an assembly knowledge base.

[0006] In one embodiment, the pre-trained action library includes: the action parameters of multiple actions, the strategy for selecting the optimal action parameters for each action in any environment, and the value function; The assembly knowledge base includes: a component ontology library, an assembly relationship graph, and an assembly process knowledge tree; The component ontology library is used to define the geometric attributes, physical attributes, and functional attributes of parts; The assembly relationship graph is used to describe the assembly relationships between various parts; The assembly process knowledge tree is used to represent the sequence of assembly processes.

[0007] In one embodiment, based on the image and the text description, determining the joint feature encoding includes: Preprocessing the image to obtain an initial image vector; the preprocessing includes: image block division, block embedding, and position encoding; Performing feature extraction on the initial image vector to obtain a visual feature vector; Encoding the text description through a BERT model to obtain a text feature vector; Performing weighted integration on the visual feature vector and the text feature vector to obtain the joint feature encoding.

[0008] In one embodiment, based on the joint feature encoding and the assembly knowledge base, determining the assembly retrieval result includes: Filtering out all parts from the part ontology database whose cosine similarity with the joint feature encoding is greater than a first preset threshold to obtain multiple relevant parts; Based on the assembly relationship graph, determining the assembly relationships between the relevant parts; Based on the assembly process knowledge tree and the assembly relationships between the relevant parts, determining the assembly process sequence; Determining the relevant parts, the assembly relationships between the parts in the relevant parts, and the assembly process sequence as the assembly retrieval result.

[0009] In one embodiment, based on the pre-trained action library, respectively determining multiple executable action sequences for each assembly subtask in each assembly subtask sequence includes: Determining any assembly subtask as the subtask to be determined; Determine the cosine similarity between the sub-task to be determined and each action in the pre-trained action library respectively; Input all actions with the cosine similarity corresponding to the sub-task to be determined greater than the second preset threshold and the assembly retrieval result into the large language model to obtain multiple executable action sequences for the sub-task to be determined.

[0010] In one embodiment, determining the similarity score and the optimal feasibility score corresponding to each action in each executable action sequence respectively includes: Determine any assembly sub-task as the current sub-task, determine any executable action sequence of the current assembly sub-task as the current action sequence, and determine any action in the current action sequence as the current action; Calculate the similarity score between the current action and the current sub-task, and determine the similarity score corresponding to the current action as the similarity score between the current action and the current sub-task; Based on the policy in the pre-trained action library, determine multiple action parameter sequences of the current action sequence; Based on the value function in the pre-trained action library, determine the feasibility scores of the current action under each action parameter sequence respectively; Based on the feasibility scores of the current action under each action parameter sequence, determine the optimal feasibility score corresponding to the current action.

[0011] In one embodiment, determining the total score of the corresponding executable action sequence respectively based on the similarity scores and the optimal feasibility scores corresponding to all actions in each executable action sequence includes: Multiply the similarity score and the optimal feasibility score corresponding to each action in each executable action sequence respectively to obtain the comprehensive score of the corresponding action; Multiply the comprehensive scores of all actions in each executable action sequence respectively to obtain the total score of the corresponding executable action sequence.

[0012] In one embodiment, determining the optimal action parameter sequence based on the pre-trained action library and the optimal action sequence includes: Based on the current environment at the assembly site, the first action in the optimal action sequence and the policy, determine the optimal action parameters of the first action in the optimal action sequence; Based on the optimal action parameters of the first action in the optimal action sequence, determine the optimal action parameters of each action in the optimal action sequence in turn, so as to determine the optimal action parameter sequence.

[0013] In a second aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the assembly robot task planning method described in any one of the above.

[0014] According to the specific embodiments provided in this application, the following technical effects are disclosed in this application: This application discloses an assembly robot task planning method and device. First, multi-modal data of the assembly robot is obtained; the multi-modal data includes: images of the assembly site and text descriptions of the assembly tasks; based on the images and text descriptions, a joint feature encoding is determined; based on the joint feature encoding and the assembly knowledge base, an assembly retrieval result is determined; then, based on the multi-modal data, the assembly retrieval result, and a large language model, multiple assembly sub-task sequences of the assembly task are determined; secondly, based on a pre-trained action library, multiple actionable action sequences for each assembly sub-task in each assembly sub-task sequence are respectively determined; subsequently, similarity scores and optimal feasibility scores corresponding to each action in each actionable action sequence are respectively determined; based on the similarity scores and optimal feasibility scores corresponding to all actions in each actionable action sequence, the total score of the corresponding actionable action sequence is determined; the actionable action sequence with the maximum total score is determined as the optimal action sequence of the assembly robot; thirdly, based on the pre-trained action library and the optimal action sequence, an optimal action parameter sequence is determined; finally, based on the pre-trained action library, the optimal action parameter sequence is converted into the movements of the joints of the assembly robot to perform the assembly task. This application decomposes the assembly task based on a large language model combined with the RAG technology to obtain multiple assembly sub-task sequences, improving the task understanding and decomposition ability of the assembly robot task planning; determining the optimal action sequence of the assembly robot based on the similarity score and the optimal feasibility score, improving the feasibility of the assembly robot; the method of this application can be used to implement the assembly robot task planning, improving the adaptability of the assembly robot task planning. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0016] Figure 1 It is a schematic flowchart of the assembly robot task planning method provided by an embodiment of this application.

[0017] Figure 2 It is a schematic diagram of the assembly robot task planning architecture provided by an embodiment of this application.

[0018] Figure 3 It is a schematic diagram of the neural network structure corresponding to the value function.

[0019] Figure 4Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0020] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0021] The purpose of the present application is to provide a method and device for task planning of an assembly robot arm, aiming to improve the task understanding and decomposition ability, feasibility and adaptability of the task planning of the assembly robot arm.

[0022] To make the above objects, features and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0023] In an exemplary embodiment, as Figure 1 and Figure 2 shown, a method for task planning of an assembly robot arm is provided, including Step 01 - Step 10.

[0024] Step 01: Obtain multi-modal data of the assembly robot arm; the multi-modal data includes: images of the assembly site and text descriptions of the assembly tasks.

[0025] As an alternative implementation manner, before Step 01, it further includes: defining a pre-trained action library and an assembly knowledge base.

[0026] As an alternative implementation manner, the pre-trained action library includes: action parameters of multiple actions, strategies for selecting optimal action parameters for each action in any environment, and value functions.

[0027] The assembly knowledge base includes: a component ontology library, an assembly relationship graph, and an assembly process knowledge tree.

[0028] The component ontology library is used to define the geometric, physical, and functional attributes of parts.

[0029] The assembly relationship graph is used to describe the assembly relationships between various parts.

[0030] The assembly process knowledge tree is used to represent the sequence of assembly processes.

[0031] Specifically, when defining the pre-trained action library, first determine the actions required in the assembly task , such as: grasping, aligning, inserting, placing, moving, etc. And each action needs to define the following three attributes: ① Action parameters : The action parameters required for each action execution, which are used for the execution of each action's underlying planner. For example, the parameters required for grasping are the coordinates and poses of the object to be grasped; the parameters required for insertion are the insertion direction and force, etc.

[0032] ② Strategy : The strategy for each action to select the optimal action parameters in the current environment, and the form can be any method of reinforcement learning, learning-based methods such as neural networks, etc.

[0033] ③ A value function : The value function for whether the final execution of any action parameters selected by each action in any environment is successful. If the execution is successful, then ; if the execution fails, then .

[0034] After defining the three attributes, it is also necessary to conduct detailed training or solution for each part until an attribute that conforms to the actual situation is obtained. First, for the action parameters , it is also necessary to define a underlying motion planner for executing the action parameters . This motion planner can use robotic arm path planning algorithms such as reinforcement learning, Rapidly-exploring Random Tree (RRT), and Probabilistic Roadmap (PRM).

[0035] Secondly, it is necessary to train the value function of each action . The training effect of the value function is crucial for the generation of the strategy. Since the value of can only be 0 or 1 in the definition, can be regarded as a binary classifier. The input is the environment of the assembly site corresponding to each moment, including the assembly image (the image of the assembly site) and the part coordinates. A neural network as shown in Figure 3 is designed to represent . The structure is as follows. Image processing includes a convolutional layer, a pooling layer, a convolutional layer, a pooling layer, and three convolutional layers connected in sequence, which are used to convert the assembly image into a high-dimensional vector; the hidden layer includes two hidden layers with 256 neurons each, and finally the input is binary-classified through the softmax layer to obtain the value of the value function.

[0036] For the training data of the value function, it is generated through the following steps: ① Define the value range of the action parameters.

[0037] ② Sample the data of each scenario, including images and part coordinates (collected in the simulation environment).

[0038] ③For the action parameters, uniform sampling is performed within a given range.

[0039] ④Simulate the sampled environment and action parameters, and judge whether the action is successful to obtain the value function.

[0040] Make a slight modification to this sampling strategy to maintain at least a 40% success-failure ratio, because for more challenging actions such as pulling and pushing, uniform sampling rarely produces successful results. Approximately 1 million triples were collected for each action, with 800,000 used for training and the remaining 200,000 used for validation.

[0041] Finally, train the policy for each action , where the policy here refers to the selection of action parameters, rather than the policy of choosing a path given a goal in general path planning, but the policy of choosing the planning goal, that is, the action parameters. At each assembly stage, there will be a corresponding specific assembly environment, and a policy for selecting the optimal action parameters according to the assembly environment is trained. Since the value function has been trained, directly according to the learned value function model, the value function is transformed into the optimal policy through the maximum entropy model, that is: .

[0042] Among them, is the optimal policy; is the expectation of the training set; is the training set used in the policy training process, including sample pairs of the environment and action parameters ; is the entropy regularization coefficient used to balance the two optimization objectives of the policy, that is, maximizing the value of the value function and maximizing the entropy of the policy. The robustness and diversity of the policy can be improved through entropy regularization.

[0043] When defining the assembly knowledge base, in order to assist the large model in understanding the specified assembly environment, it is necessary to first define the knowledge of parts, processes, environments, etc. required in the assembly site. The assembly knowledge is mainly divided into the following categories: ①Component ontology library: used to define the geometric attributes (CAD models, dimensions, etc.), physical attributes (materials, masses, friction coefficients, etc.), functional attributes (component types, fit tolerance grades) of parts, etc.

[0044] ②Assembly relationship graph: used to describe the assembly relationships between various parts, which can be represented by a knowledge graph. Each edge can be used to represent the fit type, fit tolerance grade, assembly constraint relationships (sequence constraint, geometric constraint, mechanical constraint) between parts, etc.

[0045] ③Assembly process knowledge tree: It is used to represent the sequence of processes during assembly. Using a tree structure can better represent the assembly sequence among multiple parts. At the same time, tools, assembly sequence, assembly methods, etc. required for each process can also be defined through the knowledge tree.

[0046] After the definition is completed, it exists in graph databases such as neo4j in the form of a knowledge graph, which is convenient for fine-tuning the retrieval ability of the large model in specific tasks.

[0047] Finally, a pre-trained action library containing action parameters, strategies, a value function, and a defined assembly knowledge base are defined.

[0048] Step 02: Determine the joint feature encoding based on the image and text description.

[0049] As an alternative implementation, Step 02 includes Step 021 - Step 024.

[0050] Step 021: Preprocess the image to obtain an initial image vector; the preprocessing includes: image block division, block embedding, and position encoding.

[0051] Specifically, assume the input image is an RGB image with a size of 3×224×224. During preprocessing, first, divide the image into non-overlapping image blocks with a size of 14×14. A total of 256 image blocks are obtained. Then, flatten each image block with a size of 14×14 into a 196-dimensional vector, and finally obtain a 768×196 vector. After that, add a learnable label with an initial value of 0 before each image block to obtain a 768×197 vector. Finally, to express the position information of each image block, perform position encoding on the positions of the image blocks and add it to the previously added learnable label. After preprocessing, a 768×197 initial image vector is finally obtained.

[0052] Step 022: Extract features from the initial image vector to obtain a visual feature vector.

[0053] Specifically, use a ViT model composed of 12 stacked transformer encoders to extract features from the initial image vector. Among them, each transformer encoder contains a multi-head attention mechanism (12 heads of 64 dimensions) to capture global dependencies. Then, pass through a fully connected neural network (768→1024→768). A normalization layer is added before both of them to stabilize the training process. Finally, use the first row marker vector (with a dimension of 768×1) of the final output as the final output visual feature vector.

[0054] Step 023: Encode the text description through the BERT model to obtain a text feature vector.

[0055] Specifically, the text description is segmented, and special tags ([CLS] and [SEP]) are added at the beginning and end, and a maximum text length is specified. If the text description is too long, it will be trimmed, and if it is too short, it will be padded. Finally, the words in the text description are converted to numbers in the dictionary to obtain a vector representing the text description with a dimension of the maximum text length. This is the end of the preprocessing part. After that, the vector is input into the pre-trained BERT model. Finally, the 768-dimensional vector obtained by the BERT model for [CLS] is used as the text feature vector.

[0056] Step 024: Perform weighted integration on the visual feature vector and the text feature vector to obtain a joint feature code.

[0057] Specifically, first, the correlation between each visual feature in the visual feature vector and each text feature in the text feature vector is calculated through the dot product function, and then a softmax layer is used to obtain the attention weight to indicate the degree of correlation between the two, namely: .

[0058] in, for and The degree of correlation between for and The dot product between ; The first visual features; is the first text features.

[0059] Then, using the obtained weights , weighted integration of each text feature in the text feature vector is performed to obtain the joint feature encoding. The formula is: .

[0060] in, is the first A joint feature.

[0061] Step 03: Determine assembly retrieval results based on the joint feature coding and assembly knowledge base.

[0062] As an optional implementation, step 03 includes steps 031 to 034.

[0063] Step 031: From the part ontology database, screen out all parts whose cosine similarity with the combined feature code is greater than the first preset threshold to obtain multiple relevant parts.

[0064] Specifically, first, vectorize the part ontology database. Use BERT to vectorize the text information in the part ontology database simultaneously, and vectorize the part ontology library to obtain the vectors of each part.

[0065] Calculate the cosine similarity between each part and the combined feature code based on the vectors of each part and the combined feature code respectively.

[0066] The first preset threshold is set to 60%.

[0067] Step 032: Based on the assembly relationship graph, determine the assembly relationships between the relevant parts.

[0068] Step 033: Based on the assembly process knowledge tree and the assembly relationships between the relevant parts, determine the assembly process sequence.

[0069] Specifically, match the assembly relationships between the relevant parts in the assembly process knowledge tree. By comparing the pre-assembly and post-assembly sequences in the assembly process knowledge tree, if assembly relationship 1 and assembly relationship 2 are found under the same assembly process knowledge tree (a certain assembly process), if assembly relationship 1 is the parent node of assembly relationship 2, then assembly relationship 1 must precede assembly relationship 2, and assembly relationship 1 and assembly relationship 2 are strongly correlated. If assembly relationship 1 and assembly relationship 2 are sibling nodes of a tree, then the pre-assembly and post-assembly sequences of assembly relationship 1 and assembly relationship 2 are not strictly distinguished, but they are strongly correlated. If assembly relationship 1 and assembly relationship 2 do not belong to the same tree, then assembly relationship 1 and assembly relationship 2 are weakly correlated and the pre-assembly and post-assembly sequences do not need to be distinguished.

[0070] Step 034: Determine the relevant parts, the assembly relationships between the parts in each relevant part, and the assembly process sequence as the assembly retrieval result.

[0071] For example, the relevant parts obtained by screening through cosine similarity are the first-stage gear, gear shaft, second-stage gear, bearing, etc. Then, in the assembly relationship graph, the assembly relationship between the first-stage gear and the gear shaft is coaxial and interference fit, etc. Finally, sort out the pre-assembly and post-assembly sequences of each assembly relationship, such as the assembly of the bearing and the base should be after the gear and the gear shaft, to obtain the assembly retrieval result, and structure it into the following structure: Relevant parts: ["First-stage gear" (radius: 30mm, number of teeth: 15...), "Gear shaft" (radius: 5mm...), "Bearing" (inner diameter: 5mm...),...].

[0072] Assembly relationship: [[“First - stage gear”, “Gear shaft”, “Coaxial, H7 / h6 interference fit...”],...].

[0073] Assembly process sequence: [[“Assembly of the first - stage gear and the gear shaft” (first), “Assembly of the bearing and the base” (later)],...].

[0074] Step 04: Based on multi - modal data, assembly retrieval results, and large - language models, determine multiple assembly subtask sequences for the assembly task.

[0075] Specifically, combine multi - modal data, assembly retrieval results with prompt words and input them into the large - model. The example format during input is as follows: Assembly task description: Assembly of the two - stage reducer.

[0076] Multi - modal data.

[0077] Related parts: [“First - stage gear” (radius: 30mm, number of teeth: 15...), “Gear shaft” (radius: 5mm...), “Bearing” (inner diameter: 5mm...),...].

[0078] Assembly relationship: [[“First - stage gear”, “Gear shaft”, “Coaxial, H7 / h6 interference fit...”],...].

[0079] Assembly process sequence: [[“Assembly of the first - stage gear and the gear shaft” (first), “Assembly of the bearing and the base” (later)],...].

[0080] By using the Retrieval - Augmented Generation (RAG) technology to enhance the domain - knowledge retrieval ability of the large - language model (that is, before inputting into the large - language model, first use cosine similarity to screen relevant parts and obtain the assembly retrieval results step by step, and then input the assembly retrieval results and multi - modal data into the large - language model), it solves the problem of insufficient knowledge of general large - language models in professional assembly scenarios and improves the accuracy of task decomposition and process rationality.

[0081] Step 05: Based on the pre - trained action library, determine multiple executable action sequences for each assembly subtask in each assembly subtask sequence.

[0082] As an optional implementation, Step 05 includes Step 051 - Step 053.

[0083] Step 051: Determine any assembly subtask as the subtask to be determined.

[0084] Step 052: Determine the cosine similarity between the subtask to be determined and each action in the pre - trained action library respectively.

[0085] Step 053: Input all the actions and assembly retrieval results corresponding to the subtasks to be determined with cosine similarity greater than the second preset threshold into the large language model to obtain multiple actionable action sequences for the subtasks to be determined.

[0086] Step 06: Determine the similarity scores and optimal feasibility scores corresponding to each action in each actionable action sequence respectively.

[0087] As an alternative implementation, Step 06 includes Step 061 - Step 065.

[0088] Step 061: Determine any assembly subtask as the current subtask, determine any actionable action sequence of the current assembly subtask as the current action sequence, and determine any action in the current action sequence as the current action.

[0089] Step 062: Calculate the similarity score between the current action and the current subtask, and determine the similarity score between the current action and the current subtask as the similarity score corresponding to the current action.

[0090] Specifically, for the th action in the current actionable action sequence of any current subtask, the calculation formula for the similarity score between the action and the current subtask is: .

[0091] Where is the similarity score between the th action and the current subtask; is the probability of generating and under the conditions of ; is the 1st to th environment; is the 1st to th action; is the vector of the th action; is the vector between the current subtasks; is the modulus of the vector; is the dot product of two vectors.

[0092] Step 063: Based on the policies in the pre-trained action library, determine multiple action parameter sequences for the current action sequence.

[0093] Specifically, for the current action sequence, sample from the probability distribution of the action parameters in the current environment at the assembly site output by the policy to obtain multiple first action parameters in the current environment, and respectively based on each first action parameter according to (i.e., the environment obtains the action parameter after assuming that the action is executed and updated to the environment , and then sampling a new action parameter according to the policy again .) Update sequentially to obtain multiple action parameter sequences.

[0094] Step 064: Based on the value function in the pre-trained action library, determine the feasibility scores of the current action under each action parameter sequence respectively.

[0095] Specifically, the calculation formula for the feasibility score of the th action under any current action parameter is: .

[0096] Wherein, is the feasibility score of the th action under any current action parameter; is the probability of and under the condition of , where is the action parameter; is the value of the value function under and ; is the th environment; is the th action.

[0097] Step 065: Based on the feasibility scores of the current action under each action parameter sequence, determine the optimal feasibility score corresponding to the current action.

[0098] Specifically, determine the maximum feasibility score among the feasibility scores of the current action under each action parameter sequence as the optimal feasibility score.

[0099] Step 07: Based on the similarity scores and the optimal feasibility scores corresponding to all actions in each actionable action sequence respectively, determine the total score of the corresponding actionable action sequence.

[0100] As an optional implementation manner, step 07 includes step 071-step 072.

[0101] Step 071: Multiply the similarity score and the optimal feasibility score corresponding to each action in each actionable action sequence to obtain the comprehensive score of the corresponding action.

[0102] Specifically, the calculation formula for the comprehensive score of any action is: .

[0103] Where, is the comprehensive score of the th action .

[0104] Step 072: Multiply the comprehensive scores of all actions in each actionable action sequence respectively to obtain the total score of the corresponding actionable action sequence.

[0105] Step 08: Determine the optimal action sequence of the assembly robot arm as the actionable action sequence with the largest total score.

[0106] Step 09: Determine the optimal action parameter sequence based on the pre-trained action library and the optimal action sequence.

[0107] As an alternative implementation, Step 09 includes Step 091 - Step 092.

[0108] Step 091: Determine the optimal action parameters of the first action in the optimal action sequence based on the current environment at the assembly site, the first action in the optimal action sequence, and the strategy.

[0109] Specifically, based on the current environment at the assembly site, the first action in the optimal action sequence, and the strategy, determine multiple initial action parameters of the first action in the optimal action sequence, and determine the optimal action parameters of the first action in the optimal action sequence as the initial action parameter with the highest probability.

[0110] Step 092: Based on the optimal action parameters of the first action in the optimal action sequence, sequentially determine the optimal action parameters of each action in the optimal action sequence, thereby determining the optimal action parameter sequence.

[0111] Specifically, according to , based on the optimal action parameters of the first action in the optimal action sequence, sequentially determine the optimal action parameters of each action in the optimal action sequence, thereby determining the optimal action parameter sequence.

[0112] Step 10: Based on the pre-trained action library, convert the optimal action parameter sequence into the movements of each joint of the assembly robot arm to perform the assembly task.

[0113] Specifically, substitute the optimal action parameter sequence into the pre-trained action library in step 1 to execute the underlying planner, convert the optimal action parameter sequence into the movement of each joint of the robotic arm, and finally complete the implementation from the complex assembly site to the assembly task.

[0114] Furthermore, the underlying planner provides real-time feedback of environmental information (such as part coordinates, images of the assembly site) to support dynamic adjustment of action parameters and adapt to the uncertainties and disturbances during the assembly process.

[0115] In an exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the task planning method for an assembly robotic arm.

[0116] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the task planning method for an assembly robotic arm is implemented.

[0117] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 4 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, a task planning method for an assembly robotic arm is implemented.

[0118] Those skilled in the art can understand that Figure 4 the structure shown in

[0119] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0120] In the various embodiments provided in this application, the databases involved can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the various embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logics, data processing logics based on quantum computing, etc., and are not limited thereto.

[0121] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0122] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0123] In this article, specific examples are used to illustrate the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A task planning method for an assembly robotic arm, characterized in that The described assembly robotic arm task planning method includes: Obtain multi-modal data of the assembly robotic arm; the multi-modal data includes: images of the assembly site and text descriptions of the assembly tasks; Based on the images and the text descriptions, determine the joint feature encoding; Based on the joint feature encoding and the assembly knowledge base, determine the assembly retrieval result; Based on the multi-modal data, the assembly retrieval result, and the large language model, determine multiple assembly sub-task sequences for the assembly task; Based on the pre-trained action library, respectively determine multiple executable action sequences for each assembly sub-task in each assembly sub-task sequence; Respectively determine the similarity scores and optimal feasibility scores corresponding to each action in each executable action sequence; Based on the similarity scores and optimal feasibility scores corresponding to all actions in each executable action sequence respectively, determine the total score of the corresponding executable action sequence; Determine the executable action sequence with the maximum total score as the optimal action sequence of the assembly robotic arm; Based on the pre-trained action library and the optimal action sequence, determine the optimal action parameter sequence; Based on the pre-trained action library, transform the optimal action parameter sequence into the motion of each joint of the assembly robotic arm to perform the assembly task.

2. The assembly robotic arm task planning method according to claim 1, wherein Before obtaining the multi-modal data of the assembly robotic arm, it further includes: defining a pre-trained action library and an assembly knowledge base.

3. The method for task planning of an assembly robotic arm according to claim 2, wherein The pre-trained action library includes: action parameters of multiple actions, the strategy and value function for each action to select the optimal action parameters in any environment; The assembly knowledge base includes: a component ontology library, an assembly relationship graph, and an assembly process knowledge tree; The component ontology library is used to define the geometric attributes, physical attributes, and functional attributes of parts; The assembly relationship graph is used to describe the assembly relationships between various parts; The assembly process knowledge tree is used to represent the sequence of assembly processes.

4. The method for task planning of an assembly robotic arm according to claim 3, wherein Based on the images and the text descriptions, determining the joint feature encoding includes: Preprocess the images to obtain an initial image vector; the preprocessing includes: image block division, block embedding, and position encoding; Extract features from the initial image vector to obtain a visual feature vector; Encode the text description through a BERT model to obtain a text feature vector; Perform weighted integration on the visual feature vector and the text feature vector to obtain the joint feature encoding.

5. The task planning method of the assembly robotic arm according to claim 4, wherein Based on the joint feature encoding and the assembly knowledge base, determining the assembly retrieval result includes: From the part ontology database, screen out all parts whose cosine similarity with the joint feature encoding is greater than a first preset threshold to obtain multiple relevant parts; Based on the assembly relationship graph, determine the assembly relationships between the relevant parts; Based on the assembly process knowledge tree and the assembly relationships between the relevant parts, determine the assembly process sequence; Determine each relevant part, the assembly relationships between the parts in each relevant part, and the assembly process sequence as the assembly retrieval result.

6. The assembly robot task planning method according to claim 5, characterized in that Based on the pre-trained action library, respectively determining multiple executable action sequences for each assembly sub-task in each assembly sub-task sequence includes: Determine any assembly sub-task as the sub-task to be determined; Respectively determine the cosine similarity between the sub-task to be determined and each action in the pre-trained action library; All actions with a cosine similarity greater than the second preset threshold corresponding to the subtasks to be determined and the assembly retrieval results are input into the large language model to obtain multiple actionable action sequences for the subtasks to be determined.

7. The task planning method for an assembly robotic arm according to claim 6, characterized in that Respectively determine the similarity scores and optimal feasibility scores corresponding to each action in each actionable action sequence, including: Determine any assembly subtask as the current subtask, determine any actionable action sequence of the current assembly subtask as the current action sequence, and determine any action in the current action sequence as the current action; Calculate the similarity score between the current action and the current subtask, and determine the similarity score corresponding to the current action as the similarity score between the current action and the current subtask; Based on the policy in the pre-trained action library, determine multiple action parameter sequences for the current action sequence; Based on the value function in the pre-trained action library, respectively determine the feasibility scores of the current action under each action parameter sequence; Based on the feasibility scores of the current action under each action parameter sequence, determine the optimal feasibility score corresponding to the current action.

8. The assembly robotic arm task planning method according to claim 7, characterized in that Respectively determine the total scores of the corresponding actionable action sequences based on the similarity scores and optimal feasibility scores corresponding to all actions in each actionable action sequence, including: Multiply the similarity scores and optimal feasibility scores corresponding to each action in each actionable action sequence respectively to obtain the comprehensive score of the corresponding action; Multiply the comprehensive scores of all actions in each actionable action sequence respectively to obtain the total score of the corresponding actionable action sequence.

9. The method for task planning of an assembly robotic arm according to claim 8, wherein Based on the pre-trained action library and the optimal action sequence, determine the optimal action parameter sequence, including: Based on the current environment at the assembly site, the first action in the optimal action sequence, and the policy, determine the optimal action parameters of the first action in the optimal action sequence; Based on the optimal action parameters of the first action in the optimal action sequence, sequentially determine the optimal action parameters of each action in the optimal action sequence, so as to determine the optimal action parameter sequence.

10. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the assembly robotic arm task planning method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Robot assembly unit multi-objective layout optimization method based on multi-group distribution estimation algorithm

    CN109917754A

  • Dual-module cooperative robot coordinated assembly system for 3C assembly and planning method

    CN111522305A

  • Man-machine interaction assembly method and system based on multi-modal large model and reinforcement learning

    CN118744426A

  • System for controlling robot task decision-making on the basis of semantic network and knowledge base

    WO2025102453A1

Cited By

  • Robot assembly planning method and system based on multi-modal large model and robot

    CN122471336A