Natural language driven robot arm control method, apparatus, device, medium and product

By using large language models and multimodal perception technology, the robotic arm can accurately understand and efficiently execute natural language commands, solving the problems of insufficient flexibility and intelligence in existing control schemes. It is applicable to fields such as intelligent manufacturing, service robots, and medical assistance.

CN120886271BActive Publication Date: 2025-12-12BEIJING BAIXINGHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511404825.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-12-12
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

Existing robotic arm control solutions are insufficient in terms of flexibility, ease of use, and intelligence, making it difficult to adapt to dynamic task requirements and complex environments.

Method used

By employing a large language model combined with multimodal perception and adaptive control technologies, the robotic arm achieves precise understanding and efficient execution through hierarchical semantic parsing of natural language commands, multimodal perception data fusion, and task graph generation.

Benefits of technology

It improves the accuracy, adaptability, flexibility and ease of use of robotic arms in dynamic environments, and is suitable for fields such as intelligent manufacturing, service robots and medical assistance. It features high precision, strong adaptability, intelligence and high efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120886271B_ABST
    Figure CN120886271B_ABST
Patent Text Reader

Abstract

The application discloses a natural language driven mechanical arm control method, device, equipment, medium and product, and relates to the technical field of artificial intelligence and robots. The method comprises the following steps: after natural language instructions and multi-modal perception data are acquired, the natural language instructions are first decomposed into semantic embedding vectors, a large language model is used to obtain structured task representations through hierarchical semantic parsing, multi-modal perception data are fused to generate environment representations required for task execution, the structured task representations and the environment representations are then converted into a task graph, the task graph is converted into joint control instructions of a target mechanical arm, and finally the joint control instructions are sent to a joint controller for execution. Through multi-level semantic understanding, multi-modal perception fusion and adaptive execution control, the mechanical arm can accurately understand and efficiently execute complex natural language instructions, and the task execution accuracy, adaptability, flexibility, ease of use and intelligent level of the mechanical arm in a dynamic environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of interdisciplinary technology of artificial intelligence and robotics, specifically relating to a natural language-driven robotic arm control method, device, equipment, medium, and product. Background Technology

[0002] As the core execution unit in the fields of industrial automation and service robots, the development of control technology for robotic arms is crucial for improving production efficiency and human-machine collaboration.

[0003] Currently, existing robotic arm control schemes mainly include programmed control schemes, teach-in control schemes, vision-guided control schemes, and policy-based control schemes. However, these schemes have the following limitations in practical applications: (1) Programmed control schemes: rely on professionals to write fixed programs, which are difficult to adapt to dynamic task requirements and are not user-friendly for non-technical users; (2) Teach-in control schemes: are cumbersome to operate, inefficient, and unable to cope with complex or variable tasks; (3) Vision-guided control schemes: are sensitive to environmental changes, difficult to handle situations such as occlusion and changes in lighting, and are usually limited to simple tasks; (4) Policy-based control schemes: lack versatility and are difficult to extend to new tasks or new environments.

[0004] With the advancement of artificial intelligence technology, especially the breakthroughs of large language models in the field of natural language processing, how to combine large language models, multimodal perception, and adaptive control technologies to enable robotic arms to intelligently understand and execute natural language instructions, thereby significantly improving their flexibility, ease of use, and intelligence level, is a topic that urgently needs to be studied by those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a natural language-driven robotic arm control method, device, control equipment, computer-readable storage medium, and computer program product to solve the problems of limitations in flexibility, ease of use, and intelligence level of existing robotic arm control solutions.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] In a first aspect, a natural language-driven robotic arm control method is provided, executed by a control device, wherein the control device is communicatively connected to a language command input device, a multimodal sensing device, and a joint controller of the target robotic arm, and the multimodal sensing device is located in the surrounding environment of the target robotic arm;

[0008] The natural language-driven robotic arm control method includes:

[0009] Acquire natural language commands input by the language command input device and multimodal perception data collected by the multimodal perception device, wherein the multimodal perception data includes visual data, tactile data and environmental data;

[0010] The natural language instruction is decomposed into semantic embedding vectors, and a large language model based on the Transformer structure is used to perform hierarchical semantic parsing on the semantic embedding vectors to obtain a structured task representation. The hierarchical semantic parsing includes sequential semantic parsing of the intent recognition layer, semantic parsing of the parameter extraction layer, and semantic parsing of the task combination layer. The structured task representation includes the task intent obtained by the semantic parsing of the intent recognition layer, the task parameters obtained by the semantic parsing of the parameter extraction layer, at least one identified subtask obtained by the semantic parsing of the task combination layer, and the execution order of the at least one identified subtask.

[0011] The multimodal sensing data is fused to generate the environmental representation required for task execution;

[0012] The structured task representation and the environment representation are transformed into a task graph, wherein the nodes in the task graph represent the identified subtasks, and the edges in the task graph represent the execution order;

[0013] The task diagram is converted into joint control commands for the target robotic arm;

[0014] The joint control command is sent to the joint controller for execution.

[0015] Based on the above-mentioned invention, a novel scheme for natural language-driven robotic arm control is provided, combining large language models, multimodal perception, and adaptive control technologies. Specifically, after acquiring natural language instructions and multimodal perception data, the natural language instructions are first decomposed into semantic embedding vectors, and a structured task representation is obtained through hierarchical semantic parsing using a large language model. Then, the multimodal perception data is fused to generate the environmental representation required for task execution. Next, the structured task representation and environmental representation are transformed into a task graph, and the task graph is converted into joint control instructions for the target robotic arm. Finally, the joint control instructions are sent to the joint controller for execution. Through multi-level semantic understanding, multimodal perception fusion, and adaptive execution control, the robotic arm can achieve accurate understanding and efficient execution of complex natural language instructions, improving its task execution accuracy, adaptability, flexibility, ease of use, and intelligence level in dynamic environments. This is particularly suitable for applications requiring high-precision human-machine interaction, such as intelligent manufacturing, service robots, and medical assistance, facilitating practical application and promotion.

[0016] In one possible design, during the hierarchical semantic parsing process, the loss function of semantic parsing... It is expressed as follows:

[0017]

[0018] In the formula, This represents the total number of natural language instruction samples. Indicates less than or equal to positive integers, Indicates the relationship with the first The actual semantic labels corresponding to each natural language instruction sample. Indicates the first A sample of natural language instructions. This represents the model parameters of the large language model. This represents the probability distribution function predicted by the large language model.

[0019] In one possible design, the multimodal sensing data is fused, including:

[0020] A multimodal dynamic fusion network based on a self-attention mechanism is employed to optimize the modal contribution of each modal sensing data in the multimodal sensing data in real time according to task requirements and environmental changes. Then, the multimodal sensing data is fused based on all the modal contributions, wherein the modal contributions are expressed as fusion weights by the following formula:

[0021]

[0022] In the formula, This represents the total number of modalities in the multimodal sensing data. and They represent less than or equal to positive integers, In the multimodal sensing data, the first... The fusion weights of the modal sensing data This represents the natural exponential function. This indicates the preset adjustment parameters. Indicates the first The correlation between each modal sensing data and the target task and has , Indicates the first Feature vectors of modal sensing data Indicates the transpose symbol. The feature vector representing the target task, also known as the structured task representation. Indicates the first Confidence level of modal sensing data This represents the Sigmoid function. Indicates the first The sensitivity of each modal sensing data to environmental changes and has , This represents the preset sensitivity coefficient to environmental changes. This represents the environmental feature vector at the current moment. This represents the environmental feature vector from the previous time step. In the multimodal sensing data, the first... The correlation between each modal sensing data and the target task. Indicates the first The sensitivity of each modal sensing data to environmental changes.

[0023] In one possible design, the structured task representation and the environment representation are transformed into a task graph, including:

[0024] A task graph generation model based on graph neural networks is used to transform the structured task representation and the environment representation into a task graph. In this task graph, nodes represent identified subtasks, and edges represent the execution order. The task graph generation model aims to minimize the optimization objective. It is expressed as follows:

[0025]

[0026] In the formula, This represents the total number of natural language instruction samples. Indicates less than or equal to positive integers, Indicates the relationship with the first Structured task representations corresponding to each natural language instruction sample. Indicates the relationship with the first The environment representation corresponding to the nth multimodal sensing data sample, the nth The first natural language instruction sample and the first It was obtained by acquiring multiple multimodal sensing data samples together. Indicates the first The first natural language instruction sample and the first The real task graph labels corresponding to each multimodal sensing data sample. This represents the probability distribution function predicted by the task graph generation model. This represents the preset complexity penalty coefficient. Indicates less than or equal to positive integers, This represents the number of nodes in the generated task graph. This represents the number of edges in the generated task graph, where the generated task graph refers to the graph generated by the task graph generation model that connects the first edge to the second edge. The structured task representation corresponding to the first natural language instruction sample and its relation to the first... The task graph is obtained by transforming the environmental representation corresponding to each multimodal sensing data sample.

[0027] In one possible design, the task diagram is translated into joint control commands for the target robotic arm, including:

[0028] In the process of converting the task diagram into joint control commands for the target robotic arm, an inverse kinematics algorithm capable of adaptive adjustment with multiple constraints is used to solve the following optimization problem to obtain the joint angle vector used to generate the joint control commands:

[0029]

[0030] In the formula, Indicates minimization. This represents the joint angle vector. Represents the target space coordinate vector. Represents the forward kinematic function. Represents positive integers. Represents the first in the multiple constraints The constraints are in functional form, and the multiple constraints include joint limit constraints, load constraints, and / or obstacle constraints. This represents the preset constraint weight coefficient.

[0031] In one possible design, after sending the joint control command to the joint controller for execution, the method further includes:

[0032] Collect the execution status monitoring results of the target robotic arm;

[0033] A reinforcement learning-based controller is used to optimize the execution of the target robotic arm's control tasks based on the execution state monitoring results, wherein the reward function of the reinforcement learning... It is expressed as follows:

[0034]

[0035] In the formula, 、 and These represent the preset weighting coefficients. Denotes the base of the natural logarithm. This indicates the control accuracy of the target robotic arm. This indicates the speed at which the target robotic arm executes its control tasks. This indicates the number of errors in the execution of the control task of the target robotic arm.

[0036] In a second aspect, a natural language driven robotic arm control device is provided, which is suitable for being arranged in a control device, wherein the control device is communicatively connected to a language command input device, a multimodal sensing device and a joint controller of the target robotic arm, and the multimodal sensing device is located in the surrounding environment of the target robotic arm;

[0037] The natural language driven robotic arm control device includes an instruction data acquisition unit, a semantic parsing and processing unit, a perception data fusion unit, a task graph conversion unit, a control instruction conversion unit, and a control instruction sending unit.

[0038] The instruction data acquisition unit is used to acquire natural language instructions input by the language instruction input device and multimodal perception data collected by the multimodal perception device, wherein the multimodal perception data includes visual data, tactile data and environmental data.

[0039] The semantic parsing processing unit is communicatively connected to the instruction data acquisition unit. It is used to decompose the natural language instruction into semantic embedding vectors and perform hierarchical semantic parsing processing on the semantic embedding vectors using a large language model based on the Transformer structure to obtain a structured task representation. The hierarchical semantic parsing processing includes sequentially performed intent recognition layer semantic parsing processing, parameter extraction layer semantic parsing processing, and task combination layer semantic parsing processing. The structured task representation includes the task intent obtained through the intent recognition layer semantic parsing processing, the task parameters obtained through the parameter extraction layer semantic parsing processing, and at least one identified subtask obtained through the task combination layer semantic parsing processing, as well as the execution order of the at least one identified subtask.

[0040] The perception data fusion unit is communicatively connected to the instruction data acquisition unit and is used to fuse the multimodal perception data to generate an environmental representation required for task execution.

[0041] The task graph transformation unit is communicatively connected to the semantic parsing processing unit and the perceptual data fusion unit, respectively, and is used to transform the structured task representation and the environment representation into a task graph, wherein the nodes in the task graph represent the identified sub-tasks, and the edges in the task graph represent the execution order;

[0042] The control command conversion unit is communicatively connected to the task diagram conversion unit and is used to convert the task diagram into joint control commands for the target robotic arm.

[0043] The control command sending unit is communicatively connected to the control command conversion unit and is used to send the joint control command to the joint controller for execution.

[0044] Thirdly, the present invention provides a control device, comprising a storage module, a processing module, and a transceiver module connected in sequence for communication, wherein the storage module is used to store a computer program, the transceiver module is used to send and receive messages, and the processing module is used to read the computer program and execute the natural language driven robotic arm control method as described in the first aspect or any possible design in the first aspect.

[0045] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, perform the natural language-driven robotic arm control method as described in the first aspect or any possible design within the first aspect.

[0046] Fifthly, the present invention provides a computer program product, including a computer program or instructions, which, when executed by a computer, implement the natural language driven robotic arm control method as described in the first aspect or any possible design in the first aspect.

[0047] The beneficial effects of the above scheme are:

[0048] (1) This invention creatively provides a new scheme for natural language driven robotic arm control by combining large language model, multimodal perception and adaptive control technology. That is, after acquiring natural language instructions and multimodal perception data, the natural language instructions are first decomposed into semantic embedding vectors, and a structured task representation is obtained by hierarchical semantic parsing using a large language model. Then, the multimodal perception data is fused to generate the environment representation required for task execution. Then, the structured task representation and environment representation are transformed into a task graph, and the task graph is transformed into joint control instructions for the target robotic arm. Finally, the joint control instructions are sent to the joint controller for execution. In this way, through multi-level semantic understanding, multimodal perception fusion and adaptive execution control, the robotic arm can achieve accurate understanding and efficient execution of complex natural language instructions, and improve its task execution accuracy, adaptability, flexibility, ease of use and intelligence level in dynamic environments.

[0049] (2) It has high precision characteristics: it can make the control error of the end effector of the robotic arm less than 0.1mm, which is a significant improvement compared to the 1mm error of the traditional method, thus meeting the requirements of high precision tasks;

[0050] (3) It has strong adaptability: It integrates multi-source data such as vision and touch in real time through a multimodal perception fusion network and combines it with a dynamic task graph generation algorithm, which can adapt to dynamic environments and support complex scenarios such as changes in lighting, object occlusion and dynamic obstacles.

[0051] (4) It has intelligent features: Through hierarchical semantic parsing, the accuracy of understanding natural language instructions can exceed 95%, and it can handle multi-step, ambiguous or context-dependent instructions, thereby improving the naturalness and efficiency of human-computer interaction.

[0052] (5) It has the characteristics of high efficiency: By adjusting the task execution through real-time feedback, it can achieve the purpose of adaptive learning and improving the success rate of robotic arm control. Specifically, it can shorten the task completion time by about 20% and reduce the error rate to below 5%, thereby effectively improving robustness;

[0053] (6) It has scalability: The solution of the present invention can adopt a modular design and support different types of robotic arms and sensor configurations, which can be easily extended to new tasks and scenarios;

[0054] (7) It can be applied to a variety of scenarios that require natural language driving and dynamic task execution, and is widely used in intelligent manufacturing, service robots and medical assistance. For example, in intelligent manufacturing, the robotic arm can accurately grasp parts and complete assembly according to natural language instructions, effectively improving the flexibility and automation level of the production line. The robotic arm can also perform quality inspection tasks of product surface defects, and achieve high-precision and high-efficiency inspection through instruction driving. In the field of service robots, the robotic arm can respond to the user's language instructions and complete household tasks such as picking up items, organizing or cleaning, which greatly improves the convenience of home life. In public service places such as hotels or hospitals, the robotic arm can undertake the delivery of items or nursing assistance work, helping to reduce the burden of manpower. In medical assistance, the robotic arm can accurately locate and operate surgical instruments under the doctor's instructions, improving the accuracy and safety of minimally invasive surgery. At the same time, it can also be used to provide personalized grasping training or life support for rehabilitation patients, helping the rehabilitation process.

[0055] (8) It can improve the accuracy and adaptability of task execution while enhancing the ability to handle dynamic environments. Experiments and application examples show that the solution has certain practicality in intelligent manufacturing, service robots and medical assistance scenarios, and can provide a new implementation path for natural language interaction control in related fields, which is convenient for practical application and promotion. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1This is a flowchart illustrating the natural language-driven robotic arm control method provided in an embodiment of this application.

[0058] Figure 2 This is a schematic diagram illustrating the communication connection between the control device, the language command input device, the multimodal sensing device, and the joint controller of the target robotic arm provided in the embodiments of this application.

[0059] Figure 3 This is an example diagram illustrating the hierarchical semantic parsing process provided in an embodiment of this application.

[0060] Figure 4 A schematic diagram of the structure of the natural language driven robotic arm control device provided in the embodiments of this application.

[0061] Figure 5 A schematic diagram of the structure of the control device provided in the embodiment of this application. Detailed Implementation

[0062] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the present invention will be briefly introduced below in conjunction with the accompanying drawings and descriptions of the embodiments or the prior art. Obviously, the following description of the structure of the accompanying drawings is only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these embodiments without creative effort. It should be noted that the description of these embodiments is for the purpose of helping to understand the present invention, but does not constitute a limitation of the present invention.

[0063] It should be understood that although the terms "first" and "second", etc., may be used herein to describe various objects, these objects should not be limited by these terms. These terms are only used to distinguish one object from another. For example, the first object may be referred to as the second object, and similarly, the second object may be referred to as the first object, without departing from the scope of the exemplary embodiments of the invention.

[0064] It should be understood that the term "and / or" that may appear in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, or A and B exist simultaneously. Another example is A, B and / or C, which can mean that any one of A, B, and C or any combination thereof exists. The term " / and" that may appear in this document describes another relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone or A and B exist simultaneously. In addition, the character " / " that may appear in this document generally indicates that the related objects before and after it are in an "or" relationship.

[0065] Example

[0066] like Figures 1-3 As shown, the natural language-driven robotic arm control method provided in the first aspect of this embodiment can be executed, but is not limited to, by a control device with certain computing resources. The control device is communicatively connected to a language command input device, a multimodal sensing device, and a joint controller of the target robotic arm. The multimodal sensing device is located in the surrounding environment of the target robotic arm, such as... Figure 2 As shown. The language command input device is used to input the user's natural language commands, which can be conventionally implemented by combining a microphone with Automatic Speech Recognition (ASR) technology, but is not limited to. The multimodal sensing device is used to collect multimodal sensing data in real time, including but not limited to visual data, tactile data, and environmental data. Specifically, it includes but is not limited to visual sensors (e.g., RGB-D cameras) for collecting visual data, tactile sensors (e.g., force sensors) for collecting tactile data, and environmental sensors (e.g., inertial measurement unit (IMU) sensors) for collecting environmental data. The target robotic arm is the controlled object. The joint controller is used to generate corresponding control signals according to the joint control commands from the control device and transmit the control signals to the actuators of the target robotic arm for execution. Both the target robotic arm and the joint controller can be conventionally implemented using existing configurations.

[0067] like Figure 1 As shown, the natural language driven robotic arm control method includes, but is not limited to, the following steps S1 to S6.

[0068] S1. Acquire natural language instructions input by the language instruction input device and multimodal perception data collected by the multimodal perception device, wherein the multimodal perception data includes, but is not limited to, visual data, tactile data, and environmental data.

[0069] In step S1, since subsequent joint control commands are generated based on the natural language instructions and the multimodal sensing data, in order to ensure the rationality of the joint control commands, the input time of the natural language instructions and the acquisition time of the multimodal sensing data need to be synchronized (i.e., within the same unit period), preferably at the same time.

[0070] S2. The natural language instruction is decomposed into semantic embedding vectors, and a large language model based on the Transformer structure is used to perform hierarchical semantic parsing processing on the semantic embedding vectors to obtain a structured task representation. The hierarchical semantic parsing processing includes, but is not limited to, sequential semantic parsing processing of the intent recognition layer, semantic parsing processing of the parameter extraction layer, and semantic parsing processing of the task combination layer. The structured task representation includes, but is not limited to, the task intent obtained through the semantic parsing processing of the intent recognition layer, the task parameters obtained through the semantic parsing processing of the parameter extraction layer, and at least one identified subtask obtained through the semantic parsing processing of the task combination layer, as well as the execution order of the at least one identified subtask.

[0071] In step S2, as Figure 3 As shown, the process of decomposing the natural language instruction into semantic embedding vectors specifically includes, but is not limited to, performing conventional text segmentation on the natural language instruction, and then performing conventional semantic embedding on the segmentation results to obtain the semantic embedding vectors. The Transformer structure is an existing neural network architecture based on a self-attention mechanism. Its core structure includes an encoder and a decoder, primarily used for processing sequence data (such as natural language). The Large Language Model (LLM) is a type of deep learning model trained on massive amounts of text data, capable of generating natural language text or understanding language meaning. Therefore, the Large Language Model can, but is not limited to, employ existing BERT (Bidirectional Encoder Representations from Transformers) models or GPT (Generative Pre-trained Transformer) models. Transformer (generative pre-trained transformer) models, combined with domain-specific cue engineering and few-shot learning techniques (i.e., few-shot learning, a machine learning technique designed to enable models to quickly adapt to new tasks or identify new categories using only a small amount of labeled data, such as 50 or fewer), are used to perform hierarchical semantic parsing processing on the semantic embedding vectors: first, the overall intent of the instruction is identified (e.g., "grab intent," "place intent," or "move intent"), then it is refined layer by layer to specific parameters (e.g., the target object to be grasped in "grab intent," the target position to be placed in "place intent," or the movement constraints in "move intent"), and it also supports the identification of at least one identified subtask (i.e., identified subtasks) and their execution order in multi-step instructions (e.g., "put the cup on the right side of the table"), such as... Figure 3 As shown.

[0072] In step S2, specifically, during the hierarchical semantic parsing process, the loss function of semantic parsing... It is expressed as follows:

[0073]

[0074] In the formula, This represents the total number of natural language instruction samples. Indicates less than or equal to positive integers, Indicates the relationship with the first The real semantic labels corresponding to each natural language instruction sample (i.e., the actual structured task representations required). Indicates the first A sample of natural language instructions. This represents the model parameters of the large language model. This represents the probability distribution function predicted by the large language model. Furthermore, as... Figure 3 As shown, while obtaining the structured task representation, corresponding confidence evaluation information, such as task feasibility and parameter completeness, can also be output for historical backtracking.

[0075] S3. Fuse the multimodal sensing data to generate the environmental representation required for task execution.

[0076] In step S3, to enhance the adaptability of the environment representation to task requirements and environmental changes, preferably, the multimodal perception data is fused, including but not limited to: employing a multimodal dynamic fusion network based on a self-attention mechanism to optimize the modal contribution of each modal perception data in the multimodal perception data in real time according to task requirements and environmental changes, and then fusing the multimodal perception data according to all the modal contributions, wherein the modal contributions are expressed as fusion weights by the following formula:

[0077]

[0078] In the formula, This represents the total number of modalities in the multimodal sensing data. and They represent less than or equal to positive integers, In the multimodal sensing data, the first... The fusion weights of the modal sensing data This represents the natural exponential function. This indicates the preset adjustment parameters. Indicates the first The correlation between each modal sensing data and the target task and has , Indicates the first Feature vectors of modal sensing data Indicates the transpose symbol. The feature vector representing the target task, also known as the structured task representation. Indicates the first Confidence level of modal sensing data This represents the Sigmoid function. Indicates the first The sensitivity of each modal sensing data to environmental changes and has , This represents the preset sensitivity coefficient to environmental changes. This represents the environmental feature vector at the current moment. This represents the environmental feature vector from the previous time step. In the multimodal sensing data, the first... The correlation between each modal sensing data and the target task. Indicates the first The sensitivity of each modal sensing data point to environmental changes. Specifically, the adjustment parameters... The recommended range is [2.0, 10.0], and the environmental change sensitivity coefficient is... The recommended range is [0.5, 2.0], and the feature vector... and the confidence level The data can be obtained from the output of a sensing device of the corresponding modality. The current time refers to the current acquisition time of the multimodal sensing data, and the previous time refers to the previous acquisition time of the multimodal sensing data. The environmental feature vector can be obtained by performing conventional feature extraction processing on the multimodal sensing data. In addition, the Sigmoid function is an existing function, commonly used to normalize confidence scores.

[0079] S4. The structured task representation and the environment representation are transformed into a task graph, wherein the nodes in the task graph represent the identified subtasks, and the edges in the task graph represent the execution order.

[0080] In step S4, the task graph is preferably a directed graph. Specifically, converting the structured task representation and the environment representation into a task graph includes, but is not limited to, using a task graph generation model based on a graph neural network to convert the structured task representation and the environment representation into a task graph, wherein the nodes in the task graph represent the identified subtasks, the edges in the task graph represent the execution order, and the task graph generation model needs to minimize the optimization objective. It is expressed as follows:

[0081]

[0082] In the formula, This represents the total number of natural language instruction samples. Indicates less than or equal to positive integers, Indicates the relationship with the first Structured task representations corresponding to each natural language instruction sample. Indicates the relationship with the first The environment representation corresponding to the nth multimodal sensing data sample, the nth The first natural language instruction sample and the first It was obtained by acquiring multiple multimodal sensing data samples together. Indicates the first The first natural language instruction sample and the first The real task graph labels corresponding to each multimodal sensing data sample. This represents the probability distribution function predicted by the task graph generation model. This represents the preset complexity penalty coefficient. Indicates less than or equal to positive integers, This represents the number of nodes in the generated task graph. This represents the number of edges in the generated task graph, where the generated task graph refers to the graph generated by the task graph generation model that connects the first edge to the second edge. The structured task representation corresponding to the first natural language instruction sample and its relation to the first... The task graph is obtained by transforming the environment representation corresponding to each multimodal sensing data sample. The Graph Neural Network (GNN) is an algorithmic framework based on deep learning for processing graph-structured data. It performs tasks such as classification and prediction by extracting node, edge, and overall graph features, and is widely used in fields such as social networks and molecular structures. Therefore, the optimization objective can be applied in conventional applications. After optimizing the task graph generation model, the structured task representation and the environment representation can be transformed into a task graph, which can then dynamically adjust the task graph and support real-time environment changes.

[0083] S5. Convert the task diagram into joint control commands for the target robotic arm.

[0084] In step S5, to adaptively adjust constraints to ensure control accuracy and safety, preferably, the task diagram is converted into joint control commands for the target robotic arm, including but not limited to: during the conversion of the task diagram into joint control commands for the target robotic arm, an inverse kinematics algorithm capable of adaptively adjusting multiple constraints to solve the following optimization problem is used to obtain the joint angle vector used to generate the joint control commands:

[0085]

[0086] In the formula, Indicates minimization. This represents the joint angle vector. Represents the target space coordinate vector. Represents the forward kinematic function. Represents positive integers. Represents the first in the multiple constraints The constraints are in functional form, and the multiple constraints include, but are not limited to, joint limit constraints, load constraints, and / or obstacle constraints. This represents the preset constraint weight coefficients. The aforementioned forward kinematics (FK) is a fundamental problem in robotics, used to calculate the pose (position and orientation) of an end effector when joint variables are known. The aforementioned inverse kinematics is the process of solving for each joint parameter by knowing the target position and orientation of the end effector of a movable object. Therefore, the specific solution process can be derived conventionally based on existing technologies, thereby achieving the purpose of end effector localization.

[0087] S6. Send the joint control command to the joint controller for execution.

[0088] Therefore, based on the natural language-driven robotic arm control method described in steps S1 to S6 above, a new scheme for natural language-driven robotic arm control is provided, combining large language models, multimodal perception, and adaptive control technology. Specifically, after acquiring natural language instructions and multimodal perception data, the natural language instructions are first decomposed into semantic embedding vectors, and a structured task representation is obtained through hierarchical semantic parsing using a large language model. Then, multimodal perception data is fused to generate the environmental representation required for task execution. Next, the structured task representation and environmental representation are transformed into a task graph, and the task graph is converted into joint control instructions for the target robotic arm. Finally, the joint control instructions are sent to the joint controller for execution. Through multi-level semantic understanding, multimodal perception fusion, and adaptive execution control, the robotic arm can achieve accurate understanding and efficient execution of complex natural language instructions, improving its task execution accuracy, adaptability, flexibility, ease of use, and intelligence level in dynamic environments. This is particularly suitable for applications requiring high-precision human-machine interaction, such as intelligent manufacturing, service robots, and medical assistance, facilitating practical application and promotion.

[0089] Based on the technical solution of the first aspect described above, this embodiment also provides a possible design for closed-loop feedback optimization, that is, after sending the joint control command to the joint controller for execution, the method further includes, but is not limited to, the following steps S7 to S8.

[0090] S7. Collect the execution status monitoring results of the target robotic arm.

[0091] In step S7, the execution status monitoring results can be obtained by routine collection from relevant sensors and / or by routine user feedback.

[0092] S8. A reinforcement learning-based controller is used to optimize the execution of the target robotic arm's control task based on the execution state monitoring results, wherein the reward function of the reinforcement learning... It is expressed as follows:

[0093]

[0094] In the formula, 、 and These represent the preset weighting coefficients. Denotes the base of the natural logarithm. This indicates the control accuracy of the target robotic arm. This indicates the speed at which the target robotic arm executes its control tasks. This indicates the number of errors in the execution of the control task of the target robotic arm.

[0095] In step S8, reinforcement learning (RL) is a machine learning method based on the Markov decision process framework. It allows an agent to learn the optimal policy through trial and error in its interaction with the environment. The control task accuracy (in millimeters), the control task execution speed (in seconds), and the number of execution errors can all be routinely extracted from the execution state monitoring results. Therefore, reinforcement learning and the reward function can be used as a basis for this process. Based on the execution status monitoring results, the control tasks of the target robotic arm are optimized, thereby achieving adaptive learning and improving the success rate of robotic arm control.

[0096] Based on the aforementioned possible design one, task execution can be adjusted through real-time feedback to achieve adaptive learning and improve the success rate of robotic arm control. Specifically, task completion time can be shortened by about 20%, while the error rate can be reduced to below 5%, thereby effectively improving robustness.

[0097] like Figure 4 As shown, the second aspect of this embodiment provides a virtual device for implementing the natural language driven robotic arm control method described in the first aspect or possibly the first design, which is suitable for being arranged in a control device, wherein the control device is communicatively connected to a language command input device, a multimodal sensing device and a joint controller of the target robotic arm, and the multimodal sensing device is located in the surrounding environment of the target robotic arm;

[0098] The virtual device includes an instruction data acquisition unit, a semantic parsing and processing unit, a perception data fusion unit, a task graph conversion unit, a control instruction conversion unit, and a control instruction sending unit.

[0099] The instruction data acquisition unit is used to acquire natural language instructions input by the language instruction input device and multimodal perception data collected by the multimodal perception device, wherein the multimodal perception data includes visual data, tactile data and environmental data.

[0100] The semantic parsing processing unit is communicatively connected to the instruction data acquisition unit. It is used to decompose the natural language instruction into semantic embedding vectors and perform hierarchical semantic parsing processing on the semantic embedding vectors using a large language model based on the Transformer structure to obtain a structured task representation. The hierarchical semantic parsing processing includes sequentially performed intent recognition layer semantic parsing processing, parameter extraction layer semantic parsing processing, and task combination layer semantic parsing processing. The structured task representation includes the task intent obtained through the intent recognition layer semantic parsing processing, the task parameters obtained through the parameter extraction layer semantic parsing processing, and at least one identified subtask obtained through the task combination layer semantic parsing processing, as well as the execution order of the at least one identified subtask.

[0101] The perception data fusion unit is communicatively connected to the instruction data acquisition unit and is used to fuse the multimodal perception data to generate an environmental representation required for task execution.

[0102] The task graph transformation unit is communicatively connected to the semantic parsing processing unit and the perceptual data fusion unit, respectively, and is used to transform the structured task representation and the environment representation into a task graph, wherein the nodes in the task graph represent the identified sub-tasks, and the edges in the task graph represent the execution order;

[0103] The control command conversion unit is communicatively connected to the task diagram conversion unit and is used to convert the task diagram into joint control commands for the target robotic arm.

[0104] The control command sending unit is communicatively connected to the control command conversion unit and is used to send the joint control command to the joint controller for execution.

[0105] In one possible design, it also includes a monitoring result collection unit and a reinforcement learning control unit that are connected in communication.

[0106] The monitoring result collection unit is used to collect the execution status monitoring results of the target robotic arm after the joint control command is sent to the joint controller for execution;

[0107] The reinforcement learning control unit is used to optimize the execution of the target robotic arm's control task based on the execution state monitoring results using a reinforcement learning-based controller, wherein the reward function of the reinforcement learning... It is expressed as follows:

[0108]

[0109] In the formula, 、 and These represent the preset weighting coefficients. Denotes the base of the natural logarithm. This indicates the control accuracy of the target robotic arm. This indicates the speed at which the target robotic arm executes its control tasks. This indicates the number of errors in the execution of the control task of the target robotic arm.

[0110] The working process, working details and technical effects of the aforementioned device provided in the second aspect of this embodiment can be found in the natural language driven robotic arm control method described in the first aspect or possible design, and will not be repeated here.

[0111] like Figure 5 As shown, the third aspect of this embodiment provides a control device for executing the natural language-driven robotic arm control method as described in the first aspect or a possible design, including a storage module, a processing module, and a transceiver module connected in sequence. The storage module stores a computer program, the transceiver module sends and receives messages, and the processing module reads the computer program and executes the natural language-driven robotic arm control method as described in the first aspect or a possible design. Specifically, the storage module may include, but is not limited to, random-access memory (RAM), read-only memory (ROM), flash memory, first-in-first-out (FIFO) memory, and / or first-in-last-out (FILO) memory, etc.; the processing module may, but is not limited to, use a microprocessor of the STM32F105 series. Furthermore, the control device may also include, but is not limited to, a power supply module, a display screen, and other necessary components.

[0112] The working process, working details and technical effects of the control device provided in the third aspect of this embodiment can be found in the natural language driven robotic arm control method described in the first aspect or possible design, and will not be repeated here.

[0113] This fourth aspect of the embodiment provides a computer-readable storage medium storing instructions comprising the natural language-driven robotic arm control method as described in the first aspect or possible design one. Specifically, the computer-readable storage medium stores instructions that, when executed on a computer, perform the natural language-driven robotic arm control method as described in the first aspect or possible design one. The computer-readable storage medium refers to a data storage medium, which may include, but is not limited to, floppy disks, optical disks, hard disks, flash memory, USB flash drives, and / or Memory Sticks. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0114] The working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the fourth aspect of this embodiment can be found in the natural language driven robotic arm control method as described in the first aspect or possible design, and will not be repeated here.

[0115] This fifth aspect of the embodiment provides a computer program product, including a computer program or instructions, which, when executed by a computer, implement the natural language-driven robotic arm control method as described in the first aspect or possible design. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.

[0116] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, and refinements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A natural language-driven robotic arm control method, characterized in that, The control device is executed by a control device, which is communicatively connected to a language command input device, a multimodal sensing device, and a joint controller of the target robotic arm, wherein the multimodal sensing device is located in the surrounding environment of the target robotic arm. The natural language-driven robotic arm control method includes: Acquire natural language commands input by the language command input device and multimodal perception data collected by the multimodal perception device, wherein the multimodal perception data includes visual data, tactile data and environmental data; The natural language instruction is decomposed into semantic embedding vectors, and a large language model based on the Transformer structure is used to perform hierarchical semantic parsing on the semantic embedding vectors to obtain a structured task representation. The hierarchical semantic parsing includes sequential semantic parsing of the intent recognition layer, semantic parsing of the parameter extraction layer, and semantic parsing of the task combination layer. The structured task representation includes the task intent obtained by the semantic parsing of the intent recognition layer, the task parameters obtained by the semantic parsing of the parameter extraction layer, at least one identified subtask obtained by the semantic parsing of the task combination layer, and the execution order of the at least one identified subtask. The multimodal sensing data is fused to generate the environmental representation required for task execution; The transformation of the structured task representation and the environment representation into a task graph specifically includes: employing a task graph generation model based on a graph neural network to transform the structured task representation and the environment representation into a task graph, wherein nodes in the task graph represent the identified subtasks, edges in the task graph represent the execution order, and the task graph generation model has an optimization objective that needs to be minimized. It is expressed as follows: In the formula, This represents the total number of natural language instruction samples. Indicates less than or equal to positive integers, Indicates the relationship with the first Structured task representations corresponding to each natural language instruction sample. Indicates the relationship with the first The environment representation corresponding to the nth multimodal sensing data sample, the nth The first natural language instruction sample and the first It was obtained by acquiring multiple multimodal sensing data samples together. Indicates the first The first natural language instruction sample and the first The real task graph labels corresponding to each multimodal sensing data sample. This represents the probability distribution function predicted by the task graph generation model. This represents the preset complexity penalty coefficient. Indicates less than or equal to positive integers, This represents the number of nodes in the generated task graph. This represents the number of edges in the generated task graph, where the generated task graph refers to the graph generated using the task graph generation model that connects the first edge to the second edge. The structured task representation corresponding to the first natural language instruction sample and its relation to the first... The task graph is obtained by transforming the environmental representation corresponding to each multimodal sensing data sample; The task diagram is converted into joint control commands for the target robotic arm; The joint control command is sent to the joint controller for execution.

2. The natural language-driven robotic arm control method according to claim 1, characterized in that, In the process of hierarchical semantic parsing, the loss function of semantic parsing It is expressed as follows: In the formula, This represents the total number of natural language instruction samples. Indicates less than or equal to positive integers, Indicates the relationship with the first The actual semantic labels corresponding to each natural language instruction sample Indicates the first A sample of natural language instructions. This represents the model parameters of the large language model. This represents the probability distribution function predicted by the large language model.

3. The natural language-driven robotic arm control method according to claim 1, characterized in that, The fusion of the multimodal sensing data includes: A multimodal dynamic fusion network based on a self-attention mechanism is employed to optimize the modal contribution of each modal sensing data in the multimodal sensing data in real time according to task requirements and environmental changes. Then, the multimodal sensing data is fused based on all the modal contributions, wherein the modal contributions are expressed as fusion weights by the following formula: In the formula, This represents the total number of modalities in the multimodal sensing data. and They represent less than or equal to positive integers, In the multimodal sensing data, the first... The fusion weights of the modal sensing data This represents the natural exponential function. This indicates the preset adjustment parameters. Indicates the first The correlation between each modal sensing data and the target task and has , Indicates the first Feature vectors of modal sensing data Indicates the transpose symbol. The feature vector representing the target task, also known as the structured task representation. Indicates the first Confidence level of modal sensing data This represents the Sigmoid function. Indicates the first The sensitivity of each modal sensing data to environmental changes and has , This represents the preset sensitivity coefficient to environmental changes. This represents the environmental feature vector at the current moment. This represents the environmental feature vector from the previous time step. In the multimodal sensing data, the first... The correlation between each modal sensing data and the target task. Indicates the first The sensitivity of each modal sensing data to environmental changes.

4. The natural language-driven robotic arm control method according to claim 1, characterized in that, The task diagram is converted into joint control commands for the target robotic arm, including: In the process of converting the task diagram into joint control commands for the target robotic arm, an inverse kinematics algorithm capable of adaptive adjustment across multiple constraints is used to obtain the joint angle vectors used to generate the joint control commands. In the formula, Indicates minimization. This represents the joint angle vector. Represents the target space coordinate vector. Represents the forward kinematic function. Represents positive integers. Represents the first in the multiple constraints The constraints are in functional form, and the multiple constraints include joint limit constraints, load constraints, and / or obstacle constraints. This represents the preset constraint weight coefficient.

5. The natural language-driven robotic arm control method according to claim 1, characterized in that, After sending the joint control command to the joint controller for execution, the method further includes: Collect the execution status monitoring results of the target robotic arm; A reinforcement learning-based controller is used to optimize the execution of the target robotic arm's control tasks based on the execution state monitoring results, wherein the reward function of the reinforcement learning... It is expressed as follows: In the formula, 、 and These represent the preset weighting coefficients. Denotes the base of the natural logarithm. This indicates the control accuracy of the target robotic arm. This indicates the speed at which the target robotic arm executes its control tasks. This indicates the number of errors in the execution of the control task of the target robotic arm.

6. A natural language-driven robotic arm control device, characterized in that, Suitable for placement in a control device, wherein the control device is communicatively connected to a language command input device, a multimodal sensing device, and a joint controller of a target robotic arm, wherein the multimodal sensing device is located in the surrounding environment of the target robotic arm; The natural language driven robotic arm control device includes an instruction data acquisition unit, a semantic parsing and processing unit, a perception data fusion unit, a task graph conversion unit, a control instruction conversion unit, and a control instruction sending unit. The instruction data acquisition unit is used to acquire natural language instructions input by the language instruction input device and multimodal perception data collected by the multimodal perception device, wherein the multimodal perception data includes visual data, tactile data and environmental data. The semantic parsing processing unit is communicatively connected to the instruction data acquisition unit. It is used to decompose the natural language instruction into semantic embedding vectors and perform hierarchical semantic parsing processing on the semantic embedding vectors using a large language model based on the Transformer structure to obtain a structured task representation. The hierarchical semantic parsing processing includes sequentially performed intent recognition layer semantic parsing processing, parameter extraction layer semantic parsing processing, and task combination layer semantic parsing processing. The structured task representation includes the task intent obtained through the intent recognition layer semantic parsing processing, the task parameters obtained through the parameter extraction layer semantic parsing processing, and at least one identified subtask obtained through the task combination layer semantic parsing processing, as well as the execution order of the at least one identified subtask. The perception data fusion unit is communicatively connected to the instruction data acquisition unit and is used to fuse the multimodal perception data to generate an environmental representation required for task execution. The task graph transformation unit, communicatively connected to the semantic parsing processing unit and the perceptual data fusion unit, is used to transform the structured task representation and the environment representation into a task graph. Specifically, this includes employing a task graph generation model based on a graph neural network to transform the structured task representation and the environment representation into a task graph. In the task graph, nodes represent identified subtasks, edges represent the execution order, and the task graph generation model has an optimization objective that needs to be minimized. It is expressed as follows: In the formula, This represents the total number of natural language instruction samples. Indicates less than or equal to positive integers, Indicates the relationship with the first Structured task representations corresponding to each natural language instruction sample. Indicates the relationship with the first The environment representation corresponding to the nth multimodal sensing data sample, the nth The first natural language instruction sample and the first It was obtained by acquiring multiple multimodal sensing data samples together. Indicates the first The first natural language instruction sample and the first The real task graph labels corresponding to each multimodal sensing data sample. This represents the probability distribution function predicted by the task graph generation model. This represents the preset complexity penalty coefficient. Indicates less than or equal to positive integers, This represents the number of nodes in the generated task graph. This represents the number of edges in the generated task graph, where the generated task graph refers to the graph generated using the task graph generation model that connects the first edge to the second edge. The structured task representation corresponding to the first natural language instruction sample and its relation to the first... The task graph is obtained by transforming the environmental representation corresponding to each multimodal sensing data sample; The control command conversion unit is communicatively connected to the task diagram conversion unit and is used to convert the task diagram into joint control commands for the target robotic arm. The control command sending unit is communicatively connected to the control command conversion unit and is used to send the joint control command to the joint controller for execution.

7. A control device, characterized in that, The device includes a storage module, a processing module, and a transceiver module that are sequentially connected in communication. The storage module is used to store computer programs, the transceiver module is used to send and receive messages, and the processing module is used to read the computer programs and execute the natural language driven robotic arm control method as described in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that... The computer-readable storage medium stores instructions that, when executed on a computer, perform the natural language-driven robotic arm control method as described in any one of claims 1 to 5.

9. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or the instructions are executed by the computer, they implement the natural language-driven robotic arm control method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Cloud edge-end resource scheduling optimization method based on double-layer graph neural network

    CN118134029A

  • Robot control method and device based on multi-modal data fusion

    CN119260752A