Robot operation skill learning method based on big language model enhanced fuzzy semantic instruction reasoning

By introducing large language models and symbolic thinking chain technology into robot systems, fuzzy semantic instructions are transformed into symbolic operation sequences, which solves the problem that robots find it difficult to deal with fuzzy instructions, and achieves more efficient and autonomous task execution.

CN120069070APending Publication Date: 2025-05-30NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510128304.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing robot systems are difficult to accurately understand and execute when processing fuzzy semantic instructions, resulting in the inability to complete complex tasks and need to rely on complex rule definitions or human intervention, limiting the autonomy and adaptability of robots.

Method used

By introducing large language model (LLM) and tip fine-tuning technology, symbolic thinking chains are used to decompose fuzzy semantic instructions into symbolic logical expressions and sub-task operation sequences, and then skill learning is achieved by imitating expert demonstration data to achieve efficient execution of robots in complex tasks.

Benefits of technology

It improves the flexibility and efficiency of the robot when facing fuzzy, unstructured task descriptions, reduces understanding bias and execution errors, and enhances the robot's autonomy and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069070A_ABST
    Figure CN120069070A_ABST
Patent Text Reader

Abstract

The invention discloses a robot operation skill learning method and device for enhancing fuzzy semantic instruction reasoning based on a large language model, a medium and equipment. The method comprises the steps of obtaining a to-be-executed task; utilizing the symbolized thinking chain to guide the large language model to decompose the to-be-executed task into a plurality of symbolized subtask operation sequences; performing skill learning of staged sub-task operation strategies based on the sub-task operation sequence by simulating expert demonstration data; other to-be-processed tasks are processed on the basis of learned operation skills, the symbolization reasoning ability of a large language model is improved by introducing the thinking chain technology, the robot can convert fuzzy semantic instructions into symbolization logic representation and action sequences, and therefore operation can be accurately executed in complex tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of robotics and artificial intelligence, and particularly to a method, apparatus, medium, and device for learning robot operation skills based on enhancing fuzzy semantic instruction reasoning using a large language model. Background Art

[0002] With the rapid development of artificial intelligence and robotics technologies, the automated learning and execution of robot operation skills have become a key research direction. When faced with complex, multi-step tasks, current robots typically rely on clear, pre-defined task instructions. However, in real-world applications, task instructions often carry ambiguity or uncertainty, making it difficult for robots to accurately understand and execute them. For example, in home or industrial scenarios, users often give task instructions such as "Put the book on the bookshelf" or "Tidy up the table", where the specific details in these instructions are not clearly specified, and robots are prone to understanding and execution deviations when parsing these fuzzy semantic instructions, resulting in the inability to complete the task. Existing robot systems usually need to rely on complex rule definitions or human intervention when dealing with fuzzy instructions, which greatly limits the autonomy and adaptability of robots.

[0003] Large language models (LLMs), such as GPT-4, have demonstrated powerful language understanding and generation capabilities in natural language processing tasks in recent years. These models can learn complex semantic relationships from large amounts of language data and exhibit excellent capabilities in reasoning and generating natural language text. Therefore, semantic enhancement reasoning technology based on LLMs provides a new method for solving the problems of robots in dealing with fuzzy semantic instructions. By leveraging LLMs, robots can deeply understand fuzzy instructions and convert fuzzy semantics into clear, executable operation instructions, thereby enhancing their adaptability to complex tasks.

[0004] Despite the significant progress made by LLMs in the field of language understanding, applying them to robot operation tasks still faces challenges. Especially when dealing with fuzzy instructions, how to ensure that the model can accurately understand and generate instructions that conform to the task scenario remains a key issue. Summary of the Invention

[0005] The main objective of this application is to provide a method, apparatus, medium, and device for learning robot operation skills based on enhancing fuzzy semantic instruction reasoning using a large language model, aiming to use prompt tuning technology to guide the large language model to generate clear logical reasoning paths and help robots efficiently execute instructions in complex tasks.

[0006] To achieve the above object, a first aspect of the present application provides a robot operation skill learning method for enhancing fuzzy semantic instruction reasoning based on a large language model, including: This method is executed by a robot, and the method includes: obtaining a task to be executed; using a symbolic thinking chain to guide the large language model to decompose the task to be executed into a plurality of symbolic subtask operation sequences; through imitating expert demonstration data, performing skill learning of subtask operation strategies in stages based on the subtask operation sequences; and processing other tasks to be processed based on the learned operation skills.

[0007] Optionally, the task to be executed is a natural language instruction; the using a symbolic thinking chain to guide the large language model to decompose the task to be executed into a plurality of symbolic subtask operation sequences includes: using the symbolic thinking chain to guide the large language model to decompose the task to be executed into a plurality of symbolic logical expressions; and generating a symbolic subtask operation sequence based on the plurality of symbolic logical expressions.

[0008] Optionally, the using the symbolic thinking chain to guide the large language model to decompose the task to be executed into a plurality of symbolic logical expressions includes: guiding the large language model to process the natural language instruction based on the symbolic thinking chain to obtain a plurality of key objects and corresponding operation targets, and associating the plurality of key objects and the corresponding operation targets; and reasoning the relationships between the key objects according to each operation target, where the relationships between the key objects include a dependency relationship and an interaction relationship; and generating each symbolic logical expression corresponding to the relationships between the key objects.

[0009] Optionally, the using a plurality of symbolic logical expressions to guide the large language model to generate a symbolic subtask operation sequence based on the plurality of symbolic logical expressions includes: using the plurality of symbolic logical expressions to guide the large language model to process the symbolic logical expressions to obtain a plurality of subtask target conditions; and determining a subtask sequence for completing the task to be executed based on the plurality of subtask target conditions, where each subtask sequence includes; and determining the logical relationships between the key objects based on the corresponding subtask target conditions for completing each subtask.

[0010] Optionally, the symbolic logical expressions include: a single object symbolic logical expression and a multi-object symbolic logical expression, where the single object symbolic logical expression is determined based on the target object / area and the subtask target conditions, and the multi-object symbolic logical expression is determined based on the target object / multiple areas and the subtask target conditions.

[0011] Optionally, the skill learning of the sub-task operation strategy in stages based on the sub-task operation sequence by imitating the expert demonstration data includes: imitating the expert demonstration data based on the behavior cloning algorithm to obtain state-action pairs in different environments, and inputting the state-action pairs and the sub-task operation sequence into the Transformer model to obtain the dependency relationship between the states and actions of each sub-task operation sequence; processing the dependency relationship between each state and action based on the Gaussian mixture model to generate multiple predicted action distributions; and performing skill learning of the sub-task operation strategy in stages based on the multiple predicted action distributions.

[0012] Optionally, the performing skill learning of the sub-task operation strategy in stages based on the multiple predicted action distributions includes: processing the multiple predicted action distributions by using a multi-layer perceptron to obtain the optimal action prediction for each sub-task operation sequence; and performing skill learning of the sub-task operation strategy in stages based on the multiple optimal action predictions of each sub-task operation sequence.

[0013] In addition, to achieve the above object, a robot operation skill learning device for enhancing fuzzy semantic instruction reasoning based on a large language model according to the second aspect of the present application includes: an acquisition module for acquiring a task to be executed; a symbolization module for guiding the large language model to decompose the task to be executed into multiple symbolized sub-task operation sequences by using symbolic thinking chains; a strategy learning module for performing skill learning of the sub-task operation strategy in stages based on the sub-task operation sequence by imitating expert demonstration data; and a task processing module for processing other tasks to be processed based on the learned operation skills.

[0014] To achieve the above object, a computer-readable storage medium according to the third aspect of the present application includes instructions that, when running on a computer, cause the computer to execute the robot operation skill learning method for enhancing fuzzy semantic instruction reasoning based on a large language model provided in the first aspect.

[0015] To achieve the above object, an electronic device according to the fourth aspect of the present application includes: at least one processor, a memory, and an input / output unit; wherein, the memory is used for storing a computer program, and the processor is used for calling the computer program stored in the memory to execute the robot operation skill learning method for enhancing fuzzy semantic instruction reasoning based on a large language model provided in the first aspect.

[0016] A robot operation skill learning method, device, medium and equipment based on enhancing fuzzy semantic instruction reasoning by a large language model proposed in an embodiment of the present application obtains a task to be executed; uses a symbolic thinking chain to guide the large language model to decompose the task to be executed into a symbolic sequence of multiple subtask operations; through imitating expert demonstration data, conducts skill learning of subtask operation strategies in stages based on the subtask operation sequence; and processes other tasks to be processed based on the learned operation skills. By introducing the thinking chain technology, the present application improves the symbolic reasoning ability of the large language model, enables the robot to convert fuzzy semantic instructions into symbolic logical representations and action sequences, so as to accurately execute operations in complex tasks. Through the large language model and prompt fine-tuning, the robot can accurately understand and execute fuzzy instructions; the symbolic reasoning method makes the task execution process clearer and more operable, thus reducing understanding deviations and execution errors; by using the symbolic thinking chain technology, complex natural language instructions are decomposed into clear symbolic logics and operation sequences, making the task execution clearer, more orderly and easier to implement. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 FIG. is a schematic flowchart provided by an embodiment of the robot operation skill learning method based on enhancing fuzzy semantic instruction reasoning by a large language model of the present application;

[0018] Figure 2 FIG. is a process of a robot generating a task plan based on a thinking chain provided by an embodiment of the robot operation skill learning method based on enhancing fuzzy semantic instruction reasoning by a large language model of the present application;

[0019] Figure 3 FIG. is the NutAssemblySquare scene of Robosuite provided by an embodiment of the robot operation skill learning method based on enhancing fuzzy semantic instruction reasoning by a large language model of the present application;

[0020] Figure 4 FIG. is a training success rate curve provided by an embodiment of the robot operation skill learning method based on enhancing fuzzy semantic instruction reasoning by a large language model of the present application;

[0021] Figure 5 FIG. is a structural block diagram provided by an embodiment of the robot operation skill learning device based on enhancing fuzzy semantic instruction reasoning by a large language model of the present application.

[0022] The implementation, functional features and advantages of the objectives of the present application will be further described with reference to the accompanying drawings in conjunction with the embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0024] In the prior art, in order to improve the reasoning ability of LLM when processing fuzzy semantic instructions, the "Prompt Tuning" technology is introduced as an effective enhancement means.

[0025] Prompt Tuning can significantly improve the reasoning ability of large language models on specific tasks through a small number of additional parameters and carefully designed input prompts, without the need to massively modify the model structure. For the fuzzy semantic instruction reasoning task in the present invention, the Prompt Tuning technology can guide the LLM to better understand the context of the task scenario and generate more accurate operation instructions. For example, after receiving the fuzzy instruction "Put the book on the bookshelf", through the Prompt Tuning model, it can understand the structure and function of the "bookshelf" in the specific scenario, and even infer the possible user preferences, and finally convert the fuzzy semantics into the specific operation of "Put the book in the middle position of the second layer".

[0026] The introduction of Prompt Tuning has brought a significant improvement to the LLM reasoning ability of the present invention, enabling the robot to complete complex operation tasks more flexibly and efficiently when facing fuzzy and unstructured task descriptions. In addition, Prompt Tuning also has good scalability and can adapt to the needs of different tasks, further enhancing the versatility and intelligence of the robot system.

[0027] The purpose of the present invention is to provide a robot operation skill learning method for enhancing fuzzy semantic instruction reasoning based on a large language model (LLM), using the Prompt Tuning technology to guide the large language model to generate a clear logical reasoning path, and helping the robot to efficiently execute instructions in complex tasks. By introducing the Chain of Thought technology, the symbolic reasoning ability of the large language model is improved, enabling the robot to convert fuzzy semantic instructions into a sub-task operation sequence represented by symbolic logic, so as to accurately execute operations in complex tasks. The generated sub-task operation sequence and the state-action pairs of the expert demonstration data are used as the input of the Transformer model. The Transformer extracts temporal features through the self-attention mechanism and captures the dependencies between states and actions. These features are then passed to the Gaussian Mixture Model (GMM), which models the actions, captures the multi-modal characteristics of the actions, and generates multiple possible action distributions. Finally, after being processed by the Multi-Layer Perceptron (MLP), the model outputs the optimal action prediction. By imitating the expert demonstration data, the robot can execute the operation strategy and verify the effectiveness of this method in the simulation environment.

[0028] To achieve the above object, the invention content involved in the present invention is as follows:

[0029] Refer to Figure 1, the robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning provided by the first embodiment of this application can be executed by a robot. The robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning may include:

[0030] S10. Obtain the task to be executed.

[0031] S20. Use symbolic thinking chain to guide the large language model to decompose the task to be executed into a symbolic sequence of multiple subtask operations

[0032] It should be noted that the large language model includes at least one of GPT-4, Claude 3, Gemini 1.5, LLaMA3, Mistral7B, Tongyi Qianwen, Wenxin Yiyan, and iFlytek Spark.

[0033] In this application, through prompt tuning technology, the concept of "thinking chain" is introduced to guide the large language model to gradually reason and generate an executable instruction sequence in stages. The thinking chain refers to the process of decomposing complex fuzzy instructions into a series of clear logical relationships and action plans through multi-level and multi-step reasoning. The design process includes two stages: translation and planning, aiming to convert fuzzy natural language instructions into executable operation steps

[0034] S30. Through imitating expert demonstration data, perform skill learning of subtask operation strategies in stages based on the subtask operation sequence;

[0035] After the symbolic thinking chain completes reasoning and planning, by imitating expert demonstration data, the robot learns the subtask operation strategies in stages to obtain the task execution strategies of the robot in different environments.

[0036] S40. Process other tasks to be processed based on the learned operation skills.

[0037] Specifically, the main inventive content of this application includes: symbolic thinking chain prompt tuning design; semantic reasoning and symbolic operation generation of large language models. Through in-depth semantic reasoning of natural language instructions by the large language model, and through prompt tuning technology, fuzzy instructions are converted into symbolic logic and subtask operation sequences, enabling the robot to gradually understand and execute complex tasks; imitating expert demonstration data to learn operation strategies; verifying this method in a simulation environment, proving its effectiveness in complex tasks.

[0038] In summary, the robot operation skill learning method based on enhancing fuzzy semantic instruction reasoning using a large language model provided in this embodiment uses prompt fine-tuning technology to guide the large language model to generate clear logical reasoning paths, helping the robot to efficiently execute instructions in complex tasks. By introducing the thought chain technology, the symbolic reasoning ability of the large language model is improved, enabling the robot to transform fuzzy semantic instructions into a sub-task operation sequence represented by symbolic logic, so as to accurately execute operations in complex tasks. The generated sub-task operation sequence and the state-action pairs of expert demonstration data are used as the input of the Transformer model. The Transformer extracts temporal features through the self-attention mechanism and captures the dependencies between states and actions. These features are then passed to the Gaussian Mixture Model (GMM), which models the actions, captures the multi-modal characteristics of the actions, and generates multiple possible action distributions. Finally, after being processed by a Multi-Layer Perceptron (MLP), the model outputs the optimal action prediction. By imitating expert demonstration data, the robot can execute the operation strategy and verify the effectiveness of this method in a simulation environment.

[0039] In an embodiment of the present application, the task to be executed is a natural language instruction; step S20 may include the following execution process:

[0040] S201. Use the symbolic thought chain to guide the large language model to decompose the task to be executed into multiple symbolic logical expressions in symbolic form;

[0041] S202. And generate a symbolic sub-task operation sequence based on the multiple symbolic logical expressions

[0042] Specifically, step S201 may include the following execution process:

[0043] S2011. Based on the symbolic thought chain, guide the large language model to process the natural language instruction, obtain multiple key objects and corresponding operation targets, and associate the multiple key objects and corresponding operation targets;

[0044] S2012. And infer the relationships between the key objects according to each operation target, where the relationships between the key objects include dependency relationships and interaction relationships;

[0045] S2013. And generate each symbolic logical expression correspondingly based on the relationships between the above key objects.

[0046] Specifically, in the translation stage of the thought chain, the task of the thought chain in the translation stage is to deconstruct the fuzzy semantics in the natural language instruction into a logical form and provide a clear symbolic representation. The reasoning and transformation in this stage are carried out layer by layer, so that each reasoning step unfolds along the logical train of thought.

[0047] (1) Identify objects: The initial reasoning step of the chain of thought is to identify the key objects in the natural language instruction and associate these objects with the task goal. This process ensures that the robot can extract the core information from the fuzzy semantics. For example, if the task description is: pull the microwave door open. The LLM needs to understand that the target object of the operation at this time is not the "door", but the "door handle".

[0048] (2) Define relationships: After identifying the objects, the next level of reasoning in the chain of thought analyzes the relationships between these objects to clarify the interaction or dependency relationships between them. This reasoning process can help the system build the logical associations in the task and lay the foundation for subsequent task planning.

[0049] (3) Form a logical form: Through the derivation of the above objects and relationships, the system generates a symbolic logical expression, transforming the complex natural language instruction into a structured symbolic representation. This logical form serves as the input for task planning, giving the instruction execution a clear direction, such as Figure 2 the format form in the left prompt:

[0050] (<Target object / area>, <Sub-task target condition>)

[0051] In the embodiment of the present application, step S30 may include the following execution process:

[0052] S301. Use multiple symbolic logical expressions to guide the use of a large language model to process the symbolic logical expressions, and obtain multiple sub-task target conditions;

[0053] S302. And determine the sub-task sequence for completing the task to be executed based on the multiple sub-task target conditions, where each sub-task sequence includes.

[0054] Specifically, in the planning stage of the chain of thought, based on the symbolic logical form generated in the translation stage, it is further refined into executable operation steps. The chain of thought continues to guide the reasoning at this stage, decomposing the logically derived task into a series of specific sub-tasks and action steps.

[0055] (1) Decompose the task: The present invention adds sub-task target conditions, such as: grasp, place, push, pull, rotate. The system can decompose the complex task into the above five sub-task target conditions and gradually generate an executable sub-task chain.

[0056] (2) Generate the sub-task sequence: After the reasoning of the chain of thought is completed, the system generates a specific operation sequence according to the sub-task chain. This operation sequence is the minimum and reasonable actions required to complete the task.

[0057] By executing step S301 - step S302, the robot further refines the symbolized logical form generated in the translation stage into executable operation steps in this stage. The chain of thought continues to guide reasoning in this stage, decomposing the tasks derived from logic into a series of specific sub - task chains.

[0058] In an embodiment of the present application, step S40 may include the following execution process:

[0059] S404. Based on the behavior cloning algorithm, imitate the expert demonstration data to obtain state - action pairs in different environments, and input the state - action pairs and the sub - task operation sequence into the Transformer model to obtain the dependency relationship between the states and actions of each sub - task operation sequence;

[0060] S402. Process the dependency relationship between each state and action based on the Gaussian mixture model to generate multiple predicted action distributions;

[0061] S403. Based on the multiple predicted action distributions, perform skill learning on the phased sub - task operation strategy.

[0062] After the symbolized chain of thought completes reasoning and planning, the behavior cloning (BC) algorithm is introduced to learn the task execution strategy of the robot in different environments. By imitating the expert demonstration data, the robot conducts phased learning of the sub - task operation strategy.

[0063] Specifically, the present application improves the symbolized reasoning ability of the large - language model by introducing the chain - of - thought technology, enabling the robot to transform fuzzy semantic instructions into sub - task operation sequences represented by symbolized logic, so as to accurately execute operations in complex tasks. The generated sub - task operation sequence and the state - action pairs of the expert demonstration data are used as the input of the Transformer model. The Transformer extracts temporal features through the self - attention mechanism to capture the dependency relationship between states and actions. These features are then passed to the Gaussian mixture model (GMM), and the GMM models the actions to capture the multi - modal characteristics of the actions and generate multiple possible action distributions.

[0064] In an embodiment of the present application, step S403 may include the following execution process:

[0065] S4031. Use a multi - layer perceptron to process the multiple predicted action distributions to obtain the optimal action prediction for each sub - task operation sequence;

[0066] S4032. Based on the multiple optimal action predictions of each sub - task operation sequence, perform skill learning on the phased sub - task operation strategy.

[0067] Generating multiple possible action distributions, processed by a multi-layer perceptron (MLP), the model outputs the optimal action prediction. By imitating expert demonstration data, the robot can execute the operation strategy and verify the effectiveness of this method in a simulation environment.

[0068] Exemplarily, taking the task "The gold nut goes on the gold peg" as an example, the implementation steps of robot operation skill learning based on large language model enhanced fuzzy semantic instruction reasoning are as follows:

[0069] The first step: Identify objects. By calling GPT-4, the robot identifies relevant objects (such as "gold nut" and "gold peg") and operation actions (such as "goes on") from the task described in natural language.

[0070] The second step: Define relationships. Use GPT-4 to analyze the relationships between objects. GPT-4 infers the interaction relationships between objects based on "goes on". Through the set subtask target conditions: grasp, place, push, pull, rotate), it is inferred that the subtask target conditions for completing this task are "grasp" and "place".

[0071] The third step: Form a logical form. GPT-4 generates a symbolic logical expression based on the identified objects and the relationships between them. This process involves decomposing the task described in natural language into a series of structured logical steps, in the format of (<target object / area>, <subtask target condition>).

[0072] The fourth step: Task decomposition. GPT-4 further decomposes the symbolic logic into specific executable tasks. Each task is mapped to specific steps of robot operation through the reasoning of GPT-4, such as picking up the gold nut and inserting the gold nut into the gold peg.

[0073] The fifth step: Generate a subtask operation sequence. After generating the specific tasks, GPT-4 generates an executable action sequence according to the symbolic form. The robot then executes the tasks step by step according to the action sequence generated by GPT-4, and completes each step by calling each operation function.

[0074] (gold nut,grasp)

[0075] (gold nut,place, gold peg)

[0076] The sixth step: Policy learning. On the basis of completing the symbolic tasks, imitation learning is introduced. By imitating expert demonstration data, the robot learns the subtask operation strategy in stages.

Specific Embodiment

[0078] The demonstration data of 200 jack tasks was collected and trained for 1000 rounds, Figure 3 which is the simulation environment scenario. Figure 4 For the comparison between adding the task planning module and not adding the task planning module, it can be seen that the learning speed of the fuzzy semantic instruction task planning module is significantly faster in the first 300 rounds.

[0079] Reference Figure 5 , based on the above embodiments, the present application further provides a robot operation skill learning device for enhancing fuzzy semantic instruction reasoning based on a large language model. The robot operation skill learning device 100 includes a symbolization module 101, a symbolization module 102, a policy learning module 103, and a task processing module 104. Among them,

[0080] The acquisition module 101 is used to acquire the task to be executed;

[0081] The symbolization module 102 is used to guide the large language model to decompose the task to be executed into a symbolic multi-subtask operation sequence by using symbolic thinking chains;

[0082] The policy learning module 103 learns the skills of the sub-task operation strategy in stages based on the sub-task operation sequence by imitating the expert demonstration data;

[0083] The task processing module 104 is used to process other tasks to be processed based on the learned operation skills.

[0084] Based on the above embodiments, the present application further provides a computer-readable storage medium, which includes instructions. When it runs on a computer, it enables the computer to execute the robot operation skill learning method for enhancing fuzzy semantic instruction reasoning provided in any of the foregoing embodiments.

[0085] Based on the above embodiments, the present application further provides an electronic device. The electronic device includes: at least one processor, a memory, and an input / output unit; among them, the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the robot operation skill learning method for enhancing fuzzy semantic instruction reasoning provided in the foregoing embodiments.

[0086] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present application by the same token.

Claims

1. A robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning, characterized in that: The method is performed by a robot, and the method comprises: Get the tasks to be executed; Using a symbolic thinking chain to guide the large language model to decompose the task to be performed into a plurality of symbolic subtask operation sequences; By imitating expert demonstration data, the skill learning of subtask operation strategies is carried out in stages based on the subtask operation sequence; Process the other pending tasks based on the learned operation skills.

2. The robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning as claimed in claim 1, characterized in that: The task to be executed is a natural language instruction; The step of utilizing the symbolic thinking chain to guide the large language model to decompose the task to be performed into a plurality of symbolic subtask operation sequences includes: Using the symbolized thinking chain to guide the large language model to decompose the task to be executed into a plurality of symbolized symbolized logical expressions; And generate a symbolic subtask operation sequence based on the multiple symbolic logic expressions.

3. The robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning as claimed in claim 2, characterized in that: The step of using the symbolic thinking chain to guide the large language model to decompose the task to be performed into a plurality of symbolic logical expressions includes: Guiding the large language model to process the natural language instruction based on the symbolic thinking chain to obtain a plurality of key objects and corresponding operation targets, and associating the plurality of key objects and the corresponding operation targets; and inferring the relationship between the key objects according to the operation objectives, wherein the relationship between the key objects includes a dependency relationship and an interaction relationship; And, based on the relationship between the key objects, each symbolized logical expression is generated accordingly.

4. The robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning as claimed in claim 2, characterized in that: The step of using the plurality of symbolic logic expressions to guide the large language model to generate a symbolic subtask operation sequence based on the plurality of symbolic logic expressions includes: Using the plurality of the symbolic logical expressions to guide the use of the large language model to process the symbolic logical expressions, to obtain a plurality of subtask target conditions; And based on the multiple subtask target conditions, a subtask sequence for completing the task to be performed is determined, wherein each subtask sequence includes.

5. The robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning as claimed in claim 2, characterized in that: The symbolic logical expression includes: A single-object symbolic logical expression and a multi-object symbolic logical expression, wherein the single-object symbolic logical expression is determined based on the target object / area and the sub-task target conditions, and the multi-object symbolic logical expression is determined based on the target object / multiple areas and the sub-task target conditions.

6. The robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning as claimed in claim 1, characterized in that: The skill learning of the subtask operation strategy in stages based on the subtask operation sequence by imitating the expert demonstration data includes: Based on the behavior cloning algorithm, imitate the expert demonstration data to obtain state-action pairs in different environments, and input the state-action pairs and the subtask operation sequence into the Transformer model to obtain the dependency relationship between the state and action of each subtask operation sequence; Processing the dependency between each of the states and actions based on a Gaussian mixture model to generate a plurality of predicted action distributions; Based on the multiple predicted action distributions, skill learning of sub-task operation strategies is performed in stages.

7. The robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning as claimed in claim 6, characterized in that: The skill learning of sub-task operation strategies in stages based on the plurality of predicted action distributions includes: Processing a plurality of the predicted action distributions using a multi-layer perceptron to obtain an optimal action prediction for each of the subtask operation sequences; Based on multiple optimal action predictions of each of the subtask operation sequences, skill learning of the subtask operation strategies is performed in stages.

8. A robot operation skill learning device based on large language model enhanced fuzzy semantic instruction reasoning, characterized in that: include: The acquisition module is used to obtain tasks to be executed; A symbolization module, used for guiding the large language model to decompose the task to be executed into a plurality of symbolized subtask operation sequences by using a symbolic thinking chain; The strategy learning module learns the skills of subtask operation strategies in stages based on the subtask operation sequence by imitating expert demonstration data; The task processing module is used to process other tasks to be processed based on the learned operation skills.

9. A computer-readable storage medium, characterized in that: It includes instructions, which, when running on a computer, enable the computer to execute the robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning as described in any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor, memory, and input-output unit; The memory is used to store computer programs, and the processor is used to call the computer programs stored in the memory to execute the robot operation skill learning method based on large language model enhanced fuzzy semantic instruction reasoning according to any one of claims 1 to 7.

Citation Information

Cited By

  • Robot low-code editable task sequence generation system based on large model

    CN120370756A

  • A low-code editable task sequence generation system for robots based on large models

    CN120370756B

  • Humanoid robot industrial task scene generation method and system based on large language model

    CN120735021A