Intelligent planning method for task sequence based on language vision big model and knowledge graph

By combining language visual model and hierarchical knowledge graph, the input files of the PDDL planner are automatically generated, which solves the problem of difficult to combine language, visual model and complex operation task planning in the existing technology, and realizes flexible, autonomous and intelligent planning of complex daily life tasks.

CN117874258BActive Publication Date: 2025-05-13BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410113790.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-26
Publication Date
2025-05-13
Estimated Expiration
2044-01-26

AI Technical Summary

Technical Problem

The prior art is difficult to combine language and visual models with the planning of complex operational tasks, especially in complex environments, the need for state recognition between items is difficult to meet.

Method used

A scene relationship perception and prediction model based on language vision big model is adopted, combined with a hierarchical knowledge graph, and the input file of the PDDL planner is automatically generated to plan action primitive sequences.

Benefits of technology

It realizes flexible, autonomous and intelligent planning of complex daily life tasks, avoids the problem of inconsistency between human language and robot ontology language, and improves the autonomy and flexibility of robot operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117874258B_ABST
    Figure CN117874258B_ABST
Patent Text Reader

Abstract

The present invention discloses a task sequence intelligent planning method based on a language vision big model and a knowledge graph, including hierarchical knowledge graph construction, scene relationship perception and prediction model training based on a language vision big model, and PDDL autonomous action sequence planning. Relevant object and action attributes are queried and extracted from the knowledge graph constructed by the hierarchical knowledge graph, which is used to automatically generate a domain file for the PDDL planner. At the same time, based on the scene relationship perception and prediction model training of the language vision big model, the initial and target states of the relevant objects are output, and the problem file of the PDDL planner is automatically generated; the domain file and the problem file drive the PDDL autonomous action sequence planning action primitive sequence. The present invention can be applied to the field of robot intelligent planning. It only needs to input human language task instructions and perceive the visual picture of the current scene, and the robot can associate and intelligently plan related information according to the built-in knowledge graph to obtain a feasible operation layer action primitive sequence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the research fields of language vision big model design, knowledge graph construction and task intelligent planning, and specifically to a task sequence intelligent planning scheme based on language vision big model and knowledge graph, including the design of language vision big model, the construction of knowledge graph and the task sequence intelligent planning based on this. Background Art

[0002] In recent years, language large models have developed rapidly, such as ChatGPT and GPT4, which have language, knowledge, and simple reasoning capabilities, and can well approximate human intelligent behavior. Many researchers have begun to apply large models to robot operations. How to combine the ability of language and visual large models to understand scenes and tasks with the planning of complex operation tasks is a hot and difficult issue in current research.

[0003] Among the existing processing methods, the traditional idea is to use syntactic analysis in natural language understanding to obtain key information, and obtain the position and posture information of related objects through visual recognition, and organize this information into the domain file and problem file of the PDDL planner to plan the action sequence. It is difficult to realize the whole process autonomously, and the perception model has limited capabilities, which makes it difficult to meet the needs of state recognition between objects in complex environments in daily life. Recently, the idea of ​​using large models is an end-to-end idea. Language, visual perception information and robot operation teaching information are used as training inputs of the Transformer network, and the model output is directly the target position and posture of the end of the robot arm. This method is more dependent on the teaching information, and its planning flexibility for complex operation tasks needs to be verified. Another idea is to use the ChatGPT language large model to infer the action sequence to achieve the task, and then combine reinforcement learning to further plan feasible action sequences for robots, such as the SayCan model. Although this idea combines the "action" ability of reinforcement learning with the "speech" ability of large language models, what is applicable to human natural language may not be applicable to robot body language. The present invention focuses on the autonomy, intelligence and flexibility of task planning, and uses the advantages of language and visual large models in task and scene understanding to obtain the initial relationship between scene objects and predict the target relationship between objects in language tasks. By using the perception and prediction of scene relationships, combined with the knowledge graph that stores information related to robot operations, the action sequence suitable for robot operation planning is intelligently obtained, avoiding the inconsistency problem between human language and robot body language. The specific implementation is to automatically generate the input of the PDDL (Planning Domain Definition Language) planner by querying the knowledge graph, and output a sequence of action primitives that can complete the task.

[0004] The invention is characterized in that it can flexibly and autonomously extract relevant information from a knowledge base according to language visual perception, generate an input file for a PDDL planner, and is used to plan action primitive sequences. The method is concise, effective, and highly practical. Summary of the invention

[0005] Aiming at the problem that robots can flexibly convert complex and diverse task instructions in daily life into action sequences, the present invention provides a task sequence intelligent planning method based on a language vision big model and a knowledge graph. The method mainly includes three parts: hierarchical knowledge graph construction, scene relationship perception and prediction model training based on the language vision big model, and PDDL autonomous action sequence planning. Query and extract relevant object and action attributes from the knowledge graph constructed by the hierarchical knowledge graph, which is used to automatically generate the domain file of the PDDL planner. At the same time, based on the scene relationship perception and prediction model training of the language vision big model, the initial and target states of the relevant objects are output, and the problem file of the PDDL planner is automatically generated; the domain file and the problem file drive the PDDL autonomous action sequence to plan the action primitive sequence.

[0006] The knowledge graph constructed by the hierarchical knowledge graph in the present invention mainly includes an object layer and an action layer. The object layer includes the appearance attributes, operability modes and relationship descriptions of objects. The action layer includes multiple action primitives and attribute storage of skills.

[0007] In the scene relationship perception and prediction model training based on the language vision big model, we first obtain the text description of the scene image and task instructions, and input them into the VisualBERT big model pre-trained on the COCO dataset. The context features are extracted through the attention mechanism, and the potential feature connections between the scene objects in the scene image and the task sentences are obtained. After the scene relationship perception and prediction model training of the language vision big model, we finally achieve the extraction of the initial relationship between objects in the scene (described in the form of triplets) and the prediction of the task target state triplets.

[0008] In practical applications, the current language task instructions and scene images are input into the scene relationship perception and prediction model based on the language vision model, and the initial relationship triples and target state triples of the scene objects are output. According to the output information, the relevant object attributes, the relationship attributes between objects, and the relevant action attributes are queried and extracted from the knowledge graph library, and the domain file and problem file of the PDDL planner are automatically generated. Finally, the PDDL planner solves the problem online to obtain the action primitive sequence required to complete the task.

[0009] The main innovations of the present invention are (1) based on the scene relationship perception and prediction of the language vision big model, the big model is used to learn the potential feature connection between the scene image objects and the task statements, so as to describe the initial and target states of the relevant objects in the task in the form of triples; (2) according to the task requirements, the relevant objects and operation actions, skill attribute information are flexibly extracted from the knowledge graph, and the domain file and problem file of the PDDL planner are automatically generated using this information to complete the action planning autonomously. In previous studies, the domain file of the PDDL planner is usually pre-written by programmers offline, with large content and inflexibility, which makes it difficult to apply to the complex, diverse and changeable task instructions in daily life and home life; and the problem file needs to be written in the form of the object target state according to the requirements of each task instruction, which also brings difficulties to the autonomous and flexible application of the PDDL planner. In the present invention, the robot can flexibly call relevant knowledge from the knowledge base according to the task requirements, and complete the planning autonomously, which is suitable for the action sequence planning in the task-changing scenes such as daily home life.

[0010] Therefore, the advantages of the present invention can be summarized as flexible, autonomous and intelligent:

[0011] The scene relationship perception prediction model based on the language vision big model is combined with the knowledge graph to realize the perception and knowledge association of different tasks in complex daily life, which is the key to realize autonomous, flexible and intelligent task planning. The present invention can be applied to the field of robot intelligent planning. It only needs to input human language task instructions and perceive the visual picture of the current scene. The robot can associate relevant information and intelligently plan according to the built-in knowledge graph to obtain a feasible sequence of operation layer action primitives. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is the overall processing flow chart of the present invention;

[0013] Figure 2 It is a scene relationship perception prediction model based on the language vision big model;

[0014] Figure 3 It is a visualization model of hierarchical knowledge graph;

[0015] Figure 4 It is an autonomous, flexible and intelligent application implementation of the PDDL planner. DETAILED DESCRIPTION

[0016] The present invention is further described below in conjunction with the accompanying drawings.

[0017] Figure 1This is the overall processing flow chart of the present invention. The work completed offline in the present invention consists of the construction of the knowledge graph library and the training of the scene relationship perception prediction model based on the language vision large model; in the online application, the task text sentence is input, for example, "give me a glass of water", and the robot obtains the current scene picture through the visual sensor, and the text sentence and the visual scene picture are simultaneously input into the language vision perception model, and the initial state relationship of the objects in the current scene (such as the apple is on the table, the cup is on the table, etc.) and the predicted target state relationship (such as the cup is filled with water, the cup is in the service place) are output.

[0018] Based on the above information, relevant object and action attributes are queried and extracted from the knowledge graph to automatically generate the domain file of the PDDL planner. At the same time, the problem file of the PDDL planner is automatically generated according to the initial and target states of the relevant objects output by the language visual perception model. These two files can drive the PDDL planning action primitive sequence.

[0019] Figure 2 It is a scene relationship perception prediction model based on a large language vision model. In the process of model construction, the scene images and task text descriptions need to be preprocessed first and converted into a data format that the model can understand. For scene images, key visual features are extracted and labeled through convolutional neural networks. At the same time, the task text description is segmented and semantically annotated through natural language processing technology.

[0020] For the task text, it is segmented into vocabulary units (tokens). Assuming the original text is a string T, it is segmented into a series of vocabulary units t1, t2, …, t n . Convert each vocabulary unit t into the corresponding vocabulary index idx(t i ). This process can be expressed as:

[0021]

[0022] These indexes are fed into the VisualBERT model for further processing. For images, the Tokenize process includes feature extraction and encoding. The image P is segmented into a series of regions R1, R2, …, R m Each region R j The convolutional neural network is used to extract features and obtain the feature vector vec(R j ), and encode it into a format suitable for input to VisualBERT:

[0023] Tokenize(P)=[vec(R1),vec(R2),…,vec(R m )] (2)

[0024] The VisualBERT model can effectively integrate and analyze data of image and text types. The model can accurately understand the complex relationships and contexts in the scene, and has good results in scene relationship perception and prediction. This invention combines the language understanding and visual information processing capabilities of the BERT model, and effectively integrates visual and text information through a multi-head self-attention mechanism, which can be expressed by the following formula:

[0025]

[0026] Where Q, K, and V represent query, key, and value respectively. k is the dimension of the key. In this way, the model can focus on the image area that is most relevant to the task text, thereby extracting meaningful contextual features.

[0027] After pre-training is complete, the model weights are frozen to keep the knowledge learned on the COCO dataset. The processed data is then input into two model heads consisting of a softmax and a feed-forward neural network (FFN).

[0028]

[0029] f(x)=max(0,xW1+b1)W2+b2 (5)

[0030] Where W1, W2 and b1, b2 are the weights and biases of the network respectively. This step is mainly used to convert the model output into a probability distribution to classify different elements in the scene.

[0031] Then, fine-tuning training is performed on the full scene relationship graph (PSG) dataset. The PSG dataset contains rich scene relationship information, provides pixel-level annotations for each object in the scene and annotations for complex relationships between objects, and can deeply understand the complex relationships between objects in the scene. It is often used to complete scene content understanding and relationship extraction tasks. Fine-tuning training optimizes the network weights of the first model head, so that the improved VisualBERT large model can extract scene relationship triplets.

[0032] Construct a target state prediction dataset in the PSG dataset format to fine-tune the network weights of the second-layer model head. Take M home scene pictures, annotate each picture with h task instructions, and annotate the target state triplet [s_idx, o_idx, rel_id] corresponding to the task. Take the picture and task description as input, and the target state triplet as output to fine-tune the second-layer model head. The optimized two-layer model head can utilize the visual language context understanding ability of the pre-trained ViasualBERT visual language large model, reduce the resources and time required to train the model from scratch, and complete the extraction of the current state relationship triplet and the prediction of the target state triplet.

[0033] Figure 3 It is a visualization model of a hierarchical knowledge graph. The knowledge graph stores the knowledge required by robots to complete daily operations, mainly including objects in the scene and common operation actions. The object layer contains objects in daily life (such as tables, cups, water, etc.), object attributes (shape, appearance, operability, etc.), and the positional relationship between objects ( Figure 3 The words on the arrow lines between objects indicate the types of positional relationships, such as on, below, in, etc.). The action layer stores various daily operation actions (such as move, grasp, turnOn, etc.) and their attributes; in addition, the operational relationships between actions and objects are interconnected through nodes, indicating the operability of objects. In order to facilitate the subsequent PDDL (Planning Domain Definition Language) planner to automatically generate domain files, the planning attribute description of the action uses the content format of the domain file to describe the prerequisites for the execution of the action and the effect of the action execution (the following is the planning attribute storage format of the action grab). In addition, in order to facilitate the subsequent generation of motion trajectories, the trajectory attributes of the action are features extracted using the dynamic motion primitive method. In actual applications, when the start and end states of the motion are input, the trajectory attributes of the action are called to generate the corresponding motion trajectory.

[0034]

[0035] Figure 4 It is an autonomous, flexible and intelligent application implementation of the PDDL planner. PDDL planning domain definition language is a language that describes artificial intelligence planning problems. It consists of two parts: domain definition and problem definition. The domain definition describes possible actions and their effects. Each action has some prerequisites, and the action can only be executed when the conditions are met; and the effect of each action describes the change in state after the action is executed. The problem definition describes a specific planning problem, including the initial state and the target state. Input these two files to the PDDL planner to get the action primitive sequence output.

[0036] In the present invention, according to the list of related objects appearing in the scene, the action nodes associated with the object nodes are retrieved in the knowledge graph, and the attributes of the action execution preconditions and the effect after the action execution stored in the action nodes are extracted; and the positional relationship information connected between the objects and the object nodes is retrieved to generate the domain file (predicate and action) of the PDDL planner. The initial state and target state obtained in the language visual perception model are used to autonomously generate the PDDL planner problem file.

[0037] After automatically generating the PDDL solution file based on specific tasks and scenarios, the PDDL intelligent planning algorithm constructs the search space of complex planning problems through domain files and problem files, including all possible states and paths that cause state transitions by actions. The planner uses a heuristic search algorithm to explore the search space, estimate the cost of reaching the target state from the current state, solve for the optimal action primitive sequence, and complete task planning.

[0038] By using the action names in this action primitive sequence to retrieve its trajectory attribute features in the knowledge graph, the initial state and target state (described in the form of posture) of the operated object are input into the dynamic motion primitive method to obtain the motion trajectory of the operation. The robot can complete the operation task by executing the motion trajectory in sequence.

[0039] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that modifying the technical solutions described in the aforementioned embodiments, or replacing some or all of the technical features therein by equivalents, does not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An intelligent planning method for task sequences based on a language vision big model and knowledge graph, characterized in that: The method includes: hierarchical knowledge graph construction; scene relationship perception and prediction model training based on language vision big model; PDDL autonomous action sequence planning; querying and extracting relevant object and action attributes from the knowledge graph constructed by the hierarchical knowledge graph, which is used to automatically generate a domain file of the PDDL planner, and at the same time, scene relationship perception and prediction model training based on the language vision big model, and outputting the initial and target states of the relevant objects, and automatically generating a problem file of the PDDL planner; the domain file and the problem file drive the PDDL autonomous action sequence planning action primitive sequence; In the training process of the scene relationship perception and prediction model based on the language vision large model, the scene images and task text descriptions are first preprocessed and converted into a data format that the model can understand; for the scene images, key visual features are extracted and labeled through a convolutional neural network; the task text description is segmented and semantically annotated through natural language processing technology; For the task text, it is divided into vocabulary units; the original task text is a string T, which is divided into a series of vocabulary units t1, t2, ..., t i,…,t n ; Each vocabulary unit t i Convert to the corresponding vocabulary index idx(t i ); expressed as: Tokenize(T)=[idx(t1),idx(t2),…,idx(t n )] (1) The index is sent to the VisualBERT model for processing; for images, the Tokenize process includes feature extraction and encoding; the image P is divided into a series of regions R1, R2, …, R j ,…,R m ; Each region R j The convolutional neural network is used to extract features and obtain the feature vector vec(R j ), and encode it into a format suitable for input to VisualBERT: Tokenize(P)=[vec(R1),vec(R2),…vec(R j ),…,thing(R m )] (2) The VisualBERT model combines the language understanding and visual information processing capabilities of the BERT model, and effectively integrates visual and textual information through a multi-head self-attention mechanism, which is expressed as the following formula: Where Q, K, and V represent query, key, and value, respectively. k is the dimension of the key; After pre-training is completed, the model weights are frozen to keep the knowledge learned on the COCO dataset; the processed data is then input into two model heads consisting of a softmax and a feed-forward neural network ffn; f(x)=max(0,xW1+b1)W2+b2 (5) Where W1, W2 and b1, b2 are the weights and biases of the network respectively; Fine-tune the PSG dataset of the full scene relationship graph. Optimize the weights of the first model head network through fine-tuning training, so that the improved VisualBERT large model can extract scene relationship triplets. Construct a target state prediction dataset in the PSG dataset format to fine-tune the network weights of the second-layer model head; take M home scene pictures, annotate each picture with h task instructions, and annotate the target state triplet [s_idx, o_idx, rel_id] corresponding to the task; use the picture and task description as input and the target state triplet as output to fine-tune the second-layer model head; the optimized two-layer model head uses the visual language context understanding ability of the pre-trained ViasualBERT visual language large model to reduce the resources and time required to train the model from scratch, and complete the extraction of the current state relationship triplet and the prediction of the target state triplet; According to the list of related objects appearing in the home scene, the action nodes associated with the object nodes are retrieved in the knowledge graph, and the attributes of the action execution preconditions and the effect after the action execution stored in the action nodes are extracted; and the position relationship information between the objects and the object nodes is retrieved to generate the domain file of the PDDL planner, namely the predicate and action; the initial state and target state obtained from the scene relationship perception and prediction model of the language vision large model are used to autonomously generate the PDDL planner problem file; After automatically generating the PDDL solution file based on specific tasks and scenarios, the PDDL intelligent planning algorithm constructs a search space for complex planning problems through domain files and problem files, including all possible states and paths where actions cause state transitions; the planner uses a heuristic search algorithm to explore the search space, estimate the cost of reaching the target state from the current state, and solve the optimal action primitive sequence to complete task planning; the action names in the optimal action primitive sequence are used to retrieve their trajectory attribute features in the knowledge graph, and the initial state and target state or posture form of the operated object are input into the dynamic motion primitive method to obtain the motion trajectory of the operation; the robot executes the motion trajectory in sequence to complete the operation task.

2. According to claim 1, the task sequence intelligent planning method based on language vision large model and knowledge graph is characterized in that: The knowledge graph constructed by the hierarchical knowledge graph includes an object layer and an action layer; wherein the object layer contains the appearance attributes, operability modes and relationship descriptions of the objects; the action layer includes multiple action primitives and attribute storage of skills.

3. According to claim 1, the task sequence intelligent planning method based on language vision large model and knowledge graph is characterized in that: In the scene relationship perception and prediction model training based on the language vision big model, we first obtain the text description of the scene image and task instructions, and input them into the VisualBERT big model pre-trained on the COCO dataset. The context features are extracted through the attention mechanism, and the potential feature connections between the scene objects in the scene image and the task sentences are obtained. After the scene relationship perception and prediction model training of the language vision big model, the initial relationship between objects in the scene is finally extracted in the form of triples and described in the form of triples and the task target state triples are predicted.

4. The task sequence intelligent planning method based on language vision large model and knowledge graph according to claim 1 is characterized in that: In practical applications, the current language task instructions and scene images are input into the scene relationship perception and prediction model based on the language vision big model, and the initial relationship triples and target state triples of scene objects are output; based on the output information, the relevant object attributes, the relationship attributes between objects and the relevant action attributes are queried and extracted from the knowledge graph library, and used to automatically generate the domain file and problem file of the PDDL planner; finally, the PDDL planner is used to solve online to obtain the action primitive sequence required to complete the task.

5. The task sequence intelligent planning method based on language vision large model and knowledge graph according to claim 1 is characterized in that: Query and extract relevant object and action attributes from the knowledge graph to automatically generate the domain file of the PDDL planner. At the same time, according to the initial and target states of the relevant objects output by the language visual perception model, the problem file of the PDDL planner is automatically generated. The domain file and problem file drive the PDDL planning action primitive sequence.

6. The task sequence intelligent planning method based on language vision large model and knowledge graph according to claim 2 is characterized in that: The knowledge graph stores the knowledge that robots need to complete daily operations, including objects in the scene and common operating actions. The object layer includes objects in daily life, their properties, and the positional relationships between them; The action layer stores various daily operation actions and their attributes; The operational relationship between actions and objects is connected to each other through nodes, which represent the operability of objects.

Citation Information

Patent Citations

  • Low-level robot task planning method based on multi-modal knowledge graph

    CN113433941A

  • Triple information extraction method, apparatus, and device, and computer-readable storage medium

    WO2022116417A1