Action Prediction Method, Terminal, and Storage Medium Based on Key Points and Imitation Learning

Through a method based on key points and imitation learning, combined with coding, world model and trajectory prediction model, the decision accuracy and robustness in robotic arm tasks are improved, and the problem of insufficient scenario understanding of multi-task imitation learning decision model in complex environments is solved.

CN119648747BActive Publication Date: 2025-07-25BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510162679.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-07-25
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

The multi-task imitation learning decision model applied to robotic arm tasks in the prior art has poor understanding of the environment scenarios, resulting in low decision-making accuracy.

Method used

Using a method based on key points and imitation learning, we use task information and observation data to generate predictive execution actions using coding modules, world models, trajectory prediction models and attention models, and learn scene representations with environmental dynamics prediction models, and predict key points motion trajectories through the trained trajectory prediction model to assist in action decision-making.

Benefits of technology

It improves the decision-making accuracy of intelligent mechanical devices and has good robustness. Even if the object is partially blocked or the environment changes, the key point trajectory can still provide effective information, and the neural network can better fit the moving movement of the mechanical device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119648747B_ABST
    Figure CN119648747B_ABST
Patent Text Reader

Abstract

The present invention discloses a motion prediction method, a terminal, and a storage medium based on key points and imitation learning, which relates to the field of embodied intelligence technology. The present invention combines the idea of the world model with imitation learning, learns a scene representation that better conforms to the environmental context through an environmental dynamics prediction model, and predicts the key point motion trajectory through a trained trajectory prediction model to assist the action head module in making action decisions. The present invention uses the key point motion trajectory as auxiliary information, which can effectively improve the decision-making accuracy of the intelligent agent. First, the key points are not affected by environmental changes such as lighting. Even if the object is partially occluded, the key point trajectory of the remaining part can still provide the same information, so it has good robustness. Second, there is a definite mapping relationship between the motion trajectory of the key points and the motion trajectory of the intelligent mechanical device in three-dimensional space, so that the neural network can better fit the movement actions of the intelligent mechanical device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of embodied intelligence technology, and in particular, to an action prediction method, a terminal, and a storage medium based on key points and imitation learning. Background Art

[0002] Imitation learning aims to learn complex tasks by imitating the behaviors of humans or other intelligent agents, and is used to understand and construct intelligent systems with autonomous behavior capabilities. Multi-task imitation learning is a method of training an intelligent agent to master multiple skills simultaneously by imitating the expert demonstrations of multiple different tasks. Multi-task imitation learning requires the model to be able to share knowledge and capabilities between different tasks, improve the generalization ability and efficiency of the model, and is applicable to complex scenarios such as robot control and natural language processing.

[0003] For robotic arm tasks that use visual information as input information, due to the complexity of visual perception and the diversity of objects, when the robotic arm operates in a complex and diverse environment, visual perception faces many challenges, which in turn affect the grasping accuracy. For example, the object is occluded, making it difficult for the visual system to obtain complete object information; the change in light makes it difficult for the visual system to accurately identify the object boundary and texture. Scene understanding and multi-task processing also pose challenges to the model algorithm. When the robotic arm faces a complex working environment, it not only needs to identify a single object, but also needs to understand the context of the entire scene.

[0004] Currently, the multi-task imitation learning decision model applied to robotic arm tasks has poor scene understanding of the environment, resulting in a low decision-making accuracy.

[0005] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to provide an action prediction method, a terminal, and a storage medium based on key points and imitation learning for the above-mentioned defects of the existing technology, aiming to solve the problem that the multi-task imitation learning decision model applied to robotic arm tasks in the existing technology has poor scene understanding of the environment, resulting in a low decision-making accuracy.

[0007] The technical solution adopted by the present invention to solve the problem is as follows:

[0008] In a first aspect, an embodiment of the present invention provides an action prediction method based on key points and imitation learning, the method comprising:

[0009] Obtain task information and observation data of an intelligent mechanical device;

[0010] Input the task information and the observation data into a trained multi-task imitation learning decision model to obtain a predicted execution action of the intelligent mechanical device;

[0011] Among them, the multi-task imitation learning decision-making model includes:

[0012] An encoding module, configured to encode the task information and the observation data, and generate a number of semantic units according to the encoded data;

[0013] A world model, configured to generate a predicted state sequence through a decoder and an environmental dynamics prediction model according to each of the semantic units;

[0014] A trajectory prediction model, configured to generate a predicted key point trajectory according to the predicted state sequence;

[0015] An attention model, configured to generate a hidden vector according to each of the semantic units; wherein, the hidden vector is used to reflect the overall information of the environment and the task;

[0016] An action head module, configured to generate a predicted execution action of the intelligent mechanical device according to the predicted key point trajectory and the hidden vector.

[0017] In one implementation, the observation data includes image observation data of a number of camera views and state observation data of the intelligent mechanical device.

[0018] In one implementation, encoding the task information and the observation data, and generating a number of semantic units according to the encoded data, includes:

[0019] Obtaining the semantic unit corresponding to the task information through a language encoder;

[0020] Obtaining the semantic unit corresponding to the image observation data through a visual encoder and a SimNorm layer; wherein, the SimNorm layer is used to divide the encoded vector into a number of segments with a preset length, perform a softmax operation separately in each segment, and splice them back together;

[0021] Obtaining the semantic unit corresponding to the state observation data through a state encoder.

[0022] In one implementation, generating a predicted state sequence through a decoder and an environmental dynamics prediction model according to each of the semantic units, includes:

[0023] Generating reconstructed observation data through the decoder based on all the semantic units;

[0024] Generating predicted states for a number of future time steps in a loop through the environmental dynamics prediction model based on the reconstructed observation data, to obtain the predicted state sequence.

[0025] In one implementation, the training method of the multi-task imitation learning decision-making model includes:

[0026] Collect expert trajectories through the expert policy module, and establish an expert dataset according to the expert trajectories; wherein, each piece of expert data includes historical observation data, corresponding expert policy actions, and corresponding historical task information.

[0027] Take the encoding module, the world model, the attention model, and the action head module as the policy model.

[0028] Train the policy model and the trajectory prediction model through the expert dataset, and update the model parameters of the policy model and the trajectory prediction model respectively.

[0029] In one implementation, the training method of the policy model includes:

[0030] Calculate the loss value of the world model and the loss value of the action head module through the expert dataset; wherein, the loss value of the world model is determined based on the reconstruction observation loss value of the decoder and the prediction loss value of the environmental dynamics prediction model.

[0031] Determine the loss value of the policy model according to the loss values corresponding to the world model and the action head module respectively.

[0032] Update the model parameters of the policy model according to the loss value of the policy model.

[0033] In one implementation, the training method of the trajectory prediction model includes:

[0034] Judge whether the current training step number meets the training interval of the trajectory prediction model.

[0035] If it meets, obtain the predicted state sequence output by the environmental dynamics prediction model, and generate a training predicted key point trajectory through the trajectory prediction model according to this predicted state sequence.

[0036] Generate a reference predicted key point trajectory through the point tracking model based on the historical observation data of the corresponding number of frames in the expert dataset.

[0037] Calculate the loss value of the trajectory prediction model according to the training predicted key point trajectory and the reference predicted key point trajectory; update the model parameters of the trajectory prediction model according to the loss value of the trajectory prediction model.

[0038] In one implementation, the historical observation data of the corresponding number of frames in the expert dataset includes real image observation data of a continuous number of frames; generating a reference predicted key point trajectory through the point tracking model based on the historical observation data of the corresponding number of frames in the expert dataset includes:

[0039] Sample a number of feature points based on the real image observation data of the first frame;

[0040] Based on the real image observation data of each frame and the positions of each of the feature points, generate trajectory change data of each of the feature points in each frame through a point tracking model;

[0041] Calculate the change amplitude of each of the feature points according to each of the trajectory change data;

[0042] Select a number of target feature points according to each of the change amplitudes, and generate a reference prediction key point trajectory according to the trajectories of each of the target feature points.

[0043] In a second aspect, an embodiment of the present invention further provides a terminal, the terminal includes a memory and at least one processor; the memory stores a program; the program contains instructions for executing the action prediction method based on key points and imitation learning as described in any one of the above; the processor is used to execute the program.

[0044] In a third aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor to implement the steps of the action prediction method based on key points and imitation learning as described in any one of the above.

[0045] Advantages of the present invention: The embodiment of the present invention combines the idea of a world model with imitation learning, learns a scene representation that better conforms to the environmental context through an environmental dynamics prediction model, and predicts the key point motion trajectory through a trained trajectory prediction model to assist the action head module in making action decisions. The present invention uses the key point motion trajectory as auxiliary information, which can effectively improve the decision-making accuracy of the intelligent agent. First, the key points are not affected by environmental changes such as light. Even if the object is partially occluded, the key point trajectories of the remaining parts can still provide the same information, so it has good robustness. Second, there is a definite mapping relationship between the motion trajectory of the key points and the motion trajectory of the intelligent mechanical device in three-dimensional space, so it can enable the neural network to better fit the movement actions of the intelligent mechanical device. Description of the Drawings

[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] Figure 1It is a schematic flowchart of the action prediction method based on key points and imitation learning provided by an embodiment of the present invention.

[0048] Figure 2 It is a schematic diagram of the overall algorithm architecture of the action prediction method based on key points and imitation learning provided by an embodiment of the present invention.

[0049] Figure 3 It is a schematic logic diagram of key point trajectory prediction provided by an embodiment of the present invention.

[0050] Figure 4 It is a schematic block diagram of the principle of the terminal provided by an embodiment of the present invention. Detailed implementation manners

[0051] The present invention discloses an action prediction method, a terminal, and a storage medium based on key points and imitation learning. To make the objectives, technical solutions, and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.

[0052] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0053] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the technical field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0054] In view of the above defects of the prior art, the present invention provides a motion prediction method based on key points and imitation learning. The method includes obtaining task information and observation data of an intelligent mechanical device; inputting the task information and the observation data into a trained multi-task imitation learning decision model to obtain a predicted execution motion of the intelligent mechanical device. The multi-task imitation learning decision model includes: an encoding module for encoding the task information and the observation data and generating a plurality of semantic units according to the encoded data; a world model for generating a predicted state sequence through a decoder and an environmental dynamics prediction model according to each semantic unit; a trajectory prediction model for generating a predicted key point trajectory according to the predicted state sequence; an attention model for generating a hidden vector according to each semantic unit, where the hidden vector is used to reflect the overall information of the environment and the task; and an action head module for generating a predicted execution motion of the intelligent mechanical device according to the predicted key point trajectory and the hidden vector. The present invention combines the idea of the world model with imitation learning, learns a scene representation that better conforms to the environmental context through the environmental dynamics prediction model, and predicts the key point motion trajectory through the trained trajectory prediction model to assist the action head module in making action decisions. The present invention uses the key point motion trajectory as auxiliary information, which can effectively improve the decision-making accuracy of the intelligent agent. First, the key points are not affected by environmental changes such as lighting. Even if the object is partially occluded, the key point trajectory of the remaining part can still provide the same information, so it has good robustness. Second, there is a definite mapping relationship between the key point motion trajectory and the motion trajectory of the intelligent mechanical device in three-dimensional space, so that the neural network can better fit the movement actions of the intelligent mechanical device.

[0055] As Figure 1 shown, the method specifically includes the following steps:

[0056] Step S100: Obtain task information and observation data of the intelligent mechanical device.

[0057] Step S200: Input the task information and the observation data into the trained multi-task imitation learning decision model to obtain a predicted execution motion of the intelligent mechanical device.

[0058] Specifically, the task information can also be referred to as language input or task description, which is used to reflect the current task of the intelligent mechanical device, such as task type and task objective. The intelligent mechanical device can be a robotic arm. The observation data of the intelligent mechanical device includes image observation data from several camera views and its own state observation data. The task information and the observation data are the basic data for the intelligent mechanical device to make autonomous decision-making actions. Inputting the task information and the observation data into a pre-constructed and trained multi-task imitation learning decision model, the multi-task imitation learning decision model automatically decides the next action of the intelligent mechanical device.

[0059] Further, the multi-task imitation learning decision model includes:

[0060] An encoding module for encoding the task information and the observation data, and generating a number of semantic units according to the encoded data;

[0061] A world model for generating a predicted state sequence according to each of the semantic units through a decoder and an environmental dynamics prediction model;

[0062] A trajectory prediction model for generating a predicted key point trajectory according to the predicted state sequence;

[0063] An attention model for generating a hidden vector according to each of the semantic units; wherein, the hidden vector is used to reflect the overall information of the environment and the task;

[0064] An action head module for generating a predicted execution action of the intelligent mechanical device according to the predicted key point trajectory and the hidden vector.

[0065] Specifically, as Figure 2 shown, the multi-task imitation learning decision model mainly includes five functional modules, an encoding module, a world model (i.e., the dynamics model in Figure 2 ), a trajectory prediction model, an attention model, and an action head module. Among them, the role of the encoding module is to encode both the observation data and the task information into vectors in the latent space, so as to obtain a number of semantic units (Tokens). The world model includes a decoder and an environmental dynamics model, and its role is to imagine the future state trajectory at each time step, so as to obtain a predicted state sequence. The trajectory prediction model is pre-trained by supervised learning and can accurately map the predicted state sequence to the motion trajectory of the key points, so as to obtain a predicted key point trajectory, and use this predicted key point trajectory as auxiliary information for subsequent action decisions. The attention model (Transformer) is used to generate a hidden vector that can reflect the overall information of the environment and the task according to all semantic units, and input it into the action head module for action decision-making. The action head module combines the hidden vector output by the attention model and the predicted key point trajectory for analysis, and decides the action to be executed by the intelligent mechanical device, that is, obtains a predicted execution action, and sends it to the intelligent mechanical device for execution. In this embodiment, the trained multi-task imitation learning decision model can be deployed to the corresponding scenario, provide observation data at each time step, and the model outputs the action to be executed by the intelligent mechanical device.

[0066] In one implementation, encoding the task information and the observation data, and generating a number of semantic units, includes:

[0067] Obtain the semantic unit corresponding to the task information through a language encoder;

[0068] Obtain the semantic unit corresponding to the image observation data through a visual encoder and a SimNorm layer; wherein, the SimNorm layer is used to divide the encoded vector into several segments of a preset length, perform a softmax operation separately in each segment, and then splice them back together;

[0069] Obtain the semantic unit corresponding to the state observation data through a state encoder.

[0070] Specifically, in this embodiment, it is necessary to map the task information and the observation data into semantic units (Tokens) that can be used by the Transformer. The image observation data passes through the visual encoder and the SimNorm layer to generate visual Tokens, the state observation data (or called the state vector) passes through the state encoder to generate state Tokens, and the task information passes through the language encoder to generate language Tokens. The three will subsequently be jointly input into the attention model (Transformer). Among them, the language encoder can adopt a pre-trained lightweight sentence embedding model to encode different task descriptions into embedding vectors of a fixed dimension.

[0071] For example, the observation data at each step , , respectively represent the image observations from different camera perspectives and the body state vector observations of the robotic arm. For each task, the corresponding task description is encoded by a pre-trained model, denoted as Lang. The input data at each time step are respectively mapped into a number of vector encodings through the language encoder and the visual encoder. The vector encodings obtained by the visual encoder are then passed through the SimNorm layer and, together with other encodings, obtain a number of Tokens that can be input into the Transformer.

[0072] In one implementation manner, according to each of the semantic units, a predicted state sequence is generated through a decoder and an environmental dynamics prediction model, including:

[0073] Based on all the semantic units, the reconstructed observation data is generated through the decoder;

[0074] Based on the reconstructed observation data, the environmental dynamics prediction model cyclically generates the predicted states at several future time steps to obtain the predicted state sequence.

[0075] Specifically, the decoder and the environmental dynamics prediction model in this embodiment have been pre-trained and have learned the complex mapping relationship between the input and output. Therefore, in actual applications, the decoder can learn to reconstruct the observations based on the previously generated semantic units (Tokens) and output relatively accurate reconstructed observation data. As Figure 3 shown, the environmental dynamics prediction model obtains the current state based on the reconstructed observation data output by the decoder , and imagines future states backward , etc. (s represents the state, and t represents the time step), so as to obtain a predicted state sequence.

[0076] In one implementation, the training method of the multi-task imitation learning decision model includes:

[0077] Collect expert trajectories through the expert policy module, and establish an expert data set according to the expert trajectories; wherein, each expert data includes historical observation data, corresponding expert policy actions, and corresponding historical task information;

[0078] Regard the encoding module, the world model, the attention model, and the action head module as the policy model;

[0079] Train the policy model and the trajectory prediction model through the expert data set, and update the model parameters of the policy model and the trajectory prediction model respectively.

[0080] Specifically, this embodiment will train the multi-task imitation learning decision model according to the expert data and through the behavior cloning algorithm. Before training, first collect the expert trajectories of the intelligent mechanical device (such as a robotic arm) in completing various tasks through the expert policy, and record and save the expert trajectories in the form of observation-action pairs. Each observation-action pair is an expert data, which contains historical observation data, expert policy actions, and historical task information with corresponding relationships. Among them, the historical observation data and historical task information can be used as the input data for subsequent model training, and the expert policy actions can be used as the true labels for subsequent model training to evaluate the model performance. During training, use the previously generated expert data set to train the multi-task imitation learning decision model, so that the multi-task imitation learning decision model can learn multiple tasks by imitating expert behaviors. Regard the encoding module, the world model including the decoder and the environmental dynamics prediction model, the attention model, and the action head module as a policy model. When parameter optimization is required, optimize the parameters of the policy model and the trajectory prediction model respectively.

[0081] For example, the training process of the multi-task imitation learning decision model includes: initializing the parameters of the Transformer backbone, the action head module, the world model, and the trajectory prediction model. Among them, the world model includes an environmental dynamics prediction model and a decoder for learning to reconstruct observations. Denote the parameters of the policy model as , and the parameters of the trajectory prediction model as . Read the collected expert trajectories into the memory Buffer and set the hyperparameters. The hyperparameters include the batch size B and the setting parameters of the SimNorm operation (performed in the aforementioned SimNorm layer) after the image encoder encodes the observations. The SimNorm operation divides the encoded vector into L segments of length d, performs a softmax operation separately in each segment, and then reassembles them. Therefore, its setting parameters include: the division length d, the length H of the policy training imaginary trajectory (i.e., the number of time steps predicted backward during the training of the environmental dynamics prediction model), the number of key points used , and the length T of the key point trajectory time (i.e., the length of the key point trajectory), where T is less than or equal to H. It should be noted that the SimNorm operation is only applied to the visual encoder and not to the language encoder and the state encoder. Train the multi-task imitation learning decision model using the expert dataset. If the current training step has reached the maximum number of steps , then end the training; otherwise, randomly select B groups of data from the Buffer and continue to train and update the current multi-task imitation learning decision model.

[0082] In one implementation, the training method of the policy model includes:

[0083] Calculate the loss value of the world model and the loss value of the action head module using the expert dataset; among them, the loss value of the world model is determined based on the reconstruction observation loss value of the decoder and the prediction loss value of the environmental dynamics prediction model;

[0084] Determine the loss value of the policy model according to the loss values corresponding to the world model and the action head module respectively;

[0085] Update the model parameters of the policy model according to the loss value of the policy model.

[0086] Specifically, the loss value of the policy model is composed of the loss value of the world model and the loss value of the action head module. The numerical size of the loss value can reflect the gap between the prediction result and the true result. The loss value of the world model is composed of the reconstruction observation loss value of the decoder and the prediction loss value of the environmental dynamics prediction model. The reconstruction observation loss value can reflect the gap between the reconstructed observation data output by the decoder and the true observation data, while the prediction loss value can reflect the gap between the predicted state of the environmental dynamics prediction model and the future true state. During the training process, the model parameters are updated guided by the loss value of the policy model.

[0087] For example, for the world model, the decoder (Decoder) and the environmental dynamics prediction model (Dynamic) are trained separately. The calculation formula for the loss value of the world model is:

[0088] ;

[0089] In the formula, represents the encoder, represents the loss of the reconstructed observation, represents the prediction loss of the state sequence, and the state sequence can also be called the state trajectory. First, the Token obtained in the above steps is decoded by the Decoder to reconstruct the observation, and the reconstruction observation loss value is calculated with the true value of the observation, and this item is used as the auxiliary loss for learning the environmental dynamics because the rich supervision signals provided by the reconstructed image can enable the network to effectively learn the environmental representation. Second, the future state (the state can also be represented by the Token) is predicted cyclically to obtain the state trajectory , that is, a series of states, and the true label corresponding to the state trajectory is generated using the hidden state encoded by the future true observation, enabling the environmental dynamics prediction model to learn to accurately predict the future state.

[0090] For the attention model (Transformer), all the aforementioned semantic units (including language Tokens, image Tokens, and state Tokens) are input into the Transformer, and a hidden vector containing the overall environmental and task information is obtained through the attention mechanism.

[0091] For the trajectory prediction model, the predicted key point trajectory output by the trajectory prediction model is denoted as .

[0092] For the action head module, the action head module receives the hidden vector and the predicted key point trajectory denoted as as these two inputs, and outputs the predicted action of the intelligent mechanical device (robotic arm) as the action to be executed by the machine:

[0093] .

[0094] Compare the predicted action with the corresponding expert policy action in the expert data, and calculate the loss value of the action head module :

[0095] ;

[0096] where represents the expert data; represents the task objective, i.e., the above-mentioned Lang; represents the current state.

[0097] In summary, the loss function of the policy model:

[0098] ;

[0099] where is the weight of the action loss term.

[0100] In one implementation, the training method of the trajectory prediction model includes:

[0101] Determine whether the current training step satisfies the training interval of the trajectory prediction model;

[0102] If satisfied, obtain the predicted state sequence output by the environmental dynamics prediction model, and generate a training predicted key point trajectory through the trajectory prediction model according to the predicted state sequence;

[0103] Based on the historical observation data corresponding to the number of frames in the expert dataset, generate a reference predicted key point trajectory through the point tracking model;

[0104] Calculate the loss value of the trajectory prediction model according to the training predicted key point trajectory and the reference predicted key point trajectory;

[0105] Update the model parameters of the trajectory prediction model according to the loss value of the trajectory prediction model.

[0106] Specifically, for the model parameter update process of the trajectory prediction model, first, the current training step is judged. When the training interval of the trajectory prediction model is satisfied, the model parameters of the trajectory prediction model are updated. Using the environmental dynamics prediction model trained in the foregoing steps, the predicted states at T time steps are cyclically generated. The states of this time series are the predicted state sequences, which are used as the input data of the trajectory prediction model. The trajectory prediction model outputs the training predicted key point trajectory based on this input data. At the same time, the Co-tracker is used to track the key points of the real image observation data in the next T frames to obtain the reference predicted key point trajectory. Among them, the Co-tracker is a pre-trained tracking model. Taking the reference predicted key point trajectory as the true label (i.e., the supervision signal) of the trajectory prediction model, by comparing the reference predicted key point trajectory and the training predicted key point trajectory output by the trajectory prediction model, the performance of the trajectory prediction model can be evaluated, and the loss value of the trajectory prediction model can be calculated, and the model parameters are updated guided by this loss value.

[0107] For example, as Figure 3 shown, the current state is , and the environmental dynamics prediction model is used to imagine the future states backward , , etc. These are input into the trajectory prediction model trained by the Co-tracker, and the trajectory prediction model outputs the coordinate trajectories of the key points in the observed image, that is, the predicted key point trajectory is obtained. Among them, is a custom hyperparameter.

[0108] In one implementation, the historical observation data corresponding to the number of frames in the expert dataset includes the real image observation data of several consecutive frames; based on the historical observation data corresponding to the number of frames in the expert dataset, generating the reference predicted key point trajectory through the Co-tracker includes:

[0109] Sampling a number of feature points according to the real image observation data of the first frame;

[0110] Based on the real image observation data of each frame and the positions of the feature points, generating the trajectory change data of each feature point in each frame through the Co-tracker;

[0111] Calculating the change amplitude of each feature point according to the trajectory change data;

[0112] Selecting a number of target feature points according to the change amplitudes, and generating the reference predicted key point trajectory according to the trajectories of the target feature points.

[0113] Specifically, the specific process of using the Co-tracker to track key points in the real image observation data for the next T frames is as follows: First, sample feature points (or initial points, sampling points) in the first frame image, input the continuous T-frame images and the positions of the feature points into the point tracking model, and obtain the trajectory changes of the feature points in each frame. Second, calculate the total length of the trajectory changes of all feature points in T time steps, and rank all feature points according to the change amplitude. Finally, select several feature points with the largest change amplitude as key points , that is, obtain a set of target feature points, and determine the reference prediction key point trajectory according to the trajectories corresponding to these target feature points, as the true label (or prediction label) of the trajectory prediction model.

[0114] The advantages of the present invention are as follows:

[0115] 1. The present invention uses a pre-trained lightweight sentence embedding model to encode different task descriptions into embedding vectors of a fixed dimension. And through the multi-task imitation learning decision model assisted by key points, it effectively fits the expert strategy, and the multi-task imitation learning decision model can process various inputs such as language and images, and performs well in terms of convergence speed, task success rate, etc.

[0116] 2. While learning the expert strategy through the behavior cloning algorithm, the present invention constructs a world model to realize the trajectory prediction of future states, can predict the movement trajectory of key points in the observation at each time step, and uses this information as the auxiliary input of the action head module for the final decision.

[0117] 3. By adding a world model, the present invention can standardize the scene representation, combine dynamic trajectory-assisted decision-making information, and can better realize the autonomous decision-making of the robotic arm in a multi-task scenario.

[0118] Based on the above embodiments, the present invention also provides a terminal, and its principle block diagram can be as Figure 4 shown. The terminal includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes the action prediction method based on key points and imitation learning. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.

[0119] Those skilled in the art can understand, Figure 4The principle block diagram shown only shows the block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0120] In one implementation, more than one program is stored in the memory of the terminal, and is configured to be executed by more than one processor. The more than one program includes instructions for performing an action prediction method based on key points and imitation learning.

[0121] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0122] In summary, the present invention discloses an action prediction method, a terminal, and a storage medium based on key points and imitation learning, which relate to the field of embodied intelligence technology. The method includes obtaining task information and observation data of an intelligent mechanical device; inputting the task information and the observation data into a trained multi-task imitation learning decision model to obtain a predicted execution action of the intelligent mechanical device; wherein, the multi-task imitation learning decision model includes: an encoding module for encoding the task information and the observation data and generating a plurality of semantic units according to the encoded data; a world model for generating a predicted state sequence through a decoder and an environmental dynamics prediction model according to each of the semantic units; a trajectory prediction model for generating a predicted key point trajectory according to the predicted state sequence; an attention model for generating a hidden vector according to each of the semantic units; wherein, the hidden vector is used to reflect the overall information of the environment and the task; and an action head module for generating the predicted execution action of the intelligent mechanical device according to the predicted key point trajectory and the hidden vector. The present invention combines the idea of the world model with imitation learning, learns a scene representation that better conforms to the environmental context through the environmental dynamics prediction model, and predicts the key point movement trajectory through the trained trajectory prediction model to assist the action head module in making action decisions. The present invention uses the key point movement trajectory as auxiliary information, which can effectively improve the decision-making accuracy of the intelligent agent. First, the key points are not affected by environmental changes such as light. Even if the object is partially occluded, the key point trajectory of the remaining part can still provide the same information, so it has good robustness. Second, there is a definite mapping relationship between the movement trajectory of the key points and the movement trajectory of the intelligent mechanical device in three-dimensional space, so that the neural network can better fit the movement action of the intelligent mechanical device.

[0123] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.

Claims

1. A motion prediction method based on key points and imitation learning, characterized in that, The method includes: Obtaining task information and observation data of the intelligent mechanical device; the observation data includes image observation data of a number of camera views and status observation data of the intelligent mechanical device; Inputting the task information and the observation data into a trained multi-task imitation learning decision model to obtain a predicted execution action of the intelligent mechanical device; Among them, the multi-task imitation learning decision model includes: An encoding module for encoding the task information and the observation data, and generating a number of semantic units according to the encoded data, including: obtaining the semantic unit corresponding to the task information through a language encoder; obtaining the semantic unit corresponding to the image observation data through a visual encoder and a SimNorm layer; where the SimNorm layer is used to divide the encoded vector into a number of segments of a preset length, perform a softmax operation separately in each segment, and splice them back; obtaining the semantic unit corresponding to the status observation data through a status encoder; A world model for generating a predicted state sequence through a decoder and an environmental dynamics prediction model according to each semantic unit, including: generating reconstructed observation data through the decoder based on all the semantic units; obtaining the current state through the environmental dynamics prediction model based on the reconstructed observation data, and imagining future states backward, and cyclically generating predicted states at a number of future time steps to obtain the predicted state sequence; A trajectory prediction model for generating a predicted key point trajectory according to the predicted state sequence; An attention model for generating a hidden vector according to each semantic unit; where the hidden vector is used to reflect the overall information of the environment and the task; An action head module for generating a predicted execution action of the intelligent mechanical device according to the predicted key point trajectory and the hidden vector; The training method of the multi-task imitation learning decision model includes: Collecting expert trajectories through an expert policy module, and establishing an expert data set according to the expert trajectories, so that the multi-task imitation learning decision model learns multiple tasks by imitating expert behaviors; where each expert data includes historical observation data, corresponding expert policy actions, and corresponding historical task information; Regarding the encoding module, the world model, the attention model, and the action head module as a policy model; Training the policy model and the trajectory prediction model through the expert data set, and updating the model parameters of the policy model and the trajectory prediction model respectively.

2. The action prediction method based on key points and imitation learning according to claim 1, wherein The training method of the policy model includes: Calculating the loss value of the world model and the loss value of the action head module through the expert data set; where the loss value of the world model is determined based on the reconstruction observation loss value of the decoder and the prediction loss value of the environmental dynamics prediction model; Determining the loss value of the policy model according to the loss values corresponding to the world model and the action head module respectively; Updating the model parameters of the policy model according to the loss value of the policy model.

3. The action prediction method based on key points and imitation learning according to claim 1, characterized in that, The training method of the trajectory prediction model includes: Judging whether the current training step number satisfies the training interval of the trajectory prediction model; If the condition is satisfied, obtain the predicted state sequence output by the environmental dynamics prediction model, and generate a training predicted key point trajectory through the trajectory prediction model according to the predicted state sequence; Based on the historical observation data corresponding to the number of frames in the expert dataset, generate a reference predicted key point trajectory through the point tracking model; According to the training predicted key point trajectory and the reference predicted key point trajectory, calculate the loss value of the trajectory prediction model; update the model parameters of the trajectory prediction model according to the loss value of the trajectory prediction model.

4. The action prediction method based on key points and imitation learning according to claim 3, wherein, The historical observation data corresponding to the number of frames in the expert dataset includes real image observation data of several consecutive frames; Based on the historical observation data corresponding to the number of frames in the expert dataset, generating a reference predicted key point trajectory through the point tracking model includes: Sample a number of feature points according to the real image observation data of the first frame; Based on the real image observation data of each frame and the positions of the feature points, generate trajectory change data of each feature point in each frame through the point tracking model; According to the trajectory change data, calculate the change amplitude of each feature point; Select a number of target feature points according to the change amplitudes, and generate a reference predicted key point trajectory according to the trajectories of the target feature points.

5. A terminal, characterized in that, The terminal includes a memory and at least one processor; the memory stores a program; the program includes instructions for executing the action prediction method based on key points and imitation learning as described in any one of claims 1-4; the processor is used to execute the program.

6. A computer-readable storage medium having multiple instructions stored thereon, characterized in that, The instructions are suitable for being loaded and executed by the processor to implement the steps of the action prediction method based on key points and imitation learning as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Model based on semi-supervised keypoints

    CN116210030A

  • Model deep reinforcement learning method based on random Transform model

    CN117454965A