Mechanical arm control method based on potential actions
Through the robotic arm manipulation method based on potential actions, the potential action extraction module and forward model are used to train potential strategy models and specific action models, and the problem of insufficient flexibility and adaptability of robotic arm manipulation technology in changing environments and tasks is solved, achieving more efficient task execution and lower data acquisition costs.
Patent Information
- Application Number
- CN202510274934.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing robotic arm handling technologies rely on precise programming or complex sensors, lack flexibility and adaptability, especially in variable environments and tasks to reduce efficiency. At the same time, imitation learning methods rely on a large amount of high-quality expert demonstration data, which has high acquisition costs and insufficient generalization capabilities.
Using a robotic arm manipulation method based on potential actions, data preprocessing is performed by selecting video clips of tasks performed by human first perspective, a potential action extraction module and forward model are constructed, and a potential strategy model and specific action model are trained to reduce dependence on high-quality expert demonstration data.
It improves the adaptability of the robotic arm in new tasks and dynamic environments, reduces the cost and difficulty of data acquisition, and improves the efficiency and success rate of task execution.
Smart Images

Figure CN120038750A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of manipulator operation, and particularly to a manipulator control method based on potential actions. Background Art
[0002] Existing manipulator control technologies usually rely on precise programming or complex sensor systems to execute tasks. Although these methods can complete specific operation tasks, they often lack sufficient flexibility and adaptability when facing changing environments and tasks. For example, in an industrial production line, a manipulator usually needs to be precisely programmed for a specific task. Once the task requirements or the environment change, the manipulator may not be able to adapt quickly, resulting in a decrease in efficiency or even task failure. In addition, traditional manipulator learning methods require a large amount of high-quality labeled data and a complex model training process, which not only increases costs but also limits the widespread deployment of manipulators in practical applications.
[0003] In recent years, with the development of deep learning technologies, imitation learning has become an effective manipulator control method. Imitation learning is a technique for training an agent by observing and imitating the behavior of an expert. Its core idea is to use expert demonstration data to let the agent learn how to take appropriate actions in different states to complete a specific task. By analyzing the action sequences and decision-making processes when humans execute tasks, a manipulator can execute complex tasks without explicit programming instructions. This method reduces the dependence on precise programming to a certain extent and enables the manipulator to more naturally adapt to complex tasks.
[0004] However, the existing imitation learning methods still have the following deficiencies:
[0005] 1) Dependence on a large amount of high-quality expert demonstration data: Existing imitation learning methods usually require a large amount of high-quality expert demonstration data to train the model. These data not only need to be precisely labeled but also need to cover various possible task scenarios and operation steps. However, obtaining such high-quality data is costly and often difficult to achieve in practical applications. For example, in some highly specialized or extreme environments, the acquisition of expert demonstration data may be restricted by operation risks, technical limitations, or insufficient resources. In addition, even if the data can be obtained, the diversity and coverage of the data may be insufficient, resulting in limited performance of the model in practical applications.
[0006] 2) Insufficient generalization ability: Existing imitation learning methods often exhibit poor generalization ability when faced with new tasks or new environments. Since model training relies on demonstration data in specific scenarios, when task requirements or environmental conditions change, the model may not be able to effectively adapt. For example, in an industrial scenario, a robotic arm may need to handle objects of different shapes, sizes, or materials, and existing methods often require retraining or adjusting the model when faced with these changes, resulting in low efficiency. This lack of generalization ability limits the application potential of robotic arms in diverse tasks and dynamic environments. Summary of the Invention
[0007] To at least to some extent solve one of the technical problems existing in the prior art, an object of the present invention is to provide a robotic arm control method, device, and medium based on latent actions.
[0008] The first technical solution adopted by the present invention is as follows:
[0009] A robotic arm control method based on latent actions, comprising the following steps:
[0010] Select video clips of a human performing a task from the first-person perspective and perform data preprocessing in combination with robotic arm operation videos;
[0011] Construct a latent action extraction module and a forward model; wherein, the latent action extraction module is used to extract the latent action for converting the current observation image to the future observation image based on the current observation image and the future observation image, and the forward model is used to generate the future observation image based on the current observation image and the extracted latent action;
[0012] Use the preprocessed video data to train the latent action extraction module and the forward model;
[0013] Construct a latent policy model and train the latent policy model in combination with the trained forward model; the latent policy model is used to predict the latent action from the current observation image and the specified task;
[0014] Based on the latent policy model, train a specific action model; the specific action model is used to take the latent action output by the latent policy model as the subtask target and combine the current observation to output the specific robotic arm action;
[0015] Given a task description, obtain the current observation image of the robotic arm, input it into the trained latent policy model to obtain the latent action vector; input the task description, the current observation image, and the latent action vector into the trained specific action model to obtain the specific action and let the robotic arm execute it.
[0016] Furthermore, the latent action extraction module includes multiple attention layers, each of which consists of a self-attention layer and a feed-forward layer, for extracting high-level feature representations from input data; after the attention layer, a codebook containing multiple learnable vectors is designed, each vector representing a unique latent action, for mapping the extracted features to a specific action space;
[0017] Among them, in the first attention layer, the input is the concatenation of the feature vectors obtained after the current observation image and the future observation image are encoded by the image encoder; the input of each subsequent attention layer is the output of the previous attention layer. Through layer-by-layer transmission and optimization, more abstract and task-related potential action vectors are gradually extracted; then the dimensions of the feature vector and the potential action vector are aligned through a multi-layer perceptron; finally, the potential action vector closest to the feature vector is selected in the codebook as the potential action.
[0018] Furthermore, the image encoder is implemented using a pre-trained Vision Transformer (ViT);
[0019] The operation process of the potential action extraction module is as follows:
[0020]
[0021] F i+1 =FFN(SelfAtten(F ii )+F i )
[0022] z e =MLP(F 12 )
[0023]
[0024] In the formula, I t and I t+1 Respectively represent the current observation image and future observation image of the input, ViT represents the image encoder, represents the concatenation operation, F represents the feature vector of the network output, FFN represents the feedforward layer, SelfAtten represents the self-attention layer, MLP represents the multi-layer perceptron, z e represents the feature vector aligned with the potential action vector, C represents the codebook, c j represents the learnable latent action vector in the codebook, ||·|| 2 represents the Euclidean norm, and z represents the potential action vector extracted by the potential action extraction module.
[0025] Further, the forward model includes multiple attention layers and a downsampling layer. Each attention layer consists of a cross-attention layer, a self-attention layer, and a feed-forward layer, and is used to predict future observed images from the current observed image and the latent action;
[0026] Among them, in the first cross-attention layer, the input is the feature vector obtained after the current observed image passes through the image encoder and the latent action vector extracted from the latent action extraction module. After passing through the cross-attention layer, a fused feature vector is generated and input into the self-attention layer. Then, the feature vector obtained after passing through the feed-forward layer transformation and the latent action vector are input into the next cross-attention layer, and the above process is repeated; after passing through multiple attention layers, the obtained feature vector is finally downsampled and then recombined to obtain the predicted future observed image.
[0027] Further, the operation process of the forward model is as follows:
[0028] F 0 = ViT(I t )
[0029] F i+1 = FFn(SelfAtten(CrossAtten(F i ,z e +sg[z - z e ) + F i )
[0030]
[0031] In the formula, I t represents the input current observed image, ViT represents the image encoder, F represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, CrossAtten represents the cross-attention layer, sg[·] represents the gradient clipping operation, DownSampled represents the downsampling layer, represents the future observed image predicted by the forward model.
[0032] Further, training the latent action extraction module and the forward model using the preprocessed video data includes:
[0033] Extract the current observed image and the future observed image from the preprocessed video data, and input them as training data into the latent action extraction module;
[0034] The latent action extraction module extracts the latent action vector from the current observed image and the future observed image, and this vector contains the semantic information of the action from the current observation to the future observation;
[0035] The extracted potential action vector and the current observed image are used as the input of the forward model. The forward model combines the detailed features of the current observed image and the high-level semantic information of the potential action to predict the future observed image;
[0036] Calculate the loss function based on the predicted future observed image and the real future observed image, and optimize the parameters of the potential action extraction module and the forward model;
[0037] Among them, the calculation formula of the loss function is as follows:
[0038]
[0039] L total =L recon +L embedding +αL commit
[0040] In the formula, L recon represents the image reconstruction loss, MSE represents the mean square error, I t+1 represents the real future observed image, represents the future observed image predicted by the forward model, L embedding represents the encoding loss for training the codebook, α is its corresponding weight coefficient, L commit represents the focus error to make the output closer to the codebook, and L total is the final loss function.
[0041] Furthermore, the potential policy model includes multiple attention layers. Each attention layer consists of a cross-attention layer, a self-attention layer, and a feed-forward layer, and is used to predict potential actions from the current observed image and the specified task;
[0042] Among them, in the first cross-attention layer, the input is the vector obtained by passing the task description through the text encoder and the feature vector obtained by passing the current observed image through the image encoder; after passing through the cross-attention layer, the generated fused feature vector is input into the self-attention layer, and then the feature vector obtained after passing through the feed-forward layer transformation and the potential action vector are input into the next cross-attention layer, repeating the above process; after passing through multiple attention layers, the potential action is output;
[0043] The potential action output by the potential policy model and the current observed image are used as the input of the trained forward model. The forward model outputs the predicted future observed image, and the loss function is calculated based on the predicted future observed image and the real future observed image to train the network parameters of the potential policy model.
[0044] Furthermore, the text encoding is the CLIP model, and its parameters are frozen during training;
[0045] The operation process of the potential policy model is as follows:
[0046] F 0 = CLIP(T)
[0047] F i+1 = FFN(SelfAtten(CrossAtten(F i , ViT(I t ) + F i )
[0048] z e = MLP(F 12 )
[0049]
[0050] In the formula, T represents the input task text description, CLIP represents the text encoder, F represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, MLP represents the multi-layer perceptron, z e represents the feature vector aligned with the potential action vector, C represents the codebook, c j represents the potential action vector in the learned codebook, ||·|| 2 represents the Euclidean norm, and z represents the potential action vector output by the potential policy model;
[0051] After the potential policy model obtains the potential action vector, it is used together with the current observed image as the input of the forward model to obtain the predicted future observed image:
[0052] F' 0 = ViT(I t )
[0053] F' i+1 = FFN(SelfAtten(CrossAtten(F' i , z e + sg[z - z e ) + F' i )
[0054]
[0055] In the formula, I t represents the input current observed image, ViT represents the image encoder, F' represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, CrossAtten represents the cross-attention layer, sg[·] represents the gradient truncation operation, z eThe feature vector after alignment with the latent action vector is denoted as, z represents the latent action vector output by the latent policy model, and DownSampled represents the downsampling layer. represents the future observed image predicted by the forward model;
[0056] The calculation process of the loss function is as follows:
[0057]
[0058] L total = L recon + L commit
[0059] In the formula, L recon represents the image reconstruction loss, MSE represents the mean square error, and I t+1 represents the true future observed image, represents the future observed image predicted by the forward model, and L commit represents the focus error for making the output closer to the codebook, and L total is the final loss function.
[0060] Furthermore, the specific action model includes multiple attention layers. Each attention layer consists of a self-attention layer and a feed-forward layer. Taking the latent action output by the latent policy model as the sub-task target and combining the current observation, it outputs the specific robotic arm action.
[0061] Among them, in the first attention layer, the input is the concatenation of the feature vector obtained by encoding the current observed image through the image encoder, the vector obtained by encoding the task description through the text encoder, and the latent action vector output by the latent policy model; the input of each subsequent attention layer is the output of the previous attention layer; finally, the predicted specific action is obtained through a multi-layer perceptron.
[0062] Calculate the loss function according to the predicted specific action and the true specific action, and train the network parameters of the specific action model.
[0063] Furthermore, the image encoder is implemented using a pre-trained Vision Transformer, and the text encoder is implemented using a pre-trained CLIP model, and their parameters are frozen during training.
[0064] The operation process of the specific action model is as follows:
[0065]
[0066] F i+1 = FFN(SelfAtten(F ii ) + F i )
[0067]
[0068] In the formula, I t represents the current observed image of the input, ViT represents the image encoder, represents the concatenation operation, T represents the input task text description, CLIP represents the text encoder, F represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, and MLP represents the multi-layer perceptron, represents the specific actions of the robotic arm predicted by the specific action model (x, y, z, w x , w y , w z , g), where x, y, and z represent the displacements of the end of the robotic arm on the x, y, and z axes, and w x , w y , w z represents the rotation angles of the end of the robotic arm on the x, y, and z axes, and g represents the opening and closing of the gripper;
[0069] The calculation process of the loss function is as follows:
[0070]
[0071] In the formula, L recon represents the action prediction loss, MSE represents the mean square error, a represents the true specific action, represents the specific action predicted by the specific action model.
[0072] The second technical solution adopted by the present invention is:
[0073] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a robotic arm control method based on potential actions as described above.
[0074] The fourth technical solution adopted by the present invention is:
[0075] A computer-readable storage medium, and at least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a robotic arm control method based on potential actions as described above.
[0076] The fifth technical solution adopted by the present invention is:
[0077] A computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the above-mentioned robotic arm manipulation method based on potential actions.
[0078] The beneficial effects of the present invention are as follows: By introducing a potential policy model and using potential actions to capture the commonalities between tasks, the robotic arm can quickly adapt to new tasks and dynamic environments, thereby improving the efficiency and success rate of task execution. Additionally, by combining large-scale network video data for training, the present invention reduces the dependence on high-quality expert demonstration data, lowers the cost and difficulty of data acquisition, and simultaneously broadens the diversity of data sources. Description of the Drawings
[0079] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0080] Figure 1 It is a flowchart of the steps of a robotic arm manipulation method based on potential actions in an embodiment of the present invention;
[0081] Figure 2 It is a schematic diagram of the training framework disclosed in an embodiment of the present invention;
[0082] Figure 3 It is a schematic diagram of the structure of the potential action extraction module and the forward model in an embodiment of the present invention;
[0083] Figure 4 It is a schematic diagram of the structure of the potential policy model and the specific action model in an embodiment of the present invention. Detailed Embodiments
[0084] The following details the embodiments of the present invention. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0085] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., it is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.
[0086] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, above, below, within, etc. are understood as including the present number. If the first and second are described only for the purpose of distinguishing technical features, it cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.
[0087] In the description of the present invention, unless otherwise clearly defined, words such as setting, installing, connecting, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meaning of the above words in the present invention in combination with the specific content of the technical solution.
[0088] In view of the above problems, the present invention proposes a manipulator control scheme based on latent actions, mainly solving the problems of the existing manipulator control algorithms' dependence on a large amount of high-quality expert demonstration data and insufficient generalization ability. The existing technologies usually rely on precise expert demonstration data, which are not only costly and difficult to obtain, but also difficult to collect in highly specialized or extreme environments. In addition, although large-scale network video data is easy to obtain, due to its significant differences from the manipulator operation scenario, it cannot be directly migrated to imitation learning, resulting in the challenges of data scarcity and insufficient generalization ability in the actual application of the existing methods. The present invention significantly reduces the dependence on high-quality expert demonstration data by introducing the concept of latent actions and combining large-scale network video data and manipulator expert demonstration data, while improving the adaptability of the manipulator in new tasks and dynamic environments.
[0089] Embodiment 1
[0090] As Figure 1 shown, this embodiment provides a manipulator control method based on latent actions, including the following steps:
[0091] S1. Select relevant video segments of humans performing tasks from the first-person perspective in large-scale network data, and perform data preprocessing in combination with manipulator operation videos.
[0092] In some embodiments, the large-scale network data is a public video dataset Something-Something, which is a collection of videos with text descriptions of people completing various tasks from a first-person perspective; the robotic arm operation video is expert demonstration data collected in the Calvin simulation environment, which contains the execution video of the task and the corresponding text description, and also records the specific actions of the robotic arm. Select videos with clear images in the video dataset, and crop the images to a uniform size of 224×224.
[0093] S2. Construct a potential action extraction module and a forward model; wherein the potential action extraction module is used to extract the potential action from the current observation image to the future observation image based on the current observation image and the future observation image, and the forward model is used to generate the future observation image based on the current observation image and the extracted potential action.
[0094] In some embodiments, the potential action extraction module is constructed as follows:
[0095] See also Figure 3 , the potential action extraction module in this embodiment is as follows Figure 3 As shown in the dotted box in the lower right corner. The core structure of this module includes 12 attention layers, each of which consists of a self-attention layer and a feedforward layer, which is used to extract high-level feature representations from the input data. After the attention layer, a codebook containing 128 learnable vectors is designed, each of which represents a unique potential action, which is used to map the extracted features to a specific action space. In the first attention layer, the input is the concatenation of the feature vectors obtained after the current observation image and the future observation image are encoded by the image encoder. In this embodiment, the pre-trained Vision Transformer (ViT) is used as the image encoder, and its parameters are frozen during training. The input of each subsequent attention layer is the output of the previous attention layer. Through layer-by-layer transmission and optimization, more abstract and task-related potential action features are gradually extracted. Then the dimensions of the feature vector and the potential action vector are aligned through a multi-layer perceptron. Finally, the model selects the potential action vector closest to the feature vector in the codebook as the potential action output.
[0096] In summary, the operation process of the potential action extraction module is:
[0097]
[0098] F i+1 =FFN(SelfAtten(F i )+F i )
[0099] z e= MLP(F 12 )
[0100]
[0101] wherein, I t and I t+1 respectively represent the current observed image and the future observed image of the input, ViT represents the image encoder, represents the concatenation operation, F represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, MLP represents the multi-layer perceptron, z e represents the feature vector aligned with the latent action vector, C represents the codebook, c j represents the learnable latent action vector in the codebook, ||·|| 2 represents the Euclidean norm, and z represents the latent action vector extracted by the latent action extraction module.
[0102] In some embodiments, the forward model is constructed as follows:
[0103] Referring to Figure 3 , the forward model in this embodiment is as shown in the upper dotted box in Figure 3 . The core structure of the forward model includes 24 attention layers and a downsampling layer. Each attention layer consists of a cross-attention layer, a self-attention layer, and a feed-forward layer, and is used to predict the future observed image from the current observed image and the latent action. In the first cross-attention layer, the input is the feature vector obtained by passing the current observed image through the image encoder and the latent action vector extracted from the latent action extraction module. After passing through the cross-attention layer, the fused feature vector is generated and input into the self-attention layer. Then, the feature vector obtained after passing through the feed-forward layer transformation and the latent action vector are input into the next cross-attention layer, and the above process is repeated. The feature vector obtained after passing through 24 attention layers is finally downsampled by a factor of 5 and recombined into a 224×224 picture, which is output as the predicted future observed image.
[0104] In summary, the operation process of the forward model is as follows:
[0105] F 0 = ViT(I t )
[0106] F i+1 = FFN(SelfAtten(CrossAtten(F i , z e + sg[z - z e ) + F i )
[0107]
[0108] In the formula, I t represents the current observed image of the input, ViT represents the image encoder, F represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, CrossAtten represents the cross-attention layer, sg[·] represents the gradient clipping operation, and DownSampled represents the downsampling layer. represents the future observed image predicted by the forward model.
[0109] S3. Using the preprocessed video data, train the latent action extraction module and the forward model to obtain a certain number of latent action vectors with high-level semantic information.
[0110] Refer to Figure 2 , the action extraction module and the forward model training process in this embodiment are as shown in the left half. First, extract the current observed image and the future observed image from the preprocessed video clip and input them into the action extraction module as training data. The action extraction module extracts latent action vectors from the current observed image and the future observed image through a multi-layer attention mechanism and codebook mapping. This vector contains the semantic information of the action transformed from the current observation to the future observation. Subsequently, the extracted latent action vector and the current observed image are used as the input of the forward model. The forward model combines the detailed features of the current observed image and the high-level semantic information of the latent action through a multi-layer cross-attention and self-attention mechanism to predict the future observed image. Subsequently, calculate the loss function and gradient according to the predicted future observed image and the real future observed image, and use the Adam backpropagation algorithm for training to train the network parameters of the action extraction module and the forward model. The calculation formula of the loss function is as follows: Figure 2
[0111]
[0112]
[0113] L total = L recon + L embedding + αL commit
[0114] In the formula, L recon represents the image reconstruction loss, MSE represents the mean square error, I t+1 represents the real future observed image, represents the future observed image predicted by the forward model, L embedding represents the encoding loss for training the codebook, α is its corresponding weight coefficient, which is taken as 0.25 in this embodiment, L commitIndicates making the output closer to the codebook focusing error, L total is the final loss function.
[0115] S4. Construct a latent policy model and train the latent policy model in combination with the trained forward model; the latent policy model is used to predict latent actions from the current observation image and the specified task;
[0116] Specifically, use the preprocessed video data to train a latent policy model. This model predicts latent actions through the current observation, and these latent actions represent high-level task decomposition and action guidance. Refer to Figure 4 , the latent policy model in this embodiment is as Figure 4 shown by the dotted box below. The core structure of the latent policy model includes 12 attention layers, and each attention layer is composed of a cross-attention layer, a self-attention layer, and a feed-forward layer, which is used to predict latent actions from the current observation image and the specified task. In the first cross-attention layer, the input is the vector obtained by passing the task description through the text encoder and the feature vector obtained by passing the current observation image through the image encoder. In this embodiment, the text is encoded as a pre-trained CLIP model, and its parameters are frozen during training. After passing through the cross-attention layer, a fused feature vector is generated and input into the self-attention layer, and then the feature vector obtained by passing through the feed-forward layer transformation and the latent action vector are input into the next cross-attention layer, repeating the above process. After 12 attention layers, latent actions are output. Subsequently, the latent actions output by the model are combined with the current 510555. After that, calculate the loss function and gradient according to the predicted future observation image and the real future observation image, and use the Adam backpropagation algorithm for training to train the network parameters of the latent policy model.
[0117] In summary, the operation process of the latent policy model is as follows:
[0118] F 0 =CLIP(T)
[0119] F i+1 =FFN(SelfAtten(CrossAtten(F i ,ViT(I t )+F i )
[0120] z e =MLP(F 12 )
[0121]
[0122] Wherein, T represents the input task text description, CLIP represents the text encoder, F represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, MLP represents the multi-layer perceptron, and z e represents the feature vector aligned with the latent action vector, C represents the codebook, and c j represents the latent action vector in the codebook learned in step S3, ||·|| 2 represents the Euclidean norm, and z represents the latent action vector output by the latent policy model.
[0123] After the latent policy model obtains the latent action vector, it is used together with the current observed image as the input of the forward model to obtain the predicted future observed image:
[0124] F′ 0 = ViT(I t )
[0125] F′ i+1 = FFN(SelfAtten(CrossAtten(F′ i ,z e +sg[z - z e ) + F′ i )
[0126]
[0127] Wherein, I t represents the input current observed image, ViT represents the image encoder, F′ represents the feature vector output by the network, FFN represents the feed-forward layer, SelfAtten represents the self-attention layer, CrossAtten represents the cross-attention layer, sg[·] represents the gradient clipping operation, z e represents the feature vector aligned with the latent action vector, z represents the latent action vector output by the latent policy model, DownSampled represents the downsampling layer, represents the future observed image predicted by the forward model.
[0128] The calculation process of the loss function is as follows:
[0129]
[0130] L total = L recon + L commit
[0131] Wherein, L recon represents the image reconstruction loss, MSE represents the mean square error, I t+1 represents the true future observed image, Denote the future observation image predicted by the forward model, L commit Denote the focus error that makes the output closer to the codebook, L total is the final loss function.
[0132] S5. Based on the latent policy model, train a specific action model; the specific action model is used to take the latent action output by the latent policy model as the sub-task target, and combine the current observation to output the specific manipulator action;
[0133] See Figure 4 , the specific action model in this embodiment is as shown in the Figure 4 upper dotted box. The core structure of the specific action model includes 12 layers of attention layers. Each layer of attention layer consists of a self-attention layer and a feed-forward layer. It takes the latent action output by the latent policy model as the sub-task target, combines the current observation, and outputs the specific manipulator action. In the first layer of attention layer, the input is the concatenation of the feature vector obtained by encoding the current observation image through the image encoder, the vector obtained by encoding the task description through the text encoder, and the latent action vector output by the latent policy model. In this embodiment, the pre-trained Vision Transformer (ViT) is used as the image encoder, and the pre-trained CLIP model is used as the text encoder, and their parameters are frozen during training. The input of each subsequent layer of attention layer is the output of the previous layer of attention layer. Then, the predicted specific action is obtained through a multi-layer perceptron. Subsequently, the loss function and gradient are calculated according to the predicted specific action and the real specific action, and the Adam backpropagation algorithm is used for training to train the network parameters of the specific action model.
[0134] In summary, the operation process of the specific action model is as follows:
[0135]
[0136] F i+1 = FFN(SelfAtten(F i ) + F i )
[0137]
[0138] In the formula, I t represents the input current observation image, ViT represents the image encoder, represents the concatenation operation, T represents the input task text description, CLIP represents the text encoder, F represents the feature vector output by the network, FFN represents the feed-forward layer, SelftAtten represents the self-attention layer, MLP represents the multi-layer perceptron, represents the specific manipulator action (x, y, z, w) predicted by the specific action modelx , w y , w z , g), where x, y, and z represent the displacements of the end - effector of the robotic arm along the x, y, and z axes, and w x , w y , w z represents the rotation angles of the end - effector of the robotic arm along the x, y, and z axes, and g represents the opening and closing of the gripper.
[0139] The calculation process of the loss function is as follows:
[0140]
[0141] In the formula, L recon represents the action prediction loss, MSE represents the mean squared error, a represents the true specific action, represents the specific action predicted by the specific action model.
[0142] S6. Given the task description, obtain the current observation image of the robotic arm, input it into the trained latent policy model to obtain the latent action vector; input the task description, the current observation image, and the latent action vector into the trained specific action model to obtain the specific action and let the robotic arm execute it.
[0143] Exemplarily, given the task description, obtain the current observation image of the robotic arm. At the same time, input it into the trained latent policy model to obtain the latent action vector. Then, input the three into the trained specific action model to obtain the specific action and let the robotic arm execute it. After that, obtain the observation image at the next moment as the new current observation image and input it into the latent policy model. Repeat the above process until the task is completed or the maximum number of steps is reached.
[0144] In summary, the method of the present invention has at least the following advantages compared with the prior art:
[0145] 1) Reducing data dependence and enhancing diversity: By combining large - scale network video data for training, the present invention reduces the dependence on high - quality expert demonstration data, reduces the cost and difficulty of data acquisition, and at the same time broadens the diversity of data sources.
[0146] 2) Introducing latent policy and enhancing adaptability: By introducing the latent policy model, the present invention captures the commonalities between tasks using latent actions, enabling the robotic arm to quickly adapt to new tasks and dynamic environments, thereby improving the efficiency and success rate of task execution.
[0147] Embodiment 2
[0148] An embodiment of the present invention further provides an electronic device, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and is loaded and executed by the processor to implement a method for manipulating a robotic arm based on potential actions as described in Figure 1 shown.
[0149] It can be understood that the memory may include a random access memory (RAM), or may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0150] The processor may include one or more processing cores. The processor uses various interfaces and circuits to connect various parts within the entire server, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by calling data stored in the memory, it executes various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a central processing unit (CPU) and a modem, etc. in one or several combinations. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately by a single chip.
[0151] Since this electronic device is the electronic device corresponding to the method for manipulating a robotic arm based on potential actions in an embodiment of the present invention, and the principle of the electronic device for solving problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and the repeated parts will not be elaborated.
[0152] Embodiment 3
[0153] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement a Figure 1 method for manipulating a robotic arm based on potential actions as shown.
[0154] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, and the storage medium includes a read-only memory (ROM), a random access memory (RAM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc memories, a magnetic disk memory, a tape memory, or any other computer-readable medium capable of carrying or storing data.
[0155] Since this storage medium is the storage medium corresponding to the method for manipulating a robotic arm based on potential actions in an embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0156] Embodiment 4
[0157] In some possible embodiments, various aspects of the method of the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a method for manipulating a robotic arm based on potential actions according to various exemplary embodiments described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0158] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0159] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples.
[0160] The above embodiments are only for illustrating the technical concept and characteristics of the present invention, and their purpose is to enable ordinary technical personnel in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for manipulating a robotic arm based on potential actions, characterized in that: The following steps are involved: Select video clips of humans performing tasks from the first-person perspective and combine them with robotic arm operation videos for data preprocessing; Constructing a potential action extraction module and a forward model; wherein the potential action extraction module is used to extract a potential action converted from the current observation image to the future observation image according to the current observation image and the future observation image, and the forward model is used to generate the future observation image according to the current observation image and the extracted potential action; Using the preprocessed video data, the latent action extraction module and the forward model are trained; Constructing a potential policy model and training the potential policy model in combination with the trained forward model; the potential policy model is used to predict potential actions from a current observed image and a specified task; Based on the potential strategy model, a specific action model is trained; the specific action model is used to use the potential action output by the potential strategy model as a subtask goal, and output a specific robot arm action in combination with the current observation; Given a task description, obtain the current observation image of the robot arm, input the trained potential policy model, and obtain the potential action vector; input the task description, current observation image and potential action vector into the trained specific action model, and obtain the specific action to be executed by the robot arm.
2. A method for manipulating a robotic arm based on potential actions according to claim 1, characterized in that: The potential action extraction module includes multiple attention layers, each of which consists of a self-attention layer and a feed-forward layer, for extracting high-level feature representations from input data; After the attention layer, a codebook containing multiple learnable vectors is designed, each of which represents a unique latent action, to map the extracted features into a specific action space. Among them, in the first attention layer, the input is the concatenation of the feature vectors obtained after the current observation image and the future observation image are encoded by the image encoder; the input of each subsequent attention layer is the output of the previous attention layer. Through layer-by-layer transmission and optimization, more abstract and task-related potential action vectors are gradually extracted; then the dimensions of the feature vector and the potential action vector are aligned through a multi-layer perceptron; finally, the potential action vector closest to the feature vector is selected in the codebook as the potential action.
3. A method for manipulating a robotic arm based on potential actions according to claim 1 or 2, characterized in that: The image encoder is implemented using a pre-trained Vision Transformer; The operation process of the potential action extraction module is as follows: F0=ViT(I t )⊕WeT(I t+1 ) F i+1 =FFN(SelfAtten(F i )+F i ) With e =MLP(F 12 ) C={c1,c2,...,c 128 } In the formula, I t and I t+1 They represent the current observation image and the future observation image of the input, ViT represents the image encoder, ⊕ represents the splicing operation, F represents the feature vector of the network output, FFN represents the feedforward layer, SelfAtten represents the self-attention layer, MLP represents the multi-layer perceptron, and z e represents the feature vector aligned with the potential action vector, C represents the codebook, c j represents the learnable potential action vector in the codebook, ||·||2 represents the Euclidean norm, and z represents the potential action vector extracted by the potential action extraction module.
4. The method for controlling a robotic arm based on potential actions according to claim 1, characterized in that: The forward model includes multiple attention layers and a downsampling layer, each of which is composed of a cross-attention layer, a self-attention layer, and a feedforward layer, for predicting future observation images from current observation images and potential actions; Among them, in the first cross-attention layer, the input is the feature vector obtained after the current observation image passes through the image encoder and the potential action vector extracted from the potential action extraction module. After passing through the cross-attention layer, a fused feature vector is generated and then input into the self-attention layer. After that, the feature vector and the potential action vector obtained by the feedforward layer transformation are input into the next cross-attention layer; after passing through multiple layers of attention layers, the obtained feature vector is finally down-sampled and reorganized to obtain the predicted future observation image.
5. A method for manipulating a robotic arm based on potential actions according to claim 1 or 4, characterized in that: The operation process of the forward model is: F0=ViT(I t ) F i+1 =FFN(SelfAtten(CrossAtten(F i ,with e +sg[zz e ])+F i ) In the formula, I t represents the current observed image of the input, ViT represents the image encoder, F represents the feature vector of the network output, FFN represents the feedforward layer, SelfAtten represents the self-attention layer, CrossAtten represents the cross-attention layer, sg[·] represents the gradient truncation operation, DownSampled represents the downsampling layer, represents the future observation image predicted by the forward model.
6. The method for controlling a robotic arm based on potential actions according to claim 1, characterized in that: The method of using the preprocessed video data to train the potential action extraction module and the forward model includes: Extracting the current observation image and the future observation image from the preprocessed video data and inputting them into the potential action extraction module as training data; The latent action extraction module extracts the latent action vector from the current observation image and the future observation image. The vector contains the semantic information of the action converted from the current observation to the future observation. The extracted latent action vector and the current observed image are used as the input of the forward model. The forward model combines the detailed features of the current observed image and the high-level semantic information of the latent action to predict the future observed image. Calculate the loss function based on the predicted future observation images and the actual future observation images, and optimize the parameters of the potential action extraction module and the forward model; The calculation formula of the loss function is as follows: L total =L recon +L embedding +αL commit Where, L recon represents the image reconstruction loss, MSE represents the mean square error, I t+1 represents the real future observation image, represents the future observation image predicted by the forward model, L embedding represents the coding loss for training the codebook, α is its corresponding weight coefficient, L commit Indicates making the output closer to the codebook focus error, L total is the final loss function.
7. The method for controlling a robotic arm based on potential actions according to claim 1, characterized in that: The latent policy model includes multiple attention layers, each of which consists of a cross-attention layer, a self-attention layer, and a feed-forward layer, for predicting latent actions from the current observed image and the specified task; Among them, in the first cross-attention layer, the input is the vector obtained by the text encoder of the task description and the feature vector obtained by the image encoder of the current observation image; after the cross-attention layer, the fused feature vector is generated and then input into the self-attention layer, and then the feature vector and the potential action vector obtained by the feedforward layer transformation are input into the next cross-attention layer; after multiple layers of attention layers, the potential action is output; The potential action output by the potential policy model and the current observation image are used as the input of the trained forward model. The forward model outputs the predicted future observation image. The loss function is calculated based on the predicted future observation image and the actual future observation image to train the network parameters of the potential policy model.
8. The method for controlling a robotic arm based on potential actions according to claim 7, characterized in that: The text is encoded into a CLIP model and its parameters are frozen during training; The operation process of the potential strategy model is: F0=CLIP(T) F i+1 =FFN(SelfAtten(CrossAtten(F i ,We(I t )+F i ) With e =MLP(F 12 ) C={c1,c2,...,c 128 } In the formula, T represents the input task text description, CLIP represents the text encoder, F represents the feature vector output by the network, FFN represents the feedforward layer, SelfAtten represents the self-attention layer, MLP represents the multi-layer perceptron, and z e represents the feature vector aligned with the potential action vector, C represents the codebook, c j represents the potential action vector in the learned codebook, ||·||2 represents the Euclidean norm, and z represents the potential action vector output by the potential policy model; After the potential policy model obtains the potential action vector, it is used together with the current observation image as the input of the forward model to obtain the predicted future observation image: F′0=ViT(I t ) F′ i+1 =FFN(SelfAtten(CrossAtten(F′ i ,with e +sg[zz e ])+F′ i ) In the formula, I t represents the current observed image of the input, ViT represents the image encoder, F′ represents the feature vector of the network output, FFN represents the feedforward layer, SelfAtten represents the self-attention layer, CrossAtten represents the cross-attention layer, sg[·] represents the gradient truncation operation, z e represents the feature vector aligned with the potential action vector, z represents the potential action vector output by the potential strategy model, DownSampled represents the downsampling layer, represents the future observation image predicted by the forward model; The calculation process of the loss function is as follows: L total =L recon +L commit Where, L recon represents the image reconstruction loss, MSE represents the mean square error, I t+1 represents the real future observation image, represents the future observation image predicted by the forward model, L commit Indicates making the output closer to the codebook focus error, L total is the final loss function.
9. The method for controlling a robotic arm based on potential actions according to claim 1, characterized in that: The specific action model includes multiple attention layers, each of which consists of a self-attention layer and a feedforward layer; Among them, in the first attention layer, the input is the concatenation of the feature vector obtained by encoding the current observation image through the image encoder, the vector obtained by the text encoder of the task description, and the potential action vector output by the potential strategy model; the input of each subsequent attention layer is the output of the previous attention layer; finally, the predicted specific action is obtained through the multi-layer perceptron; The loss function is calculated based on the predicted specific actions and the actual specific actions, and the network parameters of the specific action model are trained.
10. A method for manipulating a robotic arm based on potential actions according to claim 9, characterized in that: The image encoder is implemented by using a pre-trained Vision Transformer, and the text encoder is implemented by using a pre-trained CLIP model, and the parameters of the two are frozen during training; The calculation process of the specific action model is: F0=ViT(I t )⊕CLIP(T)⊕z F i+1 =FFN(SelfAtten(F i )+F i ) In the formula, I t represents the current observed image of the input, ViT represents the image encoder, ⊕ represents the splicing operation, T represents the input task text description, CLIP represents the text encoder, F represents the feature vector of the network output, FFN represents the feedforward layer, SelfAtten represents the self-attention layer, MLP represents the multi-layer perceptron, represents the specific action of the robot arm predicted by the specific action model (x, y, z, w x ,w y ,w z ,g), where x, y, z represent the displacement of the end of the robot arm on the three axes of x, y, and z, and w x ,w y ,w z It indicates the rotation angle of the end of the robot arm on the three axes of x, y, and z, and g indicates the opening and closing of the gripper; The calculation process of the loss function is as follows: Where, L recon represents the action prediction loss, MSE represents the mean square error, a represents the actual specific action, Represents the specific action predicted by the specific action model.
Citation Information
Patent Citations
Determining an environmental adjustment action
CN114423574A
Autonomous pickup and placement pose acquisition method for robot in disordered scene
CN118081758A
Mechanical arm control method and device based on pre-training model, equipment and medium
CN118418134A
Method and device for determining control strategy model and method and device for controlling end effector
CN119347753A
Method for determining a moment for the operation of a robot using a model and method for teaching the model.
DE102022205011A1