Mechanical arm-oriented visual language model based on meta learning and learning method thereof
Through the visual language model based on meta-learning, using frozen pre-trained models and lightweight metamappers, the gap in visual and language modal fields in multi-modal and small sample learning is solved, and efficient multi-modal task learning and computing resources are achieved.
Patent Information
- Application Number
- CN202510264198.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively bridge the domain gap between visual and linguistic modalities in multi-modal and small sample learning, and its dependence on large-scale data and computing resources leads to insufficient resource limitations and flexibility.
Using a visual language model based on meta-learning, the model is trained in a completely autoregressive manner by freezing the pre-trained visual encoder and language model, and introducing a lightweight metamapper, connecting visual and language samples.
It reduces the computational cost, improves the generalization ability and flexibility of the model, and can show high efficiency and adaptability in multimodal and small sample tasks.
Smart Images

Figure CN120220050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-modal few-shot meta-learning methods, and particularly relates to a vision-language model for a robotic arm based on meta-learning and a learning method thereof. Background Art
[0002] In recent years, with the rapid development of deep learning and neural networks, cross-modal learning has gradually become a research hotspot in both academia and industry. Learning from a small number of samples observed in a multi-modal environment and quickly learning is an important part of human intelligence. Multi-modal few-shot learning (FSL) aims to achieve cross-modal fast learning and generalization ability through limited labeled data. However, current vision and language models still face considerable challenges in dealing with limited label spaces when performing multi-modal few-shot learning tasks.
[0003] In contrast to the challenges in multi-modal scenarios, language models have made significant progress in the past few years, especially when applied to limited label spaces. This is mainly due to large-scale pre-training and the improvement of model capacity. For example, models such as BERT and GPT series have learned rich language representation capabilities through pre-training on a large amount of unlabeled data, and thus perform excellently in few-shot learning scenarios. These advancements in the field of natural language processing (NLP) have also inspired similar efforts in the vision field, giving rise to multi-modal models such as CLIP (Contrastive Language-Image Pre-training). CLIP has learned transferable visual representation capabilities through natural language supervision and achieved impressive results in few-shot and zero-shot classification tasks. However, there is a significant domain gap between the visual and language modalities, which makes it still very difficult to directly transfer few-shot learning capabilities to multi-modal scenarios.
[0004] In addition, the pre-training of these models usually relies on large-scale datasets and high computational costs, and the resource requirements limit their feasibility in practical applications. For example, although the Flamingo model has made great progress in multi-modal few-shot learning, it still has significant limitations when dealing with complex tasks. At the same time, many methods rely on manually designed task induction mechanisms, lacking flexibility and being difficult to cope with dynamic and complex task requirements.
[0005] Specifically, in the process of combining multimodal learning with robot grasping tasks, there are still the following defects: 1) Domain gap problem: There are significant differences between visual and language modalities in data expression, distribution characteristics and semantic information. This domain gap makes it difficult for existing methods to directly transfer the few-shot learning ability of the visual modality to multimodal scenarios, resulting in insufficient generalization ability. 2) Dependence on computing resources: Current mainstream multimodal models (such as CLIP) usually require large-scale data sets for pre-training, and the training process consumes a lot of computing resources. This high cost limits its feasibility in practical applications. 3) Lack of flexibility in task induction mechanisms: Existing methods usually rely on manually designed task induction mechanisms, which are difficult to cope with complex or unknown tasks. In addition, these manually designed mechanisms are difficult to adapt to dynamic and changing task requirements, limiting the adaptability of the model. 4) Limitations on model parameter updates: Most multimodal few-shot methods require fine-tuning of model parameters during the inference process, which is a significant obstacle for resource-constrained or real-time application scenarios. 5) Lack of lightweight design: Current methods over-rely on complex neural network architectures and multimodal interaction modules, ignoring the necessity of lightweight design, which limits the application of models in scenarios such as embedded devices.
[0006] Therefore, the technical problem that needs to be solved urgently is: how to more effectively bridge the domain gap between visual and language modalities, reduce dependence on large-scale data and computing resources, and achieve efficient multimodal few-sample task learning. Summary of the invention
[0007] The present invention is made to solve the above-mentioned problems, and its purpose is to provide a visual language model for robotic arms based on meta-learning and a learning method thereof.
[0008] The present invention provides a visual language model for a robotic arm based on meta-learning, which has the following characteristics: a visual encoder, wherein the visual encoder is a frozen visual encoder v φ , used to obtain visual samples; a language model, wherein the language model is a language model having a text embedding device g ψ and the generator g ω A frozen language model for obtaining a language sample; and a meta-mapper, the meta-mapper being a meta-mapper f having a trainable meta-parameter θ θ , the meta-mapper is used to connect the visual samples and the language samples, so as to train the visual language model in a fully autoregressive manner.
[0009] In the visual language model based on meta-learning for robotic arms provided by the present invention, the following features may also be provided: wherein the visual encoder is defined as a function v φ , whose parameters is fixedly obtained from the pre-trained visual encoder, with the input being the original image x and the output being the extracted visual feature v φ (x) = x1,…,x n .
[0010] In the visual language model for robotic arms based on meta-learning provided by the present invention, it may also have the following feature: among them, a set of learnable parameters is used to map the visual encoding to the latent space of the language model in the multi-modal few-shot learning setting by the meta-mapper. The parameters are the visual prefix of the language model, and its dimension d e is the same as the dimension of the language embedding in the language model. The visual prefix is added before the encoded visual feature to form the following sequence: [p1,…,p l ,x1,…,x n . Regarding this representation as a set of ordered elements, the entire set is encoded simultaneously through the self-attention mechanism, and the set multi-head attention module is used as the meta-mapper with trainable meta-parameters θ, which is defined as follows: MetaMap θ (Q,K,V) = σ(QK T )V where the dot product QK T measures the similarity between features and is used for feature weighting through the activation function σ. If the dot product of Q and K is larger, the weight of the corresponding V feature will be higher. Let Q = K = V = [p1,…,p l ,x1,…,x n , that is, the input of the meta-mapper is the feature sequence, and the output is a set of learned parameters p’1…,p’ l , which is the visual prefix of the language model: p’1…,p l = MetaMap θ ([p1,…,p l ,x1,…,x n ).
[0011] The self-attention layer, as the core component of the set multi-head attention module, extracts meaningful information from the visual features x1,…,x n through its unique pairwise similarity weighting mechanism between elements and accumulates it into p’1…,p l . The meta-parameters of the meta-mapper are learned and shared across all tasks in T meta-train on D
[0012] In the visual language model for robotic arms based on meta-learning provided by the present invention, it may also have the following feature: among them, the language model is defined as a function g ω, which parameterizes a probability distribution over the text sequence y, and the language model includes an embedding function g ψ , whose parameter ψ embeds each word token y i in the text sequence y into the word token embedding t i . Subsequently, autoregressive text generation is performed by a Transformer module. The language model receives the visual prefix p’1…,p l concatenated with the word token embeddings t1,…,t m as input and generates subsequent word tokens autoregressively based on the prefix: t i+1 = g ω ([p’1…,p’ l t1,…,t i ), i < m. Initialize ω with the parameters pre-trained from a large-scale language model and keep it completely frozen.
[0013] The present invention also provides a learning method for a vision-language model for a robotic arm based on meta-learning, having the following features, including: meta-training. In the meta-training stage, the vision-language model learns through batch-wise multi-modal few-shot tasks. The multi-modal few-shot tasks include a support set and a query set . Among them, the support set is used to adjust the task-specific parameters θ' i of the model, and the query set is used to update the meta-parameters θ; meta-testing. In the meta-testing stage, the vision-language model uses the support set for fast adaptation and evaluates the performance on the query set . The generation process adopts an open-ended autoregressive manner and realizes the output of the language model through top-k nucleus sampling.
[0014] In the learning method for a vision-language model for a robotic arm based on meta-learning provided by the present invention, it may also have the following features: Among them, during the meta-training process, task batches are sampled from the distribution of the multi-modal few-shot learning tasks. Each task consists of a support set and a query set . Define the vision-language model as a function fθ that takes an image x as input and generates y as output. The loss function optimized for each task during training is the cross-entropy loss, defined as follows: When adapting to a new task , the trainable meta-parameters θ are updated to the task-specific parameters θ' i . These task-specific parameters are calculated through N-step gradient updates as follows: where α is the step size hyperparameter, using the query set samples, based on the task-specific parameter θ' i Optimize the meta-parameter θ of the model, initialize it as the task-specific model parameter, and the optimization objective is: Meta-optimization is performed through all tasks using the following random gradient descent (SGD) update rule: where β is the step size hyperparameter.
[0015] In the learning method of the meta-learning based vision-language model for robotic arms provided by the present invention, it can also have the following feature: wherein, in the meta-test stage, evaluate multi-modal few-shot tasks containing new classes These tasks include a support set for quickly adapting to a given task by fine-tuning the model meta-parameter θ, and a query set for evaluating the performance of the model on this task. The answer generation for each sample in the query set is completed in an open-ended autoregressive manner, using top-k nucleus sampling, sampling from the language model after the visual prefix of the given image. The final performance is calculated by the average prediction accuracy of the query samples of all meta-test tasks.
[0016] Functions and effects of the invention
[0017] According to the meta-learning based vision-language model for robotic arms and its learning method involved in the present invention, the present invention utilizes the publicly available pre-trained large vision encoder and language model, constructs a bridge between the vision and language modalities through a lightweight meta-mapper network, combines multi-modal learning with the robot grasping task, and combines visual perception and action control to achieve efficient multi-modal few-shot task learning. It not only reduces the computational cost, but also improves the generalization ability and flexibility of the model, and has broad application prospects. Brief description of the drawings
[0018] Figure 1 is the architecture of the multi-modal meta-few-shot learner of the meta-learning based vision-language model for robotic arms in the embodiment of the present invention;
[0019] Figure 2 is the learning method of the meta-learning based vision-language model for robotic arms in the embodiment of the present invention. Detailed implementation manners
[0020] In the description of the present application, it should be noted that unless otherwise clearly specified and limited, the terms "installation", "connection", and "coupling" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection, an electrical connection, or a connection that allows mutual communication; it can be a direct connection, or an indirect connection through an intermediate medium, and can be the communication inside two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0021] In order to make the technical means, creative features, achieved purposes, and functions of the present invention easy to understand, the following embodiments will specifically describe a meta-learning-based visual language model for a robotic arm and its learning method in conjunction with the accompanying drawings.
[0022] Figure 1 It is the architecture of the multi-modal meta few-shot learner of the meta-learning-based visual language model for a robotic arm in the embodiments of the present invention.
[0023] As Figure 1 shown, the present invention discloses a meta-learning-based visual language model for a robotic arm, aiming to achieve fast learning and generalization of visual and language modalities through a small number of labeled samples, including: a visual encoder, a language model, and a meta mapper. In the shown example, the model generates the last word "retriever" in an autoregressive manner.
[0024] The visual encoder is a frozen visual encoder v φ , which is used to obtain visual samples.
[0025] The visual encoder is defined as a function v φ , and its parameters are fixedly obtained from the pre-trained visual encoder. The input is the original image x, and the output is the extracted visual feature v φ (x) = x1,..., x n .
[0026] The language model is a frozen language model with a text embedder g ψ and a generator g ω , which is used to obtain language samples. To avoid updating the pre-trained parameters ω of the language model, the present invention keeps it in a frozen state.
[0027] The language model is defined as a function g ω , which parameterizes a probability distribution for the text sequence y. The model includes an embedding function g ψ , and its parameters ψ embed each word token y i in the text sequence y into the word token embedding t iSubsequently, autoregressive text generation is performed by a Transformer module.
[0028] Specifically, the language model receives the visual prefix p’1…,p l concatenated with the word token embeddings t1,…,t m as input and generates subsequent word tokens in an autoregressive manner based on the prefix:
[0029] t i+1 = g ω ([p’1…,p’ l t1,…,t i ), i < m.
[0030] Since our method aims to learn in a limited token space, updating the model parameters may lead to distortion of the learned parameters. Therefore, we initialize ω with the parameters pre-trained from a large-scale language model and keep it completely frozen. The parameters of the pre-trained language model have been optimized in an autoregressive manner through the standard language modeling objective function.
[0031] The meta-mapper is the meta-mapper f with trainable meta-parameters θ θ which is used to connect the visual samples and the language samples, thus training the vision-language model in a fully autoregressive manner and having the flexibility to handle downstream generative multi-modal tasks.
[0032] For the meta-mapper to map visual encodings to the latent space of the language model in a multi-modal few-shot learning setting, we use a set of learnable parameters (a total of l) which are also called the visual prefix of the language model, and its dimension d e is the same as the dimension of the language embedding. We add this visual prefix in front of the encoded visual features to form the following sequence:
[0033] [p1,…,p l ,x1,…,x n
[0034] Specifically, we regard this representation as a set of ordered elements and encode the entire set simultaneously through the self-attention mechanism. To this end, we adopt the set multi-head attention module as the meta-mapper with trainable meta-parameters θ. Its definition is as follows:
[0035] MetaMap θ (Q,K,V) = σ(QK T )V
[0036] where the dot product QKT The similarity between measurement features is calculated and used for feature weighting through the activation function σ. Intuitively, if the dot product of Q and K is larger, the weight of the corresponding V feature will be higher. We set
[0037] Q = K = V = [p1, …, p l , x1, …, x n ,
[0038] That is, the input of the meta-mapper is a feature sequence, and the output is a set of learned parameters p’1 …, p’ l , which is the visual prefix of the language model:
[0039] p’1 …, p l = MetaMap θ ([p1, …, p l , x1, …, x n ).
[0040] As the core component of this module, the self-attention layer extracts meaningful information from the visual features x1, …, x n through its unique pairwise similarity weighting mechanism among elements, and accumulates it into p’1 …, p l . The meta-parameters of the meta-mapper are learned and shared across all tasks in T meta-train on D.
[0041] During meta-training, we sample task batches from the distribution of multi-modal few-shot learning tasks. Each task consists of a support set and a query set . Among them, the support set is used to adjust the task-specific parameters θ' i of the model; the query set is used to update the meta-parameters θ.
[0042] Here, for simplicity, we assume that the model is defined as a function fθ that takes an image x as input and generates y as output.
[0043] The loss function optimized for each task during training is the cross-entropy loss, defined as follows:
[0044]
[0045] When adapting to a new task , the trainable meta-parameters θ are updated to the task-specific parameters θ' i . These task-specific parameters are calculated through N-step gradient updates as follows:
[0046]
[0047] where α is the step-size hyperparameter. This process is called inner-loop update.
[0048] Next, using the samples of the query set optimize the meta-parameters θ of the model, initialized as task-specific model parameters, with the optimization objective being: i This process is called outer-loop optimization.
[0049]
[0050] Meta-optimization is performed over all tasks
[0051] using the stochastic gradient descent (SGD) update rule as follows: where β is the step-size hyperparameter.
[0052]
[0053] This is the learning method of the vision-language model for the robotic arm based on meta-learning in the embodiments of the present invention.
[0054] Figure 2 As
[0055] shown, for the multi-modal few-shot meta-learning task, there are two classes (ways) in the support set images, and each class is represented by one sample (shot). Given a batch of tasks T Figure 2 i i First, use the support set to update through several gradient steps to obtain the task-specific model parameters θ for each task i Then use it together with the query set samples to perform a meta-update step to update the meta-parameters θ. After meta-training, for a newly given task, further adjust the meta-learning model using the support set and measure the performance of the unseen query samples, and use the meta-trained model for inference.
[0056] Specifically, in the meta-test phase, we evaluate multi-modal few-shot tasks containing new classes These tasks include a support set for quickly adapting to the given task by fine-tuning the model meta-parameters θ, and a query set for evaluating the performance of the model on this task. The answer generation for each sample in the query set is done in an open-ended autoregressive manner. Specifically, we use top-k nucleus sampling to sample from the language model after the visual prefix of the given image.
[0057] The final performance is calculated as the average prediction accuracy of the query samples for all meta-test tasks.
[0058] Functions and effects of the embodiments
[0059] The present invention proposes a novel meta-learning method for multi-modal few-shot learning. By introducing a lightweight trainable meta-mapper, a model is designed, which serves as a bridge between a large frozen vision model and a language model. The meta-mapper observes tasks sequentially, accumulates shared meta-knowledge, and encodes it into a learnable visual prefix to guide the language model to generate relevant language explanations.
[0060] Different from previous work, the present invention uses support samples to adapt the model to unseen tasks, which has been proven to effectively improve performance without relying on manually designed task induction mechanisms. In addition, by cleverly combining the operation of large pre-trained models, this method avoids the high computational cost of training from scratch and thus has significant computational efficiency.
[0061] It should be noted that this architecture is flexible enough to further integrate additional modal functions, which will be studied in future work. Finally, experiments verify the effectiveness of this method, which outperforms existing baselines on multiple benchmark tasks and further promotes the exploration of the multi-modal few-shot meta-learning field.
[0062] Table 1
[0063]
[0064]
[0065] In Table 1, we compare the performance of the proposed method with the baseline method (Frozen) and the upper bound benchmark (ANIL) on the miniImageNet dataset, including 2-way and 5-way classification tasks in the Real-Name and Open-Ended settings, as well as the 1-shot and 5-shot sample number scenarios.
[0066] The experimental results show that the Frozen baseline method performs poorly without a task induction mechanism. For example, in the Real-Name 2-way 1-shot task, the accuracy is only 1.7%. After introducing task induction, the performance of Frozen improves, but the accuracy in complex tasks (such as Open-Ended 5-way) is still low, only 20.2%. In contrast, the proposed method outperforms the baseline method in all settings. Especially when combined with episodic training and domain-shift, the performance is greatly improved. For example, in the Real-Name 5-way 1-shot task, the accuracy of the proposed method reaches 29.0%, significantly higher than the baseline method. In addition, the proposed method also shows superior generation ability in the Open-Ended task, with an accuracy improvement of more than 10% compared to the baseline. Although ANIL performs best as an upper bound benchmark, since it is a discriminative method and not applicable to generative tasks, the proposed method shows stronger flexibility and generalization ability in generative tasks. This indicates that the proposed method can not only effectively handle few-shot learning scenarios, but also has a lower computational cost and stronger task adaptation ability.
[0067] Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.
Claims
1. A visual language model for robotic arms based on meta-learning, characterized in that: include: The visual encoder is a frozen visual encoder v φ , used to obtain visual samples; Language model, the language model is a language model with a text embedder g ψ and the generator g ω A frozen language model is used to obtain language samples; as well as Meta-mapper, the meta-mapper being a meta-mapper f with trainable meta-parameters θ θ , the meta-mapper is used to connect the visual samples and the language samples, so as to train the visual language model in a fully autoregressive manner.
2. The visual language model for robotic arms based on meta-learning according to claim 1, characterized in that: in, The visual encoder is defined as a function v φ , whose parameters It is fixed from the pre-trained visual encoder, the input is the original image x, and the output is the extracted visual feature v φ (x)=x1,…,x n .
3. The visual language model for robotic arms based on meta-learning according to claim 2, characterized in that: in, Using a set of learnable parameters The meta-mapper thus maps the visual encoding to the latent space of the language model in a multimodal few-shot learning setting, and the parameter is the visual prefix of the language model, whose dimension d e The visual prefix is added to the encoded visual features with the same dimension as the language embedding in the language model, forming the following sequence: [p1,…,p l ,x1,…,x n ] This representation is considered as a set of ordered elements, and the entire set is encoded simultaneously through the self-attention mechanism. The collective multi-head attention module is used as a meta-mapper with trainable meta-parameters θ, which is defined as follows: MetaMap θ (Q,K,V)=σ(QK T )V Among them, the dot product QK T Measures the similarity between features and uses the activation function σ to weight features. If the dot product of Q and K is larger, the weight of the corresponding V feature will be higher. Let Q=K=V=[p1,…,p l ,x1,…,x n ], That is, the input of the meta-mapper is a feature sequence, and the output is a set of learned parameters p'1…,p' l , which is the visual prefix of the language model: p'1…,p l =MetaMap θ ([p1,…,p l ,x1,…,x n ]), As the core component of the collective multi-head attention module, the self-attention layer uses its unique pairwise similarity weighting mechanism between elements to extract the visual features x1,…,x n Extract meaningful information and accumulate it into p'1…,p l The meta-parameters of the meta-mapper are passed through all tasks in T in D meta-train Learn and share on.
4. The meta-learning-based visual language model for robotic arms according to claim 3, characterized in that: in, The language model is defined as a function g ω , which parameterizes a probability distribution for the text sequence y. The language model contains an embedding function g ψ , whose parameter ψ marks each word in the text sequence y i Embedded into word token embedding t i Then, a Transformer module performs autoregressive text generation. The language model receives the visual prefixes p'1...,p l with word token embeddings t1,…,t m It takes the concatenation of as input and generates subsequent word tokens in an autoregressive manner based on the prefix: t i+1 =g ω ([p'1…,p' l t1,…,t i ]),i <m. Initialize ω using parameters pre-trained from a large-scale language model and keep it completely frozen.
5. A learning method for a visual language model for a robotic arm based on meta-learning as claimed in any one of claims 1 to 4, characterized in that: include: Meta-training, in which the visual language model is learned through batches of multimodal few-shot tasks, which include support sets and queryset Among them, the support set Task-specific parameters θ' used to tune the model i , the query set Used to update meta-parameters θ; Meta-testing, in the meta-testing phase, the visual language model uses the support set Perform fast adaptation and in the query set The performance is evaluated on the above. The generation process adopts an open autoregressive method and the language model output is achieved through top-k kernel sampling.
6. The learning method of the visual language model for a robotic arm based on meta-learning according to claim 5, characterized in that: During the meta-training process, the distribution of multimodal few-shot learning tasks The task batches are sampled from the source, each task consists of a support set and a queryset The visual language model is defined as a function fθ that accepts an image x as input and generates y as output. The loss function optimized for each task during training is the cross entropy loss, defined as follows: Adapting to new tasks When θ is θ, the trainable meta-parameter θ is updated to the task-specific parameter θ' i , these task-specific parameters are calculated through N steps of gradient updates as follows: Where α is the step size hyperparameter, Using QuerySets samples, based on the task-specific parameters θ' i The meta-parameters θ of the optimization model are initialized to the task-specific model parameters, and the optimization objective is: Meta-optimization through all tasks To do this, we use the stochastic gradient descent (SGD) update rule as follows: Where β is the step size hyperparameter.
7. The learning method of the visual language model for a robotic arm based on meta-learning according to claim 5, characterized in that: In the meta-testing phase, multimodal few-shot tasks containing new categories are evaluated These tasks include a support set is used to quickly adapt to a given task by fine-tuning the model meta-parameters θ, and a query set D i ts , used to evaluate the performance of the model on this task, the answer generation for each sample in the query set is done in an open autoregressive manner, using top-k kernel sampling, sampling from the language model after the visual prefix of the given image, and the final performance is calculated by the average prediction accuracy of the query samples of all meta-test tasks.