Universal action learning method and system for cross-domain body basic model
By using cross-domain embodied action data sets and visual language models in the field of embodied intelligence, we learn transferable action features and generate embodied agent operation instructions, solving the problem of data action heterogeneity and improving data utilization and model efficiency.
Patent Information
- Application Number
- CN202411863257.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-23
AI Technical Summary
The problem of data action heterogeneity in the field of embodied intelligence makes it difficult for existing model training to meet heterogeneous action needs, which is inefficient and costly.
By obtaining cross-domain embodied action datasets, fine-tuning pre-trained visual language models to learn transferable action features, constructing vector quantized codebooks, learning heterogeneous decoders to decode universal action encodings, and generating specific embodied agent operation instructions.
It significantly improves cross-domain data utilization, reduces the data demand and training cost of model training, and meets the use needs of rapid migration to more types of embodied agents.
Smart Images

Figure CN120031103A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning technology, and in particular to a general action learning method and system for a cross-domain embodied basic model. Background Art
[0002] In fields such as natural language processing and computer vision, base models trained on large amounts of diverse data have demonstrated excellent generalization capabilities. These successful cases highlight the advantages of general model learning compared to models trained specifically for specific tasks. Inspired by these successful experiences, developing multi-purpose embodied base models that can handle cross-task, cross-environment, and cross-embodied agent generalization has become a promising direction for building general embodied agents.
[0003] However, the training data in the field of embodied intelligence is significantly heterogeneous, which poses a huge challenge to training embodied base models. This heterogeneity is not only reflected in the visual differences caused by camera positions (e.g., wrist perspective or third-person perspective) and environmental conditions (such as lighting or background changes), but more importantly, it is reflected in the action heterogeneity between different embodied agents. For example, embodied agents such as robotic arms with different degrees of freedom, quadruped embodied agents, and self-driving cars have completely different action spaces. In addition, the heterogeneity of control interfaces (e.g., the difference between the end effector and the speed controller of a robotic arm) may cause the same action command to have different meanings at the physical level. Furthermore, even if the actions come from the same embodied agent platform, when collected by different human operators, the heterogeneity of the data is exacerbated due to the multimodality of human behavior. The above heterogeneity significantly complicates the process of training base models based on cross-different data sources.
[0004] Currently, there is no existing solution that can adequately address the problem of data action heterogeneity in the field of embodied intelligence. Most previous studies forcibly treat different action spaces as consistent and apply the same discretization or normalization techniques. This approach is prone to action space conflicts because similar action encodings in different action spaces may correspond to completely different physical meanings. Although some studies have attempted to design a physically interpretable action space by aggregating individual action spaces to accommodate a variety of embodied agent systems, this approach often requires a lot of manual engineering work and fails to effectively explore and utilize the intrinsic connections between the basic action spaces of different embodied agents. The limitations of this approach hinder the effective development of a general embodied agent basic model. Summary of the invention
[0005] The present invention provides a general action learning method and system for a cross-domain embodied basic model, so as to solve the problem that the existing model training is difficult to meet the heterogeneous action requirements, has low efficiency and high cost.
[0006] The present invention provides a general action learning method for a cross-domain embodied basic model, comprising: Obtain a cross-domain embodied action dataset; Based on the cross-domain embodied action dataset, fine-tuning the pre-trained visual language model learns transferable action features to obtain a vector quantization codebook that can represent the universal action space; Based on the vector quantization codebook, a specific heterogeneous decoder is learned to decode the universal action code and generate specific embodied intelligent agent operation instructions to complete the model deployment.
[0007] According to a general action learning method for a cross-domain embodied basic model provided by the present invention, the obtaining of a cross-domain embodied action dataset specifically includes: Collect action image data from multiple embodied agents, including images from different perspectives and backgrounds, as well as action information in different action spaces; After data integration and preprocessing, a cross-domain embodied action dataset is generated.
[0008] According to a universal action learning method for a cross-domain embodied basic model provided by the present invention, the pre-trained visual language model based on the cross-domain embodied action dataset is fine-tuned to learn transferable action feature extraction, thereby obtaining a vector quantization codebook that can represent the universal action space, specifically including: extracting generic actions from the cross-domain embodied action dataset by fine-tuning a pre-trained visual language model as a generic action extractor; A vector quantization codebook for representing a universal action space is constructed based on the universal action.
[0009] According to a general action learning method for a cross-domain embodied basic model provided by the present invention, After constructing a vector quantization codebook for characterizing a universal action space based on the universal action, the method further comprises: The vector quantization codebook includes multiple universal action codes, each code captures the universal atomic behavior between different embodied intelligent agents, and all universal atomic behaviors constitute the universal action space.
[0010] According to a general action learning method for a cross-domain embodied base model provided by the present invention, during the general action extraction process of the fine-tuned pre-trained visual language model, general actions that can complete the task are sampled from the general action space according to the current visual information and the specified task. During the training of the visual language model, the gradient estimation of the sampling process is achieved through the classification reparameterization method Gumbel-Softmax.
[0011] According to a general action learning method for a cross-domain embodied basic model provided by the present invention, The learning of a specific heterogeneous decoder based on the vector quantization codebook decodes the universal action code and generates a specific embodied intelligent agent operation instruction to complete the model deployment, specifically including: Based on the vector quantization codebook, a specific heterogeneous decoder is learned. Each heterogeneous decoder learns the mapping relationship between the universal action and the heterogeneous action. The vector quantization codebook is decoded by the heterogeneous decoder, and specific embodied intelligent agent operation instructions are generated according to the mapping relationship between general actions and heterogeneous actions to complete the model deployment.
[0012] The present invention also provides a general action learning system for a cross-domain embodied basic model, the system comprising: A data acquisition module, used to obtain a cross-domain embodied action dataset; A universal action extraction module, used to fine-tune the pre-trained visual language model based on the cross-domain embodied action dataset to learn transferable action features and obtain a vector quantization codebook that can represent the universal action space; A heterogeneous decoding module is used to learn a specific heterogeneous decoder based on the vector quantization codebook to decode the universal action code and generate specific embodied intelligent body operation instructions to complete model deployment.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the general action learning method for a cross-domain embodied basic model as described in any one of the above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a general action learning method for a cross-domain embodied basic model as described in any one of the above.
[0015] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the general action learning method for a cross-domain embodied basic model as described in any one of the above.
[0016] The present invention provides a universal action learning method and system for a cross-domain embodied basic model. The method and system learn transferable feature extraction by fine-tuning a pre-trained visual language model based on a cross-domain embodied action dataset, obtain a vector quantization codebook representing the universal action space, and learn a specific heterogeneous decoder based on the vector quantization codebook to decode the universal action encoding and generate specific embodied intelligent body operation instructions, which significantly improves cross-domain data utilization, reduces the data demand and training cost of model training, and meets the use needs of rapid migration to more types of embodied intelligent bodies. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flowchart of a general action learning method for a cross-domain embodied basic model provided by the present invention.
[0019] Figure 2 It is a system architecture diagram of a general action learning system for a cross-domain embodied basic model provided by the present invention.
[0020] Figure 3 It is a module connection diagram of a general action learning system for a cross-domain embodied basic model provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention.
[0022] Reference numerals: 110: data acquisition module; 120: general action extraction module; 130: heterogeneous decoding module; 410: processor; 420: communication interface; 430: memory; 440: communication bus. DETAILED DESCRIPTION
[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0024] Combine the following Figure 1A general action learning method for a cross-domain embodied basic model of the present invention is described, comprising: step 100, obtaining a cross-domain embodied action dataset.
[0025] Specifically, we collect action image data from multiple embodied agents, covering images from different perspectives and backgrounds as well as action information with different action spaces. After data integration and preprocessing, we generate a cross-domain embodied action dataset.
[0026] In the present invention, by acquiring various heterogeneous action images, it is possible to subsequently extract more types of general actions and adapt to different types of embodied intelligent agents.
[0027] The present invention captures universal atomic action behaviors by learning universal action encodings, which can be universal on different embodied intelligent agent platforms without being restricted by specific embodied intelligent agent types and control interfaces. For example, when different embodied intelligent agents face a target in front, they should perform similar "move forward" behaviors, although the corresponding control signals for this behavior on different embodied intelligent agents may be completely different. This kind of motion behavior transcends specific embodied intelligent agents and control constraints, making it universally applicable to different embodied intelligent agent platforms, thus providing great potential for the utilization of heterogeneous data and model generalization. In this sense, only a small amount of parameters and data are required to decode universal actions into actions for specific agents, because general motion behaviors have been captured in a universal space, and the decoder only needs to learn a small amount of deployment details for a specific agent. Therefore, during the deployment process, this method can effectively adapt to different types of embodied intelligent agents.
[0028] Step 200: fine-tune the pre-trained visual language model based on the cross-domain embodied action dataset to learn transferable action features, and obtain a vector quantization codebook that can represent a universal action space.
[0029] Specifically, it includes: extracting universal actions from the cross-domain embodied action dataset by fine-tuning a pre-trained visual language model as a universal action extractor; A vector quantization codebook for representing a universal action space is constructed based on the universal action.
[0030] Wherein, after constructing a vector quantization codebook for characterizing the universal action space based on the universal action, it includes: The vector quantization codebook includes multiple universal action codes, each code captures the universal atomic behavior between different embodied intelligent agents, and all universal atomic behaviors constitute the universal action space.
[0031] When the fine-tuned pre-trained visual language model is extracting universal actions, the model samples universal actions that can complete the task from the universal action space based on the current visual information and the specified task. During the model training process, the classification reparameterization method Gumbel-Softmax is used to estimate the gradient of the sampling process.
[0032] In the present invention, the visual language model (VLM) constructs the universal action space as a vector quantized codebook. Similar to a learnable skill library, each encoding encapsulates an atomic behavior that is sufficiently general and can be performed by different embodied agents. This design drives the VLM to recognize and exploit shared atomic behaviors across different action spaces, and it performs far better than other embodied base models, such as OpenVLA with 7B parameters, on a large number of downstream tasks. The extracted universal actions can be converted into precise and actionable commands for various embodied agents by lightweight heterogeneous decoders. These decoders take the universal action encoding as a conditional input and map the universal actions to specific control signals based on features extracted from their respective observations. The decoders can be flexibly customized according to specific requirements, for example, the structure can be adjusted to adapt to changes in the proprioceptive information or the number of camera views. Rapid adaptation to new embodied agent platforms can be achieved by simply learning lightweight decoders for new tasks. The methods and systems are comprehensively evaluated on challenging tasks, including adapting to a wide range of view changes and embodied agents that do not appear in the training data. Evaluation results show that the proposed method and system have strong transfer capabilities and good performance, demonstrating the great advantages of developing embodied base models in a general action space.
[0033] Specifically, in the process of universal action extraction, the model focuses not only on the changes in visual observations, but more on understanding the task progress. Specifically, a large visual language model is fine-tuned as a universal action extractor, which outputs the probability of selecting different universal actions given visual observations and task goals, similar to planning in the universal action space, inferring the most relevant universal action for a given task under observation. VLM is adopted to achieve this purpose due to its strong visual language reasoning ability. In addition, in the process of learning universal actions, the use of pre-trained VLM greatly improves the sample efficiency. This extractor plays an important role in cross-domain generalization, because different embodied agents are forced to use the same vector quantization codebook to capture universal and shared atomic behaviors. Classification reparameterization is used in the training process, and the Gumbel-Sofmax technique is used to achieve gradient estimation.
[0034] Step 300: Based on the vector quantization codebook, a specific heterogeneous decoder is learned to decode the universal action code, and specific embodied intelligent agent operation instructions are generated to complete the model deployment.
[0035] Specifically, a specific heterogeneous decoder is learned based on the vector quantization codebook, and each heterogeneous decoder learns the mapping relationship between the universal action and the heterogeneous action; The vector quantization codebook is decoded by the heterogeneous decoder, and specific embodied intelligent agent operation instructions are generated according to the mapping relationship between general actions and heterogeneous actions to complete the model deployment.
[0036] In this paper, in order to effectively transform highly abstract behaviors in the generic action space into precise control signals for specific embodied agents, it is crucial to integrate the details of specific embodied agents (such as control type, proprioception, and different observations). To address this problem, a series of lightweight decoders are introduced to adapt the mapping relationship between each type of embodied agent and generic actions. Each heterogeneous decoder is specially designed to learn the mapping from generic actions to specific embodied agent control signals.
[0037] Since overly complex decoding with too many parameters may overfit the data distribution of the target domain and reduce the learning efficiency of the general action space, all heterogeneous decoders are implemented as simple MLP networks that take the general action and visual features extracted from a shared visual backbone as input. By keeping the decoding lightweight, we ensure that most of the learning is done on the general action, thereby maximizing the generalization between different embodied agents.
[0038] Among them, MLP (Multilayer Perceptron) is a typical feedforward neural network, which consists of multiple levels of nodes (or "neurons"), each node is connected to all nodes in the next layer. MLP is one of the simplest fully connected feedforward neural network forms and is widely used in various machine learning tasks such as classification and regression.
[0039] A standard MLP consists of the following parts: Input layer: receives input data. Each input node represents a dimension of input features.
[0040] Hidden layer: One or more layers between the input layer and the output layer. Each layer contains a number of nodes that perform weighted summation and activation function operations. The number and size of hidden layers can be adjusted according to the complexity of the problem.
[0041] Output layer: generates model prediction results. For classification problems, the output layer usually corresponds to the number of categories; for regression problems, there may be only one output node.
[0042] The specific working principle is as follows: Forward propagation: During the training process, the input data is passed layer by layer from the input layer to the output layer. Each node in each layer calculates the weighted sum of its input signals and converts it into the output signal of the node through the activation function. This process is called forward propagation. Activation function: In order to introduce nonlinear characteristics, an activation function is applied to each node in the MLP. Common activation functions include Sigmoid, ReLU (Rectified Linear Unit), tanh, etc. Nonlinear activation functions enable MLP to approximate complex nonlinear mapping relationships. Loss function: It is used to measure the difference between the model output and the actual label. Choose an appropriate loss function according to the specific task, such as cross entropy loss is suitable for classification tasks, and mean square error is suitable for regression tasks. Back propagation algorithm: This is the core part of MLP training. By calculating the gradient of the loss relative to each weight, and using the gradient descent method to update the weight parameters to minimize the loss function. This process is iterated repeatedly until convergence.
[0043] Through full connection: Each node in each layer of MLP is connected to all nodes in the next layer, forming a dense connection pattern. Foundation of deep learning: Although there are many more complex architectures (such as convolutional neural networks CNN and recurrent neural networks RNN), MLP is still an important foundation for understanding the concepts of deep learning. Universal approximator: In theory, given enough hidden units and appropriate activation functions, MLP can approximate any continuous function, which makes it a powerful tool for solving many types of problems.
[0044] MLP is widely used in image recognition, natural language processing, speech recognition and other fields. However, as the size of the dataset increases and the performance requirements of the model increase, more specialized network structures (such as CNN for images, RNN or Transformer for sequence data) often achieve better results. However, in some cases, especially when the input data is low-dimensional and has no obvious spatial or temporal structure, MLP is still very effective.
[0045] Experiments have shown that planning in a universal action space can solve the problem of action heterogeneity. The universal action space encodes universal atomic behaviors across multiple agents, which significantly improves data utilization across domains and promotes generalization across embodied agents, making small-parameter models outperform SOTA models with larger parameters. In addition, the learned universal actions can be accurately converted to the actions of any specific embodied agent through heterogeneous decoding, allowing rapid adaptation to new embodied agents with different control interfaces and physical properties. In addition, the learned universal action extractor can also be used as a universal action marker, providing power for the construction of future large-scale embodied base models.
[0046] A universal action learning method for a cross-domain embodied basic model provided by the present invention is based on the cross-domain embodied action dataset. By fine-tuning the pre-trained visual language model to learn transferable feature extraction, a vector quantization codebook is obtained, and a specific heterogeneous decoder is learned based on the vector quantization codebook to decode the universal action encoding, and specific embodied intelligent body operation instructions are generated. This significantly improves the cross-domain data utilization rate, reduces the data demand and training cost of model training, and meets the needs of more types of embodied intelligent bodies to quickly get started.
[0047] refer to Figure 2 and Figure 3 The present invention also discloses a general action learning system for a cross-domain embodied basic model, the system comprising: A data acquisition module 110 is used to acquire a cross-domain embodied action dataset; A universal action extraction module 120, configured to fine-tune a pre-trained visual language model based on the cross-domain embodied action dataset to learn transferable action features and obtain a vector quantization codebook that can represent a universal action space; The heterogeneous decoding module 130 is used to learn a specific heterogeneous decoder based on the vector quantization codebook to decode the general action code and generate specific embodied intelligent agent operation instructions to complete the model deployment.
[0048] Among them, obtaining a cross-domain embodied action dataset specifically includes: Collect action image data from multiple embodied agents, including images from different perspectives and backgrounds, as well as action information in different action spaces; After data integration and preprocessing, a cross-domain embodied action dataset is generated.
[0049] Based on the cross-domain embodied action dataset, the pre-trained visual language model is fine-tuned to learn transferable action feature extraction, thereby obtaining a vector quantization codebook that can represent the universal action space, specifically including: extracting generic actions from the cross-domain embodied action dataset by fine-tuning a pre-trained visual language model as a generic action extractor; A vector quantization codebook for representing a universal action space is constructed based on the universal action.
[0050] After constructing a vector quantization codebook for representing the universal action space based on the universal action, the method includes: The vector quantization codebook includes multiple universal action codes, each code captures the universal atomic behavior between different embodied intelligent agents, and all universal atomic behaviors constitute the universal action space.
[0051] The fine-tuned pre-trained visual language model samples general actions that can complete the task from the general action space according to the current visual information and the specified task during the general action extraction process. During the visual language model training process, the gradient estimation of the sampling process is achieved through the classification reparameterization method Gumbel-Softmax.
[0052] Based on the vector quantization codebook, a specific heterogeneous decoder is learned to decode the universal action code, and a specific embodied intelligent agent operation instruction is generated to complete the model deployment, specifically including: Based on the vector quantization codebook, a specific heterogeneous decoder is learned. Each heterogeneous decoder learns the mapping relationship between common actions and heterogeneous actions. The vector quantization codebook is decoded by the heterogeneous decoder, and specific embodied intelligent agent operation instructions are generated according to the mapping relationship between general actions and heterogeneous actions to complete the model deployment.
[0053] Based on a universal action learning system for a cross-domain embodied basic model provided by the present invention, a pre-trained visual language model is fine-tuned based on a cross-domain embodied action dataset to learn transferable feature extraction, a vector quantization codebook is obtained, and a specific heterogeneous decoder is learned based on the vector quantization codebook to decode the universal action encoding, and generate specific embodied intelligent body operation instructions, which significantly improves the cross-domain data utilization rate, reduces the data demand and training cost of model training, and meets the needs of more types of embodied intelligent bodies to quickly get started. .
[0054] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute a general action learning method for a cross-domain embodied basic model, the method comprising: obtaining a cross-domain embodied action dataset; fine-tuning a pre-trained visual language model based on the cross-domain embodied action dataset to learn transferable action features, and obtaining a vector quantization codebook that can characterize a general action space; based on the vector quantization codebook, learning a specific heterogeneous decoder to decode the general action encoding, and generating specific embodied intelligent body operation instructions to complete model deployment.
[0055] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0056] On the other hand, the present invention also provides a computer program product, which includes a computer program, and the computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a general action learning method for a cross-domain embodied basic model provided by the above methods, and the method includes: obtaining a cross-domain embodied action dataset; fine-tuning a pre-trained visual language model based on the cross-domain embodied action dataset to learn transferable action features to obtain a vector quantization codebook that can represent the general action space; based on the vector quantization codebook, learning a specific heterogeneous decoder to decode the general action encoding, and generating specific embodied intelligent body operation instructions to complete model deployment.
[0057] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute a general action learning method for a cross-domain embodied basic model provided by the above-mentioned methods, the method comprising: obtaining a cross-domain embodied action dataset; fine-tuning a pre-trained visual language model based on the cross-domain embodied action dataset to learn transferable action features, and obtaining a vector quantization codebook that can characterize the general action space; based on the vector quantization codebook, learning a specific heterogeneous decoder to decode the general action encoding, and generating specific embodied intelligent body operation instructions to complete model deployment.
[0058] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0059] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A general action learning method for a cross-domain embodied base model, characterized in that: include: Obtain a cross-domain embodied action dataset; Based on the cross-domain embodied action dataset, fine-tuning the pre-trained visual language model learns transferable action features to obtain a vector quantization codebook that can represent the universal action space; Based on the vector quantization codebook, a specific heterogeneous decoder is learned to decode the universal action code and generate specific embodied intelligent agent operation instructions to complete the model deployment.
2. The general action learning method for a cross-domain embodied basic model according to claim 1, characterized in that: The obtaining of the cross-domain embodied action dataset specifically includes: Collect action image data from multiple embodied agents, including images from different perspectives and backgrounds, as well as action information in different action spaces; After data integration and preprocessing, a cross-domain embodied action dataset is generated.
3. The general action learning method for a cross-domain embodied basic model according to claim 1, characterized in that: The fine-tuning of the pre-trained visual language model based on the cross-domain embodied action dataset learns transferable action feature extraction, thereby obtaining a vector quantization codebook that can represent the universal action space, specifically includes: extracting generic actions from the cross-domain embodied action dataset by fine-tuning a pre-trained visual language model as a generic action extractor; A vector quantization codebook for representing a universal action space is constructed based on the universal action.
4. The general action learning method for a cross-domain embodied basic model according to claim 3, characterized in that: After constructing a vector quantization codebook for characterizing a universal action space based on the universal action, the method further comprises: The vector quantization codebook includes multiple universal action codes, each code captures the universal atomic behavior between different embodied intelligent agents, and all universal atomic behaviors constitute the universal action space.
5. The general action learning method for a cross-domain embodied basic model according to claim 3, characterized in that: The fine-tuned pre-trained visual language model samples general actions that can complete the task from the general action space according to the current visual information and the specified task during the general action extraction process. During the visual language model training process, the gradient estimation of the sampling process is achieved through the classification reparameterization method Gumbel-Softmax.
6. The general action learning method for a cross-domain embodied basic model according to claim 1, characterized in that: The learning of a specific heterogeneous decoder based on the vector quantization codebook decodes the universal action code and generates a specific embodied intelligent agent operation instruction to complete the model deployment, specifically including: Based on the vector quantization codebook, a specific heterogeneous decoder is learned. Each heterogeneous decoder learns the mapping relationship between common actions and heterogeneous actions. The vector quantization codebook is decoded by the heterogeneous decoder, and specific embodied intelligent agent operation instructions are generated according to the mapping relationship between general actions and heterogeneous actions to complete the model deployment.
7. A general action learning system for cross-domain embodied base models, characterized in that The system comprises: A data acquisition module, used to obtain a cross-domain embodied action dataset; A universal action extraction module, used to fine-tune the pre-trained visual language model based on the cross-domain embodied action dataset to learn transferable action features and obtain a vector quantization codebook that can represent the universal action space; A heterogeneous decoding module is used to learn a specific heterogeneous decoder based on the vector quantization codebook to decode the universal action code and generate specific embodied intelligent body operation instructions to complete model deployment.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the general action learning method for a cross-domain embodied basic model is implemented as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the general action learning method for a cross-domain embodied basic model as claimed in any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the general action learning method for a cross-domain embodied basic model as claimed in any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Imitation learning method, device and equipment for intelligent agent with body
CN121189522A