Information-Theory-Constrained Continuous Learning Robot Control Method
Patent Information
- Application Number
- CN202610980054.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-18
AI Technical Summary
方法解决现有技术中视觉语言动作模型在持续学习场景下存在的灾难性遗忘严重、跨模态对齐关系易退化、历史任务保持能力不足以及机器人控制稳定性下降的问题
[0036] 1. This invention, through the dual constraints of preserving historical anchor points and preserving cross-modal mutual information, can simultaneously maintain the historical task representation position and cross-modal dependency structure, effectively suppressing catastrophic forgetting during the continuous learning process of visual language action models.
Smart Images

Figure CN122769969A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a robot control method, and relates to the fields of artificial intelligence technology, robot control technology, and machine learning technology, specifically to a continuous learning robot control method based on information theory constraints. Background Technology
[0002] Current Vision-Language-Action (VLA) models can jointly process visual observations, language commands, and action sequences, demonstrating good generalization ability in complex tasks such as robot grasping, placement, opening and closing drawers, and knob operation. However, as robot applications expand from fixed tasks to open environments, models need to possess continuous or lifelong learning capabilities, that is, retaining learned task abilities while continuously receiving new tasks. For VLA models, this problem is more complex than for unimodal models because the model must not only maintain action prediction capabilities but also maintain cross-modal correspondences between visual targets, language semantics, and action generation. In practice, models are prone to catastrophic forgetting when continuously trained in task order, causing a rapid decline in performance on old tasks; more seriously, the original information alignment relationships between multiple modalities will also gradually degrade.
[0003] Existing methods for mitigating forgetting mainly include experience replay, parameter regularization, and structural extension. Experience replay methods train by mixing a small number of historical samples with new task samples to approximate the distribution of historical data. However, these methods primarily recover the marginal distribution of historical samples, making it difficult to maintain fine-grained cross-modal relationships between vision, language, and action, thus offering limited relief for structural forgetting in visual-language-action models. Parameter regularization methods, such as those based on parameter importance constraints, reduce the risk of old knowledge being overwritten by limiting changes in important parameters. However, these methods are often too rigid and struggle to adapt to non-stationary semantic changes in visual-language-action models in new tasks, thus affecting their adaptability to new tasks. Structural extension methods reduce parameter interference by adding extra modules, employing task isolation, or expert decomposition. While they can reduce conflict to some extent, they often lead to increased parameter size, higher computational costs, and insufficient shared representations between tasks, hindering the unified modeling and generalization of robot policies in open environments.
[0004] In recent years, several specialized methods have emerged for the continuous learning of visual language action models. These methods include knowledge evolution, knowledge-driven routing, context-aware meta-training, and target conditional value constraints, aimed at improving the model's ability to continuously adapt to new tasks. While these methods have shown some effectiveness in task routing, parameter utilization, and policy adaptation, their primary focus remains on sample replay, context adaptation, or value optimization. They have not yet explicitly maintained the cross-modal dependencies between visual observations, language commands, and action outputs from an information structure perspective. Meanwhile, some contrastive learning methods for visual language action models have been developed, primarily to enhance control-related representations or improve action retrieval capabilities. However, these methods are typically geared towards static or single-stage training scenarios and do not address the issue of cross-modal relationship degradation over time during continuous learning.
[0005] For visual-language action models, catastrophic forgetting manifests not only as a decrease in the success rate of old tasks but also as the diffusion and drift of cross-modal attention and representation structures. Specifically, after training on a particular task, the model can typically focus its attention relatively closely on the target object, the manipulation area, and the vicinity of the end effector related to the language command. However, upon continuing training on subsequent tasks, the cross-modal attention response for the same task often diffuses to irrelevant background areas, the response on the target object weakens, and the original conditional dependency between language and vision is gradually disrupted. This structural degradation further propagates to the action generation process, leading to unstable action prediction, decreased ability to retain historical tasks, and reduced robustness of robot control.
[0006] Therefore, the technical problem that existing technologies need to solve is: how to effectively maintain the cross-modal information structure formed by historical tasks while adapting to new tasks during the continuous learning process of visual language action models. Specifically, this includes maintaining the stability of historical sample representation anchors, maintaining the dependency between visual representations and language semantics, suppressing cross-modal attention diffusion and semantic drift, and achieving a balance between maintaining old tasks and learning new tasks, thereby improving the robot's task retention ability, generalization ability and control stability in continuous learning scenarios. Summary of the Invention
[0007] To address the problems existing in the background art, this invention provides a continuous learning robot control method based on information theory constraints. This method solves the problems of severe catastrophic forgetting, easy degradation of cross-modal alignment relationships, insufficient historical task retention, and decreased robot control stability in continuous learning scenarios of existing visual language action models.
[0008] The technical solution adopted in this invention is:
[0009] The information-theory-constrained continuous learning robot control method of the present invention includes:
[0010] The first step is to obtain the robot's initial task dataset and at least one incremental task dataset. The incremental task dataset consists of new task data introduced in the order of tasks during continuous learning. At least one representative trajectory sample is selected from the initial task dataset and written into a pre-established playback memory. The initial task dataset is then input into the pre-trained visual language action model (VLA) for training. The trained initial visual language action model is used as the student model. The model parameters of the initial visual language action model are frozen and used as the teacher model.
[0011] The second step involves constructing a joint batch dataset from the current incremental task dataset and representative trajectory samples from the replay memory, which are then input into the teacher model and student model for processing. The teacher model outputs teacher anchor representations and old cross-modal representations, while the student model outputs student representations and new cross-modal representations. Based on the teacher anchor representations and student representations, a replay anchor contrastive learning loss based on information theory constraints is constructed, and a cross-modal mutual information preservation loss is constructed based on the old cross-modal representations and new cross-modal representations.
[0012] The third step involves constructing a total loss function based on the replay anchor point comparison learning loss and the cross-modal mutual information preservation loss. The total loss function is minimized to iteratively update the model parameters of the student model until the preset number of iterations is reached or the total loss function converges, thus obtaining the currently trained student model.
[0013] The fourth step involves selecting at least one representative trajectory sample from the current incremental task dataset and writing it into the playback memory. The currently trained student model is then used as the student model for the next incremental task dataset. The model parameters of the currently trained student model are frozen and used as the teacher model for the next incremental task dataset. The same operations from the second to the fourth step are repeated until all incremental task datasets have been processed, and the final trained student model is obtained.
[0014] The fifth step is to deploy the finally trained student model into the robot's control system. When the robot performs a task, the current task data is input, processed, and the action sequence within the future time window is output to control the robot's end effector to perform the target operation task.
[0015] In the first step, both the initial task dataset and the incremental task dataset include several task samples. Both the task samples and the task data include at least the robot's multimodal observation information and language commands. The multimodal observation information includes robot body information and visual images.
[0016] In the first step, the visual-language-action model includes a language encoding network, a visual encoding network, an ontology information encoding network, a cross-modal fusion module based on cross-attention, and an action generation module. The language encoding network, visual encoding network, and ontology information encoding network process visual images, language commands, and robot ontology information, respectively, and output visual features, language features, and ontology features. Then, after processing by the cross-modal fusion module, fused features are output, and finally, after processing by the action generation module, action sequences are output. The teacher anchor representation is obtained by forward computation of samples using the frozen teacher model. Based on the cross-modal fusion modules corresponding to the teacher model and student model, teacher anchor representation and student representation are output, respectively. The fused features output by the cross-modal fusion layers corresponding to the teacher model and student model are then processed by a projection function to output old cross-modal representation and new cross-modal representation, respectively.
[0017] In the second step, during incremental task training, when the visual language action model is used as the teacher model, the teacher model remains completely frozen; when the visual language action model is used as the student model, the ontology information encoding network of the student model and... Action generation module and The model parameters of the cross-modal fusion module are not frozen and are updated iteratively, while the model parameters of the language encoding network and the visual encoding network are kept frozen.
[0018] In the third step, a replay anchor point contrastive learning loss based on information theory constraints is used. as follows:
[0019]
[0020] in, This represents the set of representative trajectory samples currently stored in the playback memory. This represents the current joint batch dataset; This indicates the number of representative trajectory samples in the current playback memory; This represents an exponential function with the natural constant e as its base. Indicates vector similarity calculation; and Let represent the student representation and teacher anchor point representation extracted by the student model and teacher model for the i-th representative trajectory sample in the current playback memory, respectively. This represents the teacher anchor representation extracted by the teacher model for the j-th sample in the current joint batch dataset; This represents the temperature parameter.
[0021] In the third step, cross-modal mutual information preservation loss Including mutual information maximization terms and edge consistency regularization ,as follows:
[0022]
[0023]
[0024]
[0025]
[0026] in, Represents mutual information calculation; and These represent the new cross-modal representations and the old cross-modal representations generated by the student model and the teacher model, respectively. Indicates the Kullback-Leibler divergence; This indicates a joint distribution, where the inputs are concatenated before being fed into a multilayer perceptron for processing. It represents a marginal distribution, which processes its own input through a shared projection function.
[0027] In the third step, the total loss function is as follows:
[0028]
[0029] in, λ1 and λ2 represent the basic behavior learning loss; λ1 and λ2 represent the weighting coefficients of the first and second losses. RAC This indicates the learning loss compared to the playback anchor point; CMI This indicates the loss in cross-modal mutual information retention.
[0030] The aforementioned basic behavioral learning loss as follows:
[0031]
[0032] in, Indicates Gaussian noise; Representing moments in the future time domain to The sequence of actions within; This indicates the action generation module. Indicates an action, Indicates network parameters; Representing a moment in the future time domain to Action sequence within Apply Gaussian noise The action sequence is then obtained through the τth denoising step of the flow matching denoising process; τ represents the flow matching time parameter. Multimodal observation information at time t Represents language instructions.
[0033] The electronic device of the present invention includes: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor invokes the program data to execute the method described above.
[0034] The present invention provides a computer-readable storage medium having program data stored thereon, which, when executed by a processor, implements the method described above.
[0035] The beneficial effects of this invention are:
[0036] 1. This invention, through the dual constraints of preserving historical anchor points and preserving cross-modal mutual information, can simultaneously maintain the historical task representation position and cross-modal dependency structure, effectively suppressing catastrophic forgetting during the continuous learning process of visual language action models.
[0037] 2. This invention maintains the ability to perform historical tasks while retaining the ability to adapt to new tasks, thereby achieving a better balance between stability and adaptability.
[0038] 3. This invention can reduce the drift in the alignment relationship between vision, language and action, improve the accuracy of action prediction and the control stability, generalization ability and robustness of the robot in continuous tasks.
[0039] 4. This invention is not only applicable to visual language action models based on flow matching, but can also be extended to other robot policy learning models with multimodal encoders and action decoders, demonstrating good versatility. Attached Figure Description
[0040] Figure 1 This is an overall flowchart of the method of the present invention;
[0041] Figure 2 This is a schematic diagram illustrating the degradation of cross-modal information structure during continuous learning.
[0042] Figure 3 A schematic diagram of the experimental setup for evaluating the continuous learning performance of the LIBERO benchmark suite in the Long and Goal categories, as well as the performance of three real-world Piper grasping and placing tasks;
[0043] Figure 4 This is a schematic diagram illustrating the forgetting rate of Long under the continuous learning LIBERO benchmark using the method of the present invention;
[0044] Figure 5 This is a performance comparison chart of real-world pick-up and place tasks under external disturbances. Detailed Implementation
[0045] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] like Figure 1 As shown, the continuous learning robot control method based on information theory constraints of the present invention is as follows:
[0047] The first step is to obtain the robot's initial task dataset and at least one incremental task dataset. The incremental task dataset is new task data introduced in the order of tasks during continuous learning. At least one representative trajectory sample is selected from the initial task dataset and written into a pre-established playback memory. The initial task dataset is input into the pre-trained visual language action model (VLA) for training. The trained initial visual language action model is used as the student model. The model parameters of the initial visual language action model are frozen and used as the teacher model.
[0048] Both the initial task dataset and the incremental task dataset include several task samples. Each task sample and task data includes at least multimodal observation information and language commands from the robot. The multimodal observation information includes robot body information and visual images. The robot body information includes one or more of the following: joint angles, gripper state, end effector pose, and chassis speed. The action sequence includes one or more of the following: robot joint control variables, end effector pose increments, gripper opening / closing commands, or chassis motion commands. Visual images are captured by the robot's main-view camera and one or two cameras at the robot's actuator end. The language commands are text-based control commands, including one or more of the following operations: grasping, placing, opening, closing, or rotating.
[0049] The replay memory is used to store historical task samples. Initially, the replay memory stores at least one representative trajectory sample from the initial task dataset, and each historical task dataset stores at least one representative trajectory sample. The representative trajectory sample is selected from the initial task dataset by any of the following methods: random selection, cluster center selection, or selection based on task coverage.
[0050] The visual-language-action model includes a language encoding network, a visual encoding network, an ontology information encoding network, a cross-modal fusion module based on cross-attention, and an action generation module. The language encoding network, visual encoding network, and ontology information encoding network process visual images, language commands, and robot ontology information, respectively, and output visual features, language features, and ontology features. These features are then processed by the cross-modal fusion module to output fused features, and finally processed by the action generation module to output an action sequence. The teacher anchor representation is obtained by forward computation of samples using a frozen teacher model. The cross-modal fusion modules corresponding to the teacher model and student model output teacher anchor representations and student representations, respectively. The fused features output by the cross-modal fusion layers corresponding to the teacher model and student model are then processed by a projection function to output old cross-modal representations and new cross-modal representations, respectively.
[0051] The second step involves constructing a joint batch dataset from the current incremental task dataset and representative trajectory samples from the replay memory, which are then input into the teacher model and student model for processing. The teacher model outputs teacher anchor representations and old cross-modal representations, while the student model outputs student representations and new cross-modal representations. Based on the teacher anchor representations and student representations, a replay anchor contrastive learning loss based on information theory constraints is constructed to maintain the structural position of historical samples in the representation space. Based on the old cross-modal representations and new cross-modal representations, a cross-modal mutual information retention loss is constructed to achieve the goal of continuous learning and anti-forgetting.
[0052] During incremental task training, when the visual language action model is used as the teacher model, the teacher model remains completely frozen; that is, the language encoding network, visual encoding network, ontology information encoding network, cross-modal fusion module, and action generation module are all frozen and used only to generate teacher anchor representations and old cross-modal representations. When the visual language action model is used as the student model, the student model's ontology information encoding network and... Action generation module and The model parameters of the cross-modal fusion module are not frozen and are updated iteratively, while the model parameters of the language encoding network and the visual encoding network are kept frozen. This reduces training costs while maintaining the stability of the backbone representation, achieving continuous learning and stability against forgetting.
[0053] The third step involves constructing a total loss function based on the replay anchor point contrastive learning loss and the cross-modal mutual information preservation loss. This total loss function is minimized to iteratively update the student model's parameters until a preset number of iterations is reached or the total loss function converges, thus obtaining the currently trained student model. This is achieved using information-theoretic constraints on the replay anchor point contrastive learning loss. as follows:
[0054]
[0055] in, This represents the set of representative trajectory samples currently stored in the playback memory. This represents the current joint batch dataset; This indicates the number of representative trajectory samples in the current playback memory; This represents an exponential function with the natural constant e as its base. Indicates vector similarity calculation; and Let represent the student representation and teacher anchor point representation extracted by the student model and teacher model for the i-th representative trajectory sample in the current playback memory, respectively. This represents the teacher anchor representation extracted by the teacher model for the j-th sample in the current joint batch dataset; This represents the temperature parameter.
[0056] Cross-modal mutual information retention loss Including mutual information maximization terms and edge consistency regularization ,as follows:
[0057]
[0058]
[0059]
[0060]
[0061] in, Represents mutual information calculation; and These represent the new cross-modal representations and the old cross-modal representations generated by the student model and the teacher model, respectively. Indicates the Kullback-Leibler divergence; This indicates a joint distribution, where the inputs are concatenated before being fed into a multilayer perceptron for processing. It represents a marginal distribution and processes its own input through a shared projection function to constrain the global statistical consistency of the new and old cross-modal representation distributions.
[0062] By maximizing the mutual information between old and new cross-modal representations and constraining the global statistical consistency of the old and new distributions through edge consistency regularization terms to maintain the statistical dependency structure between vision, language, and action, continuous learning resists forgetting by preserving knowledge of old tasks during the learning of new tasks.
[0063] The total loss function is as follows:
[0064]
[0065] in, λ1 and λ2 represent the basic behavior learning loss, which is the flow matching action prediction loss; λ1 and λ2 represent the weight coefficients of the first and second losses, and λ1>0 and λ2>0. RAC This indicates the learning loss compared to the playback anchor point; CMI This indicates the loss in cross-modal mutual information retention.
[0066] Basic behavioral learning loss as follows:
[0067]
[0068] in, Indicates Gaussian noise; Representing moments in the future time domain to The sequence of actions within; This indicates the action generation module. Indicates an action, Indicates network parameters; Representing a moment in the future time domain to Action sequence within Apply Gaussian noise The action sequence is then obtained through the τth denoising step of the flow matching denoising process; τ represents the flow matching time parameter. Multimodal observation information at time t Represents language instructions.
[0069] The fourth step involves selecting at least one representative trajectory sample from the current incremental task dataset and writing it into the playback memory. The currently trained student model is then used as the student model for the next incremental task dataset. The model parameters of the currently trained student model are frozen and used as the teacher model for the next incremental task dataset. The same operations from the second to the fourth step are repeated. The student model continues to learn the basic learning optimization objective and the anti-forgetting optimization objective until all incremental task datasets have been processed, and the final trained student model is obtained.
[0070] The fifth step is to deploy the finally trained student model into the robot's control system. When the robot performs a task, the current task data is input, processed, and the action sequence within the future time window is output to control the robot's end effector to perform the target operation task.
[0071] The optimization objective of continuous learning basic learning corresponds to the basic behavioral learning loss. The optimization objective of continuous learning anti-forgetting includes anchor-contrast learning and cross-modal mutual information structure stability. Among them, anchor-contrast learning constrains the student representation to be close to the teacher anchor, and cross-modal mutual information structure stability constrains the information dependency relationship between visual representation, text representation and action representation.
[0072] like Figure 2 As shown, this is a visualization of cross-modal structural degradation. After training on the new task, the red box shows that visual text attention has significantly diffused and drifted compared to the previous task.
[0073] To verify the effectiveness of the method of the present invention, the following comparison method is set up:
[0074] Multitask method: Jointly train on all tasks together, serving as an upper bound for continuous learning.
[0075] The Sequential method: directly fine-tunes according to the task order, without employing a forgetting mitigation mechanism.
[0076] The EWC (Elastic Weight Consolidation) method employs a parameter regularization method based on the Fisher information matrix.
[0077] ER (Experience Replay) method: The experience replay method is used to replay one trajectory for each historical task.
[0078] The info-VLA method employs the continuous learning robot control method based on information theory constraints proposed in this invention.
[0079] In practical implementation, this invention performs LIBERO-Long continuous learning performance verification, aiming to verify the method's ability to retain historical tasks and adapt to new tasks in long-term robot continuous learning tasks. For example... Figure 3 As shown, the LIBERO-Long task suite from the LIBERO benchmark was used, with a pre-trained pi0.5 as the base visual-language action model. The experimental configuration employed a B5-5N1 continuous learning setup, initially training with five basic tasks (task 0 in Table 1), followed by the introduction of five incremental tasks (tasks 1-5 in Table 1), with one new task introduced each time. For historical tasks, one trajectory from each task was randomly saved in the replay memory. During training, only the action head parameters were updated; the visual-language encoding backbone remained frozen.
[0080] Table 1. Success rate of continuous learning on the LIBERO-Long B5-5N1 benchmark
[0081]
[0082] As shown in Table 1, the method of this invention achieved the optimal success rate for old tasks and the overall success rate in each incremental stage. The Multitask method, trained jointly on the incremental task dataset and the basic task dataset, represents the theoretical upper limit of accuracy. The Sequential method, trained jointly on all tasks, directly fine-tunes the task order without employing a forgetting mitigation mechanism, and represents the lower limit of accuracy. Compared to EWC, the method of this invention has significant advantages. Compared to the best-performing baseline ER, the average task accuracy of the method of this invention increases from 72.0% to 78.7%, indicating that the method of this invention can more effectively suppress catastrophic forgetting during continuous learning and maintain the ability to learn new tasks.
[0083] In its specific implementation, this invention verifies the continuous learning performance of the LIBERO-Goal method, aiming to validate its continuous learning performance under conditions without prior task warm-up. For example... Figure 3 As shown, the LIBERO-Goal task suite from the LIBERO benchmark was used. The experimental configuration adopted the B0-5N1 setting, that is, continuous learning started directly from the pre-trained base model, and 5 tasks were introduced in sequence (i.e., tasks 1-5 in Table 2). The average accuracy of all trained tasks was tested after each task.
[0084] Table 2 Continuous Learning Success Rate on the LIBERO-Goal B0-5N1 Benchmark
[0085]
[0086] As shown in Table 2, the method of this invention still achieves the best results in more challenging continuous learning scenarios without a basic task. Compared with ER, the average task accuracy of the method of this invention is improved from 67.2% to 73.3%, further demonstrating that the method of this invention can stably maintain the cross-modal dependency structure between vision, language, and action in open task sequences.
[0087] In its specific implementation, this invention performs LIBERO-Long continuous learning forgetting verification to demonstrate the anti-forgetting ability of the method under continuous learning. For example... Figure 3 As shown, the LIBERO-Long task suite from the LIBERO benchmark was used. The experimental configuration adopted the B5-5N1 setting, that is, initial training was performed using 5 basic tasks, and then 5 incremental tasks were introduced sequentially. The average accuracy of all trained tasks was tested after each task. Figure 4As shown, the forgetting values of each method differ significantly as tasks 1 through 5 progress: the forgetting values of EWC and Sequential methods remain high, indicating the most severe forgetting; the forgetting value of ER method is at a moderate level, reaching its highest in task 5; while the forgetting value of the info-VLA method of this invention remains at an extremely low level, significantly better than the other three methods, demonstrating its excellent anti-forgetting performance, effectively preserving the learning results of previous tasks, and solving the forgetting problem in multi-task sequence learning.
[0088] This invention compares the performance of a real-world pick-and-place task under external disturbances during specific implementations, aiming to verify the effectiveness and robustness of the method's continuous learning in the real physical world. This is conducted on a real Piper robotic arm platform, such as... Figure 3 As shown, three different grasping and placing tasks were set up (block task, sachet task, and banana task). During the experiment, multiple external physical perturbations were applied to the objects to be manipulated to test the robustness of the model. The model needed to learn each task progressively: first the block task, then the sachet task, and finally the banana task. The success rate is as follows: Figure 5 As shown, the method of this invention maintains better performance than the baseline ER method throughout the entire continuous learning process. Real-world experimental results verify that the information-theory-constrained visual language action model continuous learning method of this invention enables robots to have better historical task retention, disturbance resistance, and control robustness in real-world environments.
[0089] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of this application can be implemented using various computer languages. This application is described with flowcharts of methods, systems, and computer program products according to embodiments of this application.
[0090] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, this invention is intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.
[0091] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if these modifications and variations of this application fall within the scope of the equivalent technology of this invention, this application also intends to include these modifications and variations.
Claims
1. A continuous learning robot control method based on information theory constraints, characterized in that, include: The first step is to obtain the robot's initial task dataset and incremental task dataset, select representative trajectory samples from the initial task dataset and write them into the playback memory; The initial task dataset is input into the visual language action model for training. The trained initial visual language action model is used as the student model, and the initial visual language action model is frozen and used as the teacher model. The second step involves constructing a joint batch dataset from the current incremental task dataset and the representative trajectory samples in the replay memory, which are then input into the teacher model and student model for processing. The teacher model outputs teacher anchor representations and old cross-modal representations, while the student model outputs student representations and new cross-modal representations. Based on the teacher anchor representations and student representations, a replay anchor contrastive learning loss based on information theory constraints is constructed, and a cross-modal mutual information preservation loss is constructed based on the old cross-modal representations and new cross-modal representations. The third step involves constructing a total loss function based on the replay anchor point comparison learning loss and cross-modal mutual information preservation loss. The total loss function is minimized to iteratively update the student model until the preset number of iterations is reached or the total loss function converges, thus obtaining the currently trained student model. The fourth step is to select representative trajectory samples from the current incremental task dataset and write them into the playback memory. The currently trained student model is used as the student model for the next incremental task dataset. The currently trained student model is frozen and used as the teacher model for the next incremental task dataset. The same operation in the second to fourth steps is repeated until all incremental task datasets are processed and the final trained student model is obtained. The fifth step is to deploy the finally trained student model into the robot's control system. When the robot performs a task, the current task data is input, processed, and the action sequence is output to control the robot to perform the target operation task.
2. The continuous learning robot control method based on information theory constraints according to claim 1, characterized in that: In the first step, both the initial task dataset and the incremental task dataset include several task samples. Both the task samples and the task data include the robot's multimodal observation information and language commands. The multimodal observation information includes the robot's body information and visual images.
3. The continuous learning robot control method based on information theory constraints according to claim 2, characterized in that: In the first step, the visual-language-action model includes a language encoding network, a visual encoding network, an ontology information encoding network, a cross-modal fusion module, and an action generation module. The language encoding network, visual encoding network, and ontology information encoding network process visual images, language commands, and robot ontology information, respectively, and output visual features, language features, and ontology features. Then, after processing by the cross-modal fusion module, fused features are output. Finally, after processing by the action generation module, action sequences are output. The cross-modal fusion module based on the teacher model and the student model outputs teacher anchor representations and student representations, respectively. The fused features output by the cross-modal fusion layer based on the teacher model and the student model are then processed by a projection function to output old cross-modal representations and new cross-modal representations, respectively.
4. The continuous learning robot control method based on information theory constraints according to claim 3, characterized in that: In the second step, when the visual language action model is used as the teacher model, the teacher model remains completely frozen; when the visual language action model is used as the student model, the ontology information encoding network, action generation module, and cross-modal fusion module of the student model are not frozen and are iteratively updated, while the language encoding network and visual encoding network remain frozen.
5. The continuous learning robot control method based on information theory constraints according to claim 1, characterized in that: In the third step, a replay anchor point contrastive learning loss based on information theory constraints is used. as follows: in, This represents the set of representative trajectory samples currently stored in the playback memory. This represents the current joint batch dataset; This indicates the number of representative trajectory samples in the current playback memory; This represents an exponential function with the natural constant e as its base. Indicates vector similarity calculation; and Let represent the student representation and teacher anchor point representation extracted by the student model and teacher model for the i-th representative trajectory sample in the current playback memory, respectively. This represents the teacher anchor representation extracted by the teacher model for the j-th sample in the current joint batch dataset; This represents the temperature parameter.
6. The continuous learning robot control method based on information theory constraints according to claim 1, characterized in that: In the third step, cross-modal mutual information preservation loss Including mutual information maximization terms and edge consistency regularization ,as follows: in, Represents mutual information calculation; and These represent the new cross-modal representations and the old cross-modal representations generated by the student model and the teacher model, respectively. Indicates the Kullback-Leibler divergence; This indicates a joint distribution, where the inputs are concatenated before being fed into a multilayer perceptron for processing. It represents a marginal distribution, which processes its own input through a shared projection function.
7. The continuous learning robot control method based on information theory constraints according to claim 1, characterized in that: In the third step, the total loss function is as follows: in, λ1 and λ2 represent the basic behavior learning loss; λ1 and λ2 represent the weighting coefficients of the first and second losses. RAC This indicates the learning loss compared to the playback anchor point; CMI This indicates the loss in cross-modal mutual information retention.
8. The continuous learning robot control method based on information theory constraints according to claim 7, characterized in that: The aforementioned basic behavioral learning loss as follows: in, Indicates Gaussian noise; Representing moments in the future time domain to The sequence of actions within; This indicates the action generation module; Representing a moment in the future time domain to Action sequence within Apply Gaussian noise The action sequence is then obtained through the τth denoising step of the flow matching denoising process; τ represents the flow matching time parameter. Multimodal observation information at time t Represents language instructions.
9. An electronic device, characterized in that, include: A memory and a processor are coupled to each other, wherein the memory stores program data, and the processor invokes the program data to perform the method as described in any one of claims 1-8.
10. A computer-readable storage medium storing program data thereon, characterized in that, When the program data is executed by the processor, it implements the method as described in any one of claims 1-8.