Enhanced fine tuning method and device for visual language action model, equipment and medium
Through a phased reinforcement learning method, combined with offline and online learning, the problem of low fine-tuning efficiency of visual language action models in actual environments was solved, and the stability and reliability of robot operation were improved.
Patent Information
- Application Number
- CN202510826694.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-16
AI Technical Summary
The fine-tuning methods of existing visual language action models in robot operation tasks are limited by the long training cycle and high interaction cost in the physical environment, which makes it difficult to meet task requirements in actual perception and action, and reduces the reliability of robot operation.
A phased reinforcement learning method is adopted, including offline reinforcement learning and online reinforcement learning. The visual language action model is fine-tuned by collecting demonstration data, and interacts in the actual environment to obtain exploration trajectories and environmental feedback. The task reset strategy is used to automatically restore the environment state to avoid manual intervention.
It improves the fine-tuning effect of the visual language action model, ensures the stability and reliability of the robot's operation, and achieves non-intervention automatic optimization through interaction with the real environment, thereby improving the robot's performance in actual tasks.
Smart Images

Figure CN120656039A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and medium for enhancing and fine-tuning a visual language action model. Background Art
[0002] The visual-language-action model is a multimodal model that combines visual information, language understanding, and action generation. With its ability to integrate visual information, language instructions, and action decisions, the visual-language-action model has broad application prospects in robotic manipulation tasks, enabling more intelligent and flexible robotic manipulation in a variety of fields.
[0003] In the financial sector, robots based on visual language action models can serve as intelligent customer service assistants, providing a variety of services. For example, they can recognize customer expressions using visual perception, conduct conversations with customers using voice recognition and natural language processing technologies, and then execute corresponding actions. These can include guiding customers to designated areas or helping them hand over items through joint movements.
[0004] In the field of healthcare, robots based on visual language motion models can be used as auxiliary tools for rehabilitation training, helping patients perform complex rehabilitation movement training. For example, robots can monitor patients' body movements and expressions in real time through high-precision cameras, communicate with patients through voice, and obtain real-time feedback information to perform corresponding movement guidance or movement assistance, etc.
[0005] To optimize the performance of visual language action models in robotic manipulation tasks, fine-tuning is often required to improve model performance. However, current fine-tuning methods are limited by factors such as long training cycles and high interaction costs in physical environments. These methods are typically performed in synthetic or simulated environments, lacking the ability to model real-world interactions. In robotic manipulation tasks requiring real-world perception and action, this fine-tuning approach fails to meet task requirements, reducing the reliability of the robot's operation. Summary of the Invention
[0006] In view of the above-mentioned deficiencies in the prior art, the purpose of the present invention is to provide an enhanced fine-tuning method, device, equipment and medium for visual language action models that can be applied to the medical field, financial technology or other related fields. Its main purpose is to improve the fine-tuning effect of the visual language action model and ensure the stability and reliability of the robot operation.
[0007] The technical solutions of the present invention are as follows:
[0008] A first aspect of the present invention provides a method for enhancing and fine-tuning a visual language action model, comprising:
[0009] Loading a visual language action model to be fine-tuned, wherein the visual language action model is used to output action task instructions based on visual information and language instructions to operate the robot to perform the corresponding action task;
[0010] Collecting multiple pieces of demonstration data, and performing offline reinforcement learning on the visual language action model using the demonstration data to obtain an offline fine-tuning model;
[0011] Deploying the offline fine-tuning model into a real environment, controlling the robot to interact with the environment according to the task reset strategy, and obtaining corresponding exploration trajectories and environmental feedback;
[0012] The offline fine-tuning model is subjected to online reinforcement learning according to the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model.
[0013] A second aspect of the present invention provides a device for enhancing and fine-tuning a visual language action model, comprising:
[0014] A model loading module is used to load the visual language action model to be fine-tuned, wherein the visual language action model is used to output action task instructions based on visual information and language instructions to operate the robot to perform the corresponding action task;
[0015] An offline reinforcement fine-tuning module is used to collect multiple demonstration data and perform offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model;
[0016] An online interaction module is used to deploy the offline fine-tuning model into the actual environment, control the robot to interact with the environment according to the task reset strategy, and obtain corresponding exploration trajectory and environmental feedback;
[0017] The online reinforcement fine-tuning module is used to perform online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model.
[0018] A third aspect of the present invention provides a computer device comprising at least one processor; and
[0019] a memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned enhanced fine-tuning method of the visual language action model.
[0021] A fourth aspect of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the one or more processors can execute the above-mentioned enhanced fine-tuning method of the visual language action model.
[0022] Beneficial effects: The present invention discloses a method, device, equipment and medium for enhancing fine-tuning of a visual language action model. Compared with the prior art, the embodiment of the present invention loads a visual language action model to be fine-tuned, and the visual language action model is used to output action task instructions based on visual information and language instructions to operate a robot to perform a corresponding action task; collects multiple demonstration data, and performs offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model; deploys the offline fine-tuning model into an actual environment, controls the robot to interact with the environment according to a task reset strategy, and obtains corresponding exploration trajectories and environmental feedback; performs online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model. By performing offline reinforcement learning and online reinforcement learning in stages, and automatically restoring the environment state through the task reset strategy in the online learning stage, manual intervention is avoided, so that on the basis of quickly improving the performance of the initial action strategy, non-intervention automatic optimization can also be performed through interaction with the real environment, realizing collaborative model fine-tuning to improve the stability and reliability of the robot operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the solutions in the present invention, a brief introduction is given below to the drawings required for use in describing the embodiments of the present invention. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 A schematic diagram of an application environment for the enhanced fine-tuning method of the visual language action model provided by an embodiment of the present invention;
[0025] Figure 2 A flow chart of a method for enhancing and fine-tuning a visual language action model provided by an embodiment of the present invention;
[0026] Figure 3 A flowchart of step S202 in the method for enhancing and fine-tuning a visual language action model provided by an embodiment of the present invention;
[0027] Figure 4 A flowchart of step S304 in the method for enhancing and fine-tuning a visual language action model provided by an embodiment of the present invention;
[0028] Figure 5 A flowchart of step S203 in the method for enhancing and fine-tuning a visual language action model provided by an embodiment of the present invention;
[0029] Figure 6 A flowchart of step S502 in the method for enhancing and fine-tuning a visual language action model provided by an embodiment of the present invention;
[0030] Figure 7 A schematic diagram of the functional modules of the device for enhancing and fine-tuning a visual language action model provided by an embodiment of the present invention;
[0031] Figure 8 A schematic diagram of the hardware structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0032] To make the objectives, technical solutions, and effects of the present invention more clear and distinct, the present invention is further described in detail below. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. The embodiments of the present invention are described below with reference to the accompanying drawings.
[0033] The enhanced fine-tuning method of the visual language action model provided by the embodiment of the present invention can be applied to Figure 1 In an application environment, the system includes a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0034] The user may use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0035] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0036] The server 105 may be a server that provides various services, such as a backend server that provides support for the content browsed by the user using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (for example only). The backend server may analyze and process the received user requests and other data, and feed back the processing results (such as web pages, information, or data obtained or generated according to the user request) to the terminal device. The server 105 may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server 105 may also be a server for a distributed system, or a server combined with a blockchain.
[0037] It should be noted that the method for enhancing and fine-tuning the visual language action model provided in the embodiments of the present application can generally be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the apparatus for enhancing and fine-tuning the visual language action model provided in the embodiments of the present invention can also be provided in the first terminal device 101, the second terminal device 102, or the third terminal device 103. Alternatively, the method for enhancing and fine-tuning the visual language action model provided in the embodiments of the present invention can generally be executed by the server 105. Accordingly, the apparatus for enhancing and fine-tuning the visual language action model provided in the embodiments of the present invention can generally be provided in the server 105.
[0038] It should be understood that the numbers of the above terminal devices, networks and servers are merely illustrative and any number of terminal devices, networks and servers may be provided as required.
[0039] like Figure 2 As shown, the enhanced fine-tuning method of the visual language action model provided by the embodiment of the present invention specifically includes the following steps:
[0040] S201 . Loading a visual language action model to be fine-tuned, wherein the visual language action model is used to output action task instructions based on visual information and language instructions to operate a robot to perform a corresponding action task.
[0041] In this embodiment, a pre-trained visual language action model is loaded as the visual language action model to be fine-tuned. The visual language action model to be fine-tuned can process visual input (such as images or videos) and language input (such as natural language instructions), and output action task instructions based on visual information and language instructions to operate the robot to perform corresponding action tasks, such as "moving to the target position" or "grasping an object" and so on.
[0042] Specifically, the visual language action model can adopt a Transformer-based architecture, including a visual feature extraction module, a language understanding module, a multimodal fusion module, and an action generation module. The visual feature extraction module and the language understanding module process the visual and language inputs respectively to obtain corresponding visual feature vectors and language feature vectors. The visual feature vectors and language feature vectors are input into the multimodal fusion module, and the visual features and language features are fused through a Transformer layer to obtain a fused feature vector. The fused feature vector is mapped to the action space through a fully connected network (FCN) in the action generation module to generate corresponding action instructions (such as "move to position (x, y)", "grab the object", "release the object", etc.).
[0043] For example, in the financial field, the visual language action model can be used to control the robot to generate corresponding action instructions based on the customer's language instructions (such as "I want to check my account balance") and visual information (such as recognizing the customer's identity and expression), guide the customer to the self-service terminal and assist in completing the operation.
[0044] In the field of medical health, visual language action models can be used to control robots to generate action instructions based on the rehabilitation therapist's language instructions (such as "Please ask the patient to perform arm stretching exercises") and visual information (such as the patient's body position and status), assisting patients in completing rehabilitation training.
[0045] S202: Collect multiple pieces of demonstration data, and perform offline reinforcement learning on the visual language action model using the demonstration data to obtain an offline fine-tuning model.
[0046] In this embodiment, multiple pieces of demonstration data are collected, for example, 30 or 50 pieces of demonstration data based on the fine-tuning requirements of a specific task. This demonstration data covers the key steps and common scenarios of the task to ensure that the model can learn effective strategies. Specifically, the demonstration data may include demonstration visual information, such as images or video sequences captured by a camera, which show the state of the environment during task execution; demonstration language instructions, such as natural language instructions related to the task; and demonstration motion trajectories, which are the actual standard movements of the robot when performing the action task. These movements are pre-recorded by experts (such as human operators) to demonstrate the correct way to perform the task.
[0047] The visual language model is trained offline using collected demonstration data. This means that without interacting with the real environment, the model is fine-tuned using a fixed dataset of demonstration data. This allows the model to directly learn behaviors similar to those in the demonstration data. Reinforcement learning is achieved through environmental feedback rewards based on the similarity between the robot's actual actions and the action instructions. Additional supervisory signals are provided based on the differences between actual actions and demonstration actions, allowing the model to accurately and efficiently learn a reasonable initial action strategy from a small amount of demonstration data during the offline phase. After completing offline fine-tuning, such as learning the set step size, the current model parameters are saved as the offline fine-tuned model.
[0048] For example, in the financial sector, when fine-tuning a visual language action model for an intelligent customer service robot offline, the model can collect multiple verbal commands from employees or customers, such as "Please insert your bank card," "Check your account balance," or "I want to transfer money," along with corresponding visual information, such as videos of customers using self-service terminals, and corresponding action sequences, such as the employee or customer's specific actions at the terminal, such as inserting a bank card, entering a password, or selecting a menu option. Through offline reinforcement learning, the model learns how to generate correct action commands based on these commands and visual information, such as guiding customers to correctly operate the self-service terminal.
[0049] In the healthcare field, offline fine-tuning of visual language action models for rehabilitation training robots can collect data from multiple demonstrations of rehabilitation therapists guiding patient training. This includes verbal instructions from the therapist or patient, such as "Please stretch your arms," "Relax your arms," and "Add a little resistance," visual information such as the patient's body position, expression, and movements, and corresponding action sequences, such as the patient's specific movements on rehabilitation equipment, including arm extension, flexion, and grasping. Through offline reinforcement learning, the visual language action model can learn how to initially generate correct action instructions based on these instructions and visual information, for example, to assist patients in completing the corresponding rehabilitation movements.
[0050] S203: deploying the offline fine-tuning model into an actual environment, controlling the robot to interact with the environment according to the task reset strategy, and obtaining corresponding exploration trajectories and environmental feedback.
[0051] In this embodiment, after completing the offline reinforcement learning phase, the offline fine-tuning model is deployed in a real-world environment, controlling the robot to interact with the environment. The robot's strategy and exploration behavior are then adjusted based on environmental feedback to achieve further model optimization. Specifically, when controlling the robot's interaction with the real environment by outputting action task instructions, a task reset strategy is employed to enable the robot to automatically return to its initial state after completing a task, allowing it to proceed with the next task. This eliminates the need for human intervention during the robot's autonomous exploration during the online fine-tuning phase, further improving training efficiency and reducing labor costs. The robot interacts with the environment according to the task reset strategy based on the action instructions generated by the model, collecting trajectory data generated when executing actions in the environment and the environment's reward feedback for the robot's actions, thereby obtaining corresponding exploration trajectories and environmental feedback. Through interaction with the real environment, the model learns strategies that adapt to the real environment, improving its adaptability to the real environment.
[0052] For example, in the financial sector, robots controlled by offline fine-tuned models are deployed in bank lobby areas. The models can generate action commands based on the customer's verbal instructions and visual expressions. The robots then interact with customers according to a task reset strategy based on the action task commands output by the model. This task reset strategy allows the robots to automatically return to their initial state after each service is completed, preparing for the next service. For example, after guiding a customer to complete an account query, the robot automatically returns to its initial position to prepare for the next customer. In the interactive process of continuously executing actions and receiving environmental feedback, the robot collects corresponding exploration trajectories and environmental feedback.
[0053] In healthcare, robots controlled by offline fine-tuned models are deployed in rehabilitation settings. Based on the motion task instructions output by the model, the robot interacts with the patient according to a task reset strategy, assisting the patient in rehabilitation training. This task reset strategy allows the robot to automatically return to its initial state after each training session, preparing for the next session. For example, after assisting a patient with a set of arm extension exercises, the robot automatically returns to its initial position, preparing to assist the next patient and receiving appropriate environmental feedback.
[0054] S204 , performing online reinforcement learning on the offline fine-tuning model according to the exploration trajectory, environmental feedback, and demonstration data to obtain a fine-tuned visual language action model.
[0055] In this embodiment, based on the exploration trajectory and environmental feedback obtained through interaction in the actual environment, the offline fine-tuning model is subjected to online reinforcement learning in combination with demonstration data. Specifically, in the online stage, the model controls the robot to interact autonomously with the real environment, and collects the corresponding exploration trajectory and environmental feedback. According to the environmental feedback, the value function is updated and the next exploration behavior is adjusted, thereby realizing online reinforcement learning for interaction with the real environment. In addition, in the reinforcement learning process, additional supervision signals are also provided based on the difference between the actual exploration trajectory and the demonstrated standard action to prevent the robot from performing actions far beyond the range, thereby obtaining a fine-tuned visual language action model. Through reinforcement learning in the online stage to further optimize the model's action strategy, the model can adapt to changes and complexity of the environment through autonomous exploration and learning in the real environment, thereby improving performance.
[0056] In the above embodiment, the present invention discloses a method for enhancing fine-tuning of a visual language action model, which comprises the following steps: loading a visual language action model to be fine-tuned, the visual language action model being used to output action task instructions based on visual information and language instructions to operate a robot to perform corresponding action tasks; collecting multiple demonstration data, and performing offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model; deploying the offline fine-tuning model into an actual environment, controlling the robot to interact with the environment according to a task reset strategy, and obtaining corresponding exploration trajectories and environmental feedback; performing online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback, and demonstration data to obtain a fine-tuned visual language action model. By performing offline reinforcement learning and online reinforcement learning in stages, and automatically restoring the environment state through the task reset strategy in the online learning stage, manual intervention is avoided, so that on the basis of rapidly improving the performance of the initial action strategy, non-intervention automatic optimization can also be performed through interaction with the real environment, thereby achieving collaborative model fine-tuning to improve the stability and reliability of the robot operation.
[0057] In one embodiment, Figure 3 As shown, step S202 includes:
[0058] S301, collecting multiple pieces of demonstration data and initializing the visual language action model;
[0059] S302: inputting the demonstration data into the initialized visual language action model and outputting corresponding offline action task instructions;
[0060] S303, controlling the robot to perform a corresponding offline action task according to the offline action task instruction, and obtaining an offline action trajectory;
[0061] S304, calculating the offline behavior cloning loss and the offline Q-value function loss based on the offline action trajectory, the standard action trajectory in the demonstration data, and the offline action task instruction to obtain an offline total loss;
[0062] S305 : Fine-tune the initialized visual-language-action model according to the offline total loss to obtain an offline fine-tuning model.
[0063] In this embodiment, during offline reinforcement learning, multiple pieces of demonstration data are first collected, such as video and audio recordings of bank employees interacting with customers, including the employee's verbal instructions, the customer's response, and the employee's operating steps, or video and audio recordings of a rehabilitation therapist interacting with a patient, including the therapist's verbal instructions, the patient's response, and the therapist's operating steps. By collecting multiple pieces of demonstration data, the model is ensured to learn diverse task execution strategies. A pre-trained visual-language-action model is loaded into the training environment to prepare for subsequent offline training. The demonstration data is input into the initialized visual-language-action model. The model extracts visual and linguistic features from the demonstration data and outputs corresponding offline action task instructions according to the current action strategy to control the robot to perform tasks in the offline environment.
[0064] In a simulation environment, the robot is controlled to perform actions according to the offline action task instructions generated by the model. The robot's action execution is recorded to obtain an offline action trajectory. This offline action trajectory is compared with the standard action trajectory in the demonstration data to obtain the offline behavior cloning loss. This offline action trajectory is also compared with the offline action task instructions currently output by the model, which serves as the basis for the environmental feedback, namely the reward function. If the robot's action trajectory is very similar to the task instruction, the reward function will give a higher reward; conversely, if the action trajectory deviates far from the task instruction, the reward will be lower. The corresponding Q-value function is then obtained to calculate the offline Q-value function loss, which is then used to gradually optimize the parameters of the Q-value function. The offline behavior cloning loss and the offline Q-value function loss are used to obtain the corresponding offline total loss. The initialized visual language action model is then fine-tuned, and the model parameters and Q-value function are continuously optimized to ensure that the model selects more optimal actions when making action decisions. This is done until the model's offline performance reaches the expected target, such as learning a certain number of iterations or the offline total loss is less than a preset value. This completes the offline reinforcement fine-tuning process and obtains the offline fine-tuned model. Through offline intensive fine-tuning, we ensure that the model can accurately and quickly learn strategies close to standard actions from a small amount of demonstration data.
[0065] In one embodiment, Figure 4 As shown, step S304 includes:
[0066] S401, respectively encoding the offline motion trajectory and the offline motion task instruction to obtain corresponding offline motion trajectory vectors and offline task vectors;
[0067] S402, using the similarity between the offline action trajectory vector and the offline task vector as an offline reward function, updating the Q-value function according to the offline reward function and calculating the offline Q-value function loss;
[0068] S403, calculating offline behavior cloning loss according to the difference between the offline action trajectory and the standard action trajectory in the demonstration data;
[0069] S404 : Perform a weighted summation on the offline Q-value function loss and the offline behavior cloning loss according to a first weight strategy to obtain the offline total loss.
[0070] In this embodiment, the loss function of the offline stage includes two items, namely, the offline Q-value function loss and the offline behavior cloning loss. When calculating the offline Q-value function loss, the reward function is obtained by calculating the similarity between the actual action execution trajectory and the task instructions, so that there is no need to separately train a binary classifier to determine whether the task is successful, reducing the data volume requirement.
[0071] Specifically, after the robot completes a task, its offline motion trajectory is obtained and encoded through the S3D network. At the same time, the natural language instructions of the offline motion task are also encoded through the S3D network. The two vectors obtained are dot-producted to obtain the similarity between the actual action performed by the robot and the task instruction. This similarity is used as the offline reward function for reinforcement learning, thereby avoiding the problem of additional data collection caused by training a binary classifier. The Q-value function is updated according to the offline reward function to optimize the model's strategy so that the actions it generates are more in line with the task requirements, and the offline Q-value function loss L is calculated. off Q , which represents the difference between the predicted Q value and the actual Q value. Specifically, the mean square error can be used as an offline Q value function. Using similarity as a reward signal for reinforcement learning, the model is guided to learn actions that are more consistent with the task objectives.
[0072] And calculate the difference between the offline action trajectory and the standard action trajectory in the demonstration data to obtain the offline behavior cloning loss L off BC, which can be calculated by mean square error (MSE) or other distance metrics, which is not limited in this embodiment. The difference between the action generated by the model and the standard action in the demonstration data is measured by calculating the offline behavior cloning loss. The constraint of the behavior cloning loss ensures that the action generated by the model is consistent with the action demonstrated by the expert, making it closer to the expert's demonstration. According to the task requirements and the model training stage, the corresponding weights α1 and β1 are assigned to the offline Q-value function loss and the offline behavior cloning loss according to the first weight strategy, and then the weighted sum is performed to obtain the offline total loss:
[0073] L off total =α1*L off Q +β1*L off BC
[0074] The importance of different losses is balanced through a weight strategy. The total loss of the offline phase is obtained by combining the offline Q-value function loss and the offline behavior cloning loss to comprehensively optimize the model parameters and ensure that the model can learn effective strategies in the offline phase.
[0075] In one embodiment, Figure 5 As shown, step S203 includes:
[0076] S501, deploying the offline fine-tuning model into a real environment, and generating corresponding online action task instructions based on visual information and language instructions;
[0077] S502, controlling the robot to sequentially execute a forward motion task and a backward motion task according to the online motion task instruction;
[0078] S503: When both the forward motion task and the backward motion task are completed, corresponding exploration trajectory and environment feedback are obtained.
[0079] In this embodiment, after completing the offline reinforcement learning, the offline fine-tuning model is deployed in the actual environment to control the robot to perform corresponding actions and interact with the real environment to obtain corresponding environmental feedback. Specifically, in the real environment, visual information (such as camera images) and language instructions (such as voice instructions from the client or therapist) are collected in real time, and the collected visual information and language instructions are input into the offline fine-tuning model. At this time, the offline fine-tuning model processes these inputs based on the reasoning ability obtained after the offline fine-tuning, and generates corresponding online action task instructions for controlling the robot to perform tasks. This enables the model to generate action instructions in real time based on the visual and language inputs in the real environment to adapt to changes in the real environment.
[0080] Afterwards, the robot is controlled to perform forward-backward task execution through the task reset strategy, that is, the robot is controlled to perform forward action tasks and backward action tasks in sequence according to the online action task instructions. The forward action task refers to the robot performing forward operations according to the task instructions, and the backward action task refers to the robot performing reverse operations to restore to the initial state. By using the task reset strategy, it is ensured that the robot can automatically restore to the initial state after the task is completed so as to perform the next task, thereby reducing human intervention.
[0081] During the robot's execution of forward and backward motion tasks, the robot's motion trajectory is recorded, including information such as position, speed, and motion sequence, thereby obtaining an exploration trajectory. Furthermore, feedback from the environment on the robot's motion is collected. Specifically, the environmental feedback can use a reward function similar to that used in the offline phase, i.e., by calculating the similarity between the online motion task instructions and the exploration trajectory as a reward function. The specific calculation method is the same as in the previous embodiment and will not be described in detail here. Of course, the environmental feedback obtained in the actual environment can further include other information, such as task status signals such as customer satisfaction and patient rehabilitation progress assessment, which are not limited in this embodiment. By controlling the robot to autonomously interact with the environment in the actual environment, executing automatically reset tasks, and recording exploration trajectories and environmental feedback, rich data support is provided for online reinforcement learning, thereby continuously optimizing the model strategy and improving the accuracy and reliability of the robot's task execution.
[0082] In one embodiment, Figure 6 As shown, step S502 includes:
[0083] S601, parsing the online action task instruction to generate a corresponding action sequence;
[0084] S602: sort all actions forward and backward according to the execution order of each action in the action sequence to generate corresponding forward task sequence and backward task sequence;
[0085] S603: Splice the forward task sequence and the backward task sequence and send them to the robot, and control the robot to execute the forward action task and the backward action task in sequence.
[0086] In this embodiment, when controlling the robot to interact with the environment according to the task reset strategy, the online action task instruction is parsed to generate a series of specific action instructions to form an action sequence. For example, for the instruction "grab the red object and place it in the blue box", the generated action sequence may include:
[0087] Action 1: Move to the position of the red object;
[0088] Action 2: Grab the red object;
[0089] Action 3: Move to the position of the blue box;
[0090] Action 4: Release the red object.
[0091] Regarding the execution order of each action in the action sequence, the actions in the action sequence are arranged in sequence to form a forward task sequence, that is, the above-mentioned actions 1 to 4 are arranged in sequence. At the same time, the actions in the action sequence are arranged in sequence in the reverse order of the execution order of each action to form a backward task sequence, that is, the sequence of action 4, action 3, action 2, and action 1. The forward task sequence and the backward task sequence are spliced together to form a complete task sequence. That is, Mr. Wang’s task sequence includes action 1, action 2, action 3, action 4, action 4, action 3, action 2, and action 1. The spliced task sequence is sent to the robot, and the robot is controlled to execute the forward action task and the backward action task in sequence, so that the environment returns to its initial state after execution. By splicing and sending task sequences, the robot can execute tasks in a predetermined order, ensuring the integrity and coherence of the tasks. The robot can also interact with the environment autonomously and automatically reset between different human tasks, reducing dependence on human intervention.
[0092] In one embodiment, step S204 includes:
[0093] According to the exploration trajectory, the environment feedback and the standard action trajectory in the demonstration data, the online behavior cloning loss and the online Q-value function loss are calculated to obtain the online total loss;
[0094] Fine-tuning the offline fine-tuning model according to the online total loss to optimize the motion strategy of the offline fine-tuning model, outputting optimized online motion task instructions to control the robot to interact with the environment, and obtaining new exploration trajectories and environmental feedback;
[0095] The above loss evaluation and fine-tuning interactive process is executed cyclically until the online total loss meets the preset conditions to obtain the fine-tuned visual language action model.
[0096] In this embodiment, after each interaction with the real environment, i.e., after executing a task, the robot collects the corresponding exploration trajectory and environmental feedback. This environmental feedback serves as the online reward function for online reinforcement learning, thereby obtaining a corresponding Q-value function and calculating the online Q-value function loss. The parameters of the Q-value function are then gradually optimized using the rewards obtained from the interaction. Furthermore, the exploration trajectory is compared with the standard action trajectory in the demonstration data to obtain the online behavior cloning loss. The online behavior cloning loss is then combined with the consultative Q-value function loss to obtain the corresponding online total loss. The offline fine-tuning model is then fine-tuned to optimize its action strategy. The optimized online action task instructions are then output to control the robot's interaction with the environment, obtaining new exploration trajectories and environmental feedback. This loss evaluation and fine-tuning process is then repeated, and online reinforcement learning is performed on the offline fine-tuning model until the online total loss meets a preset condition, such as being less than a preset threshold. This results in a fine-tuned visual language action model. By continuously optimizing the action strategy in the real environment through interaction with the environment based on offline reinforcement fine-tuning, a fine-tuned model with improved performance is ultimately obtained, capable of adapting to changes in diverse environments.
[0097] In one embodiment, the online total loss is obtained by calculating the online behavior cloning loss and the online Q-value function loss based on the exploration trajectory, the environment feedback, and the standard action trajectory in the demonstration data, including:
[0098] Update the Q-value function according to the environmental feedback and calculate the online Q-value function loss;
[0099] Calculating an online behavior cloning loss based on the difference between the exploration trajectory and the standard action trajectory in the demonstration data;
[0100] The online Q-value function loss and the online behavior cloning loss are weighted and summed according to a second weight strategy to obtain the online total loss.
[0101] In this embodiment, similar to the offline phase, the loss function in the online phase also includes two items, namely, the online Q-value function loss and the online behavior cloning loss. The environmental feedback obtained by interacting with the actual environment is used as the reward function, and the Q-value function is updated to optimize the model's strategy so that the actions it generates are more in line with the task requirements. The online Q-value function loss L is calculated. on Q , this loss function represents the difference between the predicted Q value and the actual Q value. Specifically, the mean square error can be used as the online Q value function, and reinforcement learning can be performed through reward signals to guide the model to learn actions that are more in line with the task objectives.
[0102] And calculate the difference between the exploration and the standard action trajectory in the demonstration data to obtain the online behavior cloning loss L onBC , which can be calculated specifically by mean square error (MSE) or other distance metrics, which is not limited in this embodiment. The difference between the robot's exploration trajectory in the actual environment and the standard action in the demonstration data is measured by calculating the online behavior cloning loss. The constraint of the behavior cloning loss ensures that the actions generated by the model are consistent with the actions demonstrated by the expert, making it closer to the expert's demonstration. According to the task requirements and the model training stage, the corresponding weights α2 and β2 are assigned to the online Q-value function loss and the online behavior cloning loss according to the second weight strategy, and the weighted sum is performed to obtain the online total loss:
[0103] L on total =α2*L on Q +β2*L on BC
[0104] The online reinforcement learning stage also uses a weight strategy to balance the importance of different losses. Unlike the offline stage, the weights of the two loss terms can be adjusted in the online stage to adapt to the fine-tuning requirements of the online stage. For example, compared with the first weight strategy in the offline stage, the weight of the online behavior cloning loss, i.e., β2, can be reduced in the second weight strategy, while the weight of the online Q-value function loss, i.e., α2, can be increased. This can avoid excessive deviations in the strategy while strengthening the guidance model to learn actions that are more in line with the task objectives.
[0105] It should be noted that there is not necessarily a certain order between the above steps. A person skilled in the art can understand, based on the description of the embodiments of the present invention, that in different embodiments, the above steps may have different execution orders, that is, they may be executed in parallel, or may be executed interchangeably, etc.
[0106] Further references Figure 7 , as a response to the above Figure 2 The present invention provides an embodiment of a device for enhancing and fine-tuning a visual language action model. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.
[0107] like Figure 7 As shown, the visual language action model enhancement and fine-tuning device 70 of this embodiment includes:
[0108] Model loading module 701, used to load the visual language action model to be fine-tuned, wherein the visual language action model is used to output action task instructions based on visual information and language instructions to operate the robot to perform the corresponding action task;
[0109] An offline reinforcement fine-tuning module 702 is configured to collect multiple pieces of demonstration data and perform offline reinforcement learning on the visual language action model using the demonstration data to obtain an offline fine-tuning model;
[0110] An online interaction module 703 is used to deploy the offline fine-tuning model into the actual environment, control the robot to interact with the environment according to the task reset strategy, and obtain corresponding exploration trajectory and environmental feedback;
[0111] The online reinforcement fine-tuning module 704 is used to perform online reinforcement learning on the offline fine-tuning model according to the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model.
[0112] The module referred to in the present invention refers to a series of computer program instruction segments that can perform specific functions. It is more suitable for describing the enhanced fine-tuning execution process of the visual language action model than a program. For the specific implementation of each module, please refer to the corresponding method embodiment above, which will not be repeated here.
[0113] In one embodiment, the offline enhancement fine-tuning module 702 includes:
[0114] A data acquisition unit, configured to acquire multiple pieces of demonstration data and initialize the visual language action model;
[0115] An offline instruction output unit, configured to input the demonstration data into the initialized visual language action model and output corresponding offline action task instructions;
[0116] An offline trajectory acquisition unit, configured to control the robot to perform a corresponding offline motion task according to the offline motion task instruction and acquire an offline motion trajectory;
[0117] An offline loss calculation unit, configured to calculate an offline behavior cloning loss and an offline Q-value function loss based on the offline action trajectory, the standard action trajectory in the demonstration data, and the offline action task instruction, and then obtain an offline total loss;
[0118] An offline fine-tuning unit is used to fine-tune the initialized visual language action model according to the offline total loss to obtain an offline fine-tuning model.
[0119] In one embodiment, the offline loss calculation unit includes:
[0120] an encoding unit, configured to encode the offline motion trajectory and the offline motion task instruction respectively to obtain a corresponding offline motion trajectory vector and an offline task vector;
[0121] a first loss calculation unit, configured to use the similarity between the offline action trajectory vector and the offline task vector as an offline reward function, update the Q-value function according to the offline reward function, and calculate the offline Q-value function loss;
[0122] a second loss calculation unit, configured to calculate an offline behavior cloning loss according to a difference between the offline action trajectory and a standard action trajectory in the demonstration data;
[0123] The offline total loss calculation unit is used to perform weighted summation on the offline Q-value function loss and the offline behavior cloning loss according to a first weight strategy to obtain the offline total loss.
[0124] In one embodiment, the online interaction module 703 includes:
[0125] An online instruction output unit deploys the offline fine-tuning model into a real environment and generates corresponding online action task instructions based on visual information and language instructions;
[0126] A task resetting control unit, configured to control the robot to sequentially execute a forward motion task and a backward motion task according to the online motion task instruction;
[0127] The interactive information acquisition unit is used to obtain corresponding exploration trajectories and environmental feedback when both the forward action task and the backward action task are completed.
[0128] In one embodiment, the task reset control unit includes:
[0129] An action parsing unit, configured to parse the online action task instructions and generate corresponding action sequences;
[0130] A sequence generation unit is used to perform forward sorting and backward sorting on all actions according to the execution order of each action in the action sequence, and generate corresponding forward task sequence and backward task sequence;
[0131] The splicing execution unit is used to splice the forward task sequence and the backward task sequence and send them to the robot, so as to control the robot to execute the forward action task and the backward action task in sequence.
[0132] In one embodiment, the online enhancement fine-tuning module 704 includes:
[0133] An online loss calculation unit, which calculates an online behavior cloning loss and an online Q-value function loss based on the exploration trajectory, the environment feedback, and the standard action trajectory in the demonstration data to obtain an online total loss;
[0134] An online fine-tuning interaction unit is used to fine-tune the offline fine-tuning model according to the online total loss to optimize the motion strategy of the offline fine-tuning model, output optimized online motion task instructions to control the robot to interact with the environment, and obtain new exploration trajectories and environmental feedback;
[0135] The model output unit is used to cyclically execute the above loss evaluation and fine-tuning interactive process until the online total loss meets the preset conditions to obtain the fine-tuned visual language action model.
[0136] In one embodiment, the online loss calculation unit includes:
[0137] A third loss calculation unit, configured to update the Q-value function according to the environmental feedback and calculate an online Q-value function loss;
[0138] a fourth loss calculation unit, configured to calculate an online behavior cloning loss based on a difference between the exploration trajectory and a standard action trajectory in the demonstration data;
[0139] An online total loss calculation unit is used to perform weighted summation on the online Q-value function loss and the online behavior cloning loss according to a second weight strategy to obtain the online total loss.
[0140] In the above embodiment, the present invention discloses a device for enhancing fine-tuning of a visual language action model, which loads a visual language action model to be fine-tuned, the visual language action model being used to output action task instructions based on visual information and language instructions to operate a robot to perform corresponding action tasks; collects multiple demonstration data, and performs offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model; deploys the offline fine-tuning model into an actual environment, controls the robot to interact with the environment according to a task reset strategy, and obtains corresponding exploration trajectories and environmental feedback; performs online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback, and demonstration data to obtain a fine-tuned visual language action model. By performing offline reinforcement learning and online reinforcement learning in stages, and automatically restoring the environment state through the task reset strategy in the online learning stage, manual intervention is avoided, so that on the basis of quickly improving the performance of the initial action strategy, non-intervention automatic optimization can also be performed through interaction with the real environment, realizing collaborative model fine-tuning to improve the stability and reliability of the robot operation.
[0141] Another embodiment of the present invention provides a computer device, such as Figure 8 As shown, the computer device 80 includes:
[0142] One or more processors 801 and memory 802, Figure 8In the description, a processor 801 is used as an example. The processor 801 and the memory 802 can be connected via a bus or other means. Figure 8 The bus connection is taken as an example.
[0143] The processor 801 is used to complete various control logics of the computer device 80. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. In addition, the processor 801 can also be any traditional processor, microprocessor, or state machine. The processor 801 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP, and / or any other such configuration.
[0144] Memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as program instructions corresponding to the method for enhancing and fine-tuning a visual language action model in the embodiments of the present invention. Processor 801 executes the non-volatile software programs, instructions, and modules stored in memory 802 to execute various functional applications and data processing of computer device 80, thereby implementing the method for enhancing and fine-tuning a visual language action model in the aforementioned method embodiments.
[0145] The memory 802 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device 80, etc. In addition, the memory 802 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 802 may optionally include a memory remotely located relative to the processor 801, and these remote memories may be connected to the computer device 80 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. One or more units are stored in the memory 802, and when executed by one or more processors 801, the steps of the enhanced fine-tuning method of the visual language action model in any of the above-mentioned method embodiments are executed.
[0146] In the above embodiment, the present invention discloses a computer device, which loads a visual language action model to be fine-tuned, the visual language action model is used to output action task instructions based on visual information and language instructions to operate a robot to perform corresponding action tasks; collects multiple demonstration data, and performs offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model; deploys the offline fine-tuning model into an actual environment, controls the robot to interact with the environment according to a task reset strategy, and obtains corresponding exploration trajectories and environmental feedback; performs online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model. By performing offline reinforcement learning and online reinforcement learning in stages, and automatically restoring the environment state through the task reset strategy in the online learning stage, manual intervention is avoided, so that on the basis of quickly improving the performance of the initial action strategy, non-intervention automatic optimization can also be performed through interaction with the real environment, realizing collaborative model fine-tuning to improve the stability and reliability of the robot operation.
[0147] An embodiment of the present invention provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by one or more processors, the steps of the enhanced fine-tuning method of the visual language action model in any of the above method embodiments are executed.
[0148] In the above embodiment, the present invention discloses a non-volatile computer-readable storage medium, which loads a visual language action model to be fine-tuned, the visual language action model being used to output action task instructions based on visual information and language instructions to operate a robot to perform corresponding action tasks; collects multiple demonstration data, and performs offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model; deploys the offline fine-tuning model into an actual environment, controls the robot to interact with the environment according to a task reset strategy, obtains corresponding exploration trajectories and environmental feedback; performs online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model. By performing offline reinforcement learning and online reinforcement learning in stages, and automatically restoring the environment state through the task reset strategy in the online learning stage, manual intervention is avoided, so that on the basis of quickly improving the performance of the initial action strategy, non-intervention automatic optimization can also be performed through interaction with the real environment, realizing collaborative model fine-tuning to improve the stability and reliability of the robot operation.
[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0150] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0151] In summary, the present invention discloses a method, apparatus, device, and medium for enhancing fine-tuning of a visual language action model, the method comprising: loading a visual language action model to be fine-tuned, the visual language action model being used to output action task instructions based on visual information and language instructions to operate a robot to perform corresponding action tasks; collecting multiple demonstration data, and performing offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model; deploying the offline fine-tuning model into an actual environment, controlling the robot to interact with the environment according to a task reset strategy, and obtaining corresponding exploration trajectories and environmental feedback; performing online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback, and demonstration data to obtain a fine-tuned visual language action model. By performing offline reinforcement learning and online reinforcement learning in stages, and automatically restoring the environment state through a task reset strategy in the online learning stage, manual intervention is avoided, so that on the basis of rapidly improving the performance of the initial action strategy, non-intervention automatic optimization can also be performed through interaction with the real environment, achieving collaborative model fine-tuning to improve the stability and reliability of the robot operation.
[0152] Of course, those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a processor, controller, etc.) through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes in the above-described method embodiments. The storage medium can be a memory, a magnetic disk, a floppy disk, a flash memory, an optical storage device, etc.
[0153] It should be noted that if any software tools or components not developed by our company appear in the examples of this application, they are for illustration purposes only and do not represent actual use. It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications shall fall within the scope of protection of the appended claims.
Claims
1. A method for enhancing fine-tuning of a visual language action model, characterized in that: include: Loading a visual language action model to be fine-tuned, wherein the visual language action model is used to output action task instructions based on visual information and language instructions to operate the robot to perform the corresponding action task; Collecting multiple pieces of demonstration data, and performing offline reinforcement learning on the visual language action model using the demonstration data to obtain an offline fine-tuning model; Deploying the offline fine-tuning model into a real environment, controlling the robot to interact with the environment according to the task reset strategy, and obtaining corresponding exploration trajectories and environmental feedback; The offline fine-tuning model is subjected to online reinforcement learning according to the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model.
2. The method for enhancing and fine-tuning a visual language action model according to claim 1, characterized in that: The collecting of multiple pieces of demonstration data and performing offline reinforcement learning on the visual language action model using the demonstration data to obtain an offline fine-tuning model includes: Collecting multiple pieces of demonstration data and initializing the visual language action model; Inputting the demonstration data into the initialized visual language action model and outputting corresponding offline action task instructions; Controlling the robot to perform the corresponding offline motion task according to the offline motion task instruction and obtaining the offline motion trajectory; According to the offline action trajectory, the standard action trajectory in the demonstration data, and the offline action task instruction, calculating the offline behavior cloning loss and the offline Q-value function loss to obtain the offline total loss; The initialized visual-language-action model is fine-tuned according to the offline total loss to obtain an offline fine-tuning model.
3. The method for enhancing and fine-tuning a visual language action model according to claim 2, characterized in that: The calculating of the offline behavior cloning loss and the offline Q-value function loss according to the offline action trajectory, the standard action trajectory in the demonstration data, and the offline action task instruction to obtain the offline total loss includes: Encoding the offline motion trajectory and the offline motion task instruction respectively to obtain corresponding offline motion trajectory vectors and offline task vectors; Using the similarity between the offline action trajectory vector and the offline task vector as an offline reward function, updating the Q-value function according to the offline reward function and calculating the offline Q-value function loss; Calculating an offline behavior cloning loss based on a difference between the offline action trajectory and a standard action trajectory in the demonstration data; The offline Q-value function loss and the offline behavior cloning loss are weighted and summed according to a first weight strategy to obtain the offline total loss.
4. The method for enhancing and fine-tuning a visual language action model according to claim 1, wherein: The offline fine-tuning model is deployed into the actual environment, and the robot is controlled to interact with the environment according to the task reset strategy to obtain the corresponding exploration trajectory and environmental feedback, including: Deploying the offline fine-tuning model into a real environment to generate corresponding online action task instructions based on visual information and language instructions; Controlling the robot to sequentially perform a forward motion task and a backward motion task according to the online motion task instruction; When both the forward action task and the backward action task are completed, the corresponding exploration trajectory and environmental feedback are obtained.
5. The method for enhancing and fine-tuning a visual language action model according to claim 4, characterized in that: The controlling the robot to sequentially execute the forward motion task and the backward motion task according to the online motion task instruction comprises: Parsing the online action task instructions to generate corresponding action sequences; According to the execution order of each action in the action sequence, all actions are sorted forward and backward to generate corresponding forward task sequences and backward task sequences; The forward task sequence and the backward task sequence are spliced and sent to the robot, and the robot is controlled to perform the forward action task and the backward action task in sequence.
6. The method for enhancing and fine-tuning a visual language action model according to claim 1, characterized in that: The performing online reinforcement learning on the offline fine-tuning model according to the exploration trajectory, environmental feedback, and demonstration data to obtain a fine-tuned visual language action model includes: According to the exploration trajectory, the environment feedback and the standard action trajectory in the demonstration data, the online behavior cloning loss and the online Q-value function loss are calculated to obtain the online total loss; Fine-tuning the offline fine-tuning model according to the online total loss to optimize the motion strategy of the offline fine-tuning model, outputting optimized online motion task instructions to control the robot to interact with the environment, and obtaining new exploration trajectories and environmental feedback; The above loss evaluation and fine-tuning interactive process is executed cyclically until the online total loss meets the preset conditions to obtain the fine-tuned visual language action model.
7. The method for enhancing and fine-tuning a visual language action model according to claim 6, characterized in that: The online total loss is obtained by calculating the online behavior cloning loss and the online Q-value function loss based on the exploration trajectory, the environment feedback, and the standard action trajectory in the demonstration data, including: Update the Q-value function according to the environmental feedback and calculate the online Q-value function loss; Calculating an online behavior cloning loss based on the difference between the exploration trajectory and the standard action trajectory in the demonstration data; The online Q-value function loss and the online behavior cloning loss are weighted and summed according to a second weight strategy to obtain the online total loss.
8. A device for enhancing and fine-tuning a visual language action model, characterized in that: include: A model loading module is used to load the visual language action model to be fine-tuned, wherein the visual language action model is used to output action task instructions based on visual information and language instructions to operate the robot to perform the corresponding action task; An offline reinforcement fine-tuning module is used to collect multiple demonstration data and perform offline reinforcement learning on the visual language action model through the demonstration data to obtain an offline fine-tuning model; An online interaction module is used to deploy the offline fine-tuning model into the actual environment, control the robot to interact with the environment according to the task reset strategy, and obtain corresponding exploration trajectory and environmental feedback; The online reinforcement fine-tuning module is used to perform online reinforcement learning on the offline fine-tuning model based on the exploration trajectory, environmental feedback and demonstration data to obtain a fine-tuned visual language action model.
9. A computer device, characterized in that: comprising at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the enhanced fine-tuning method of the visual language action model according to any one of claims 1 to 7.
10. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, which, when executed by one or more processors, enable the one or more processors to execute the enhanced fine-tuning method for the visual language action model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Offline meta-reinforcement learning model training method and device, equipment and storage medium
CN112348113A
Small sample robust imitation learning training method oriented to high-dimensional disturbance environment, electronic equipment and storage medium
CN117193008A
Decision-making method and model for offline reinforcement learning and continuous online fine tuning
CN119249360A
Control method of multi-axis mechanical arm
CN119458384A
Cited By
Vision-language-action model training method and system utilizing inference data closed-loop optimization
CN121581156A
VLA model autonomous generalization method, system, device and medium
CN121638318A
A vla model autonomous generalization method, system, device and medium
CN121638318B