Model training method, device, equipment and medium based on visual reinforcement learning
By combining the visual large language model and self-supervised loss, the parameter training of the visual reinforcement learning model is optimized, which solves the problem of the visual encoder's over-reliance on shallow features, improves the model's ability to understand complex scenes and training efficiency, and enhances decision-making performance.
Patent Information
- Application Number
- CN202511038413.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-28
AI Technical Summary
In existing model training based on visual reinforcement learning, the visual encoder relies too much on the shallow features of the task reward and cannot effectively capture the high-level semantic associations between objects and actions in the scene, resulting in a reduced ability of the model to understand complex scenes, which in turn affects training efficiency and decision-making performance.
By combining the large visual language model and self-supervised loss, constructing distillation loss and target policy loss, the parameter training process of the visual reinforcement learning model is optimized. The semantic information and reasoning ability of the large visual language model are utilized and distilled into a lightweight visual reinforcement learning model, thereby improving the model's ability to understand complex scenes and training efficiency.
It improves the model's ability to understand complex scenarios and training efficiency, enhances the model's decision-making performance and semantic generalization, reduces ineffective exploration in the environment, and improves training efficiency and decision-making performance.
Smart Images

Figure CN120543954B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training method, apparatus, equipment and medium based on visual reinforcement learning. Background Art
[0002] Vision-based reinforcement learning (VRL) is an AI technology that learns decision-making policies from high-dimensional visual observations (such as raw pixel inputs). It is widely used in scenarios such as autonomous driving, embodied intelligence, and autonomous mobile robots. Reinforcement learning models use a visual encoder to compress raw images into a low-dimensional latent space representation and a policy decoder to output control actions based on this low-dimensional latent space representation.
[0003] Related technologies typically use a temporal difference loss approach to train models based on visual reinforcement learning. Specifically, the difference between the estimated value of the current state and the actual observed immediate reward plus the estimated value of the next state is used as a loss signal to drive the model's updates to the visual encoder and policy decoder. However, this optimization approach causes the visual encoder to learn shallow features that overly rely on task rewards, failing to capture high-level semantic associations between objects and actions in the scene. This reduces the model's ability to understand complex scenes, which in turn reduces the efficiency of model training and the model's decision-making performance. Summary of the Invention
[0004] This application proposes a model training method, device, equipment and medium based on visual reinforcement learning, which can improve the efficiency of model training and improve the decision-making performance of the model.
[0005] To achieve the above objectives, a first aspect of an embodiment of the present application proposes a model training method based on visual reinforcement learning, the method comprising:
[0006] Obtaining a sample image frame and corresponding semantic category information, inputting the semantic category information into an initial first convolution kernel of a visual large language model to extract text features, obtaining first convolution kernel parameters of the first feature convolution kernel after extracting text features, and inputting the sample image frame into the first feature convolution kernel to obtain a first feature heat map;
[0007] Processing the sample image frame using an initial second feature convolution kernel included in a visual encoder of a preset visual reinforcement learning model to obtain second convolution kernel parameters and a second feature heat map of the second feature convolution kernel after processing the image;
[0008] Constructing a first distillation loss based on a difference between the first convolution kernel parameters and the second convolution kernel parameters, and constructing a second distillation loss based on a difference between the first feature heat map and the second feature heat map;
[0009] Predicting the state and reward at the next moment for the sample action data and sample state data corresponding to the second feature heat map using a preset self-supervised model, obtaining a prediction result for each sample image frame, and constructing a self-supervised loss based on multiple prediction results corresponding to multiple sample image frames;
[0010] Performing action value and action sampling calculations based on the sample action data and the sample state data using a policy decoder of the preset visual reinforcement learning model to obtain calculation results for each sample image frame, and constructing a target policy loss based on multiple calculation results corresponding to multiple sample image frames;
[0011] Based on the first distillation loss, the second distillation loss, the self-supervision loss, and the target strategy loss, the parameters of the preset visual reinforcement learning model are adjusted to obtain a target visual reinforcement learning model.
[0012] Accordingly, a second aspect of the embodiments of the present application proposes a model training device based on visual reinforcement learning, the device comprising:
[0013] an acquisition module, configured to acquire a sample image frame and corresponding semantic category information, input the semantic category information into an initial first convolution kernel of a visual large language model to extract text features, obtain first convolution kernel parameters of the first feature convolution kernel after extracting text features, and input the sample image frame into the first feature convolution kernel to obtain a first feature heat map;
[0014] a processing module, configured to process the sample image frame using an initial second feature convolution kernel included in a visual encoder of a preset visual reinforcement learning model to obtain second convolution kernel parameters and a second feature heat map of the second feature convolution kernel after processing the image;
[0015] A construction module, configured to construct a first distillation loss based on a difference between the first convolution kernel parameters and the second convolution kernel parameters, and to construct a second distillation loss based on a difference between the first feature heat map and the second feature heat map;
[0016] A prediction module, configured to predict the state and reward at the next moment for the sample action data and sample state data corresponding to the second feature heat map using a preset self-supervised model, obtain a prediction result for each sample image frame, and construct a self-supervised loss based on multiple prediction results corresponding to multiple sample image frames;
[0017] a calculation module, configured to perform action value and action sampling calculations based on the sample action data and the sample state data using a policy decoder of the preset visual reinforcement learning model, obtain a calculation result for each sample image frame, and construct a target policy loss based on multiple calculation results corresponding to multiple sample image frames;
[0018] An adjustment module is used to adjust the parameters of the preset visual reinforcement learning model based on the first distillation loss, the second distillation loss, the self-supervision loss and the target strategy loss to obtain a target visual reinforcement learning model.
[0019] In some embodiments, the self-supervised loss includes a first self-supervised loss and a second self-supervised loss, and the prediction module is further configured to:
[0020] By using a preset self-supervisory model, for each sample image frame, the state and reward at the next moment are predicted for the sample action data and sample state data corresponding to the second feature heat map, thereby obtaining the predicted state data and predicted reward data for each sample image frame at the next moment;
[0021] Obtain the target sample state data and target sample reward data at the next moment, and construct a first self-supervised loss based on the difference between the predicted state data and the target sample state data of multiple sample image frames, and construct a second self-supervised loss based on the difference between the predicted reward data and the target sample reward data of the multiple sample image frames.
[0022] In some embodiments, the target strategy loss includes a value network sub-loss and a strategy sub-loss, and the calculation module is further configured to:
[0023] The value network included in the policy decoder of the preset visual reinforcement learning model is used to calculate the action value for each sample image frame based on the sample action data and the sample state data to obtain the corresponding expected cumulative reward for the action;
[0024] Construct a value network sub-loss based on the expected cumulative rewards of multiple actions corresponding to multiple sample image frames;
[0025] For each sample image frame, obtaining an action probability distribution generated by the strategy decoder according to the sample state data, performing action sampling according to the action probability distribution to obtain sampled action data, and calculating an action probability according to the sampled action data;
[0026] A strategy sub-loss is constructed according to a plurality of action probabilities corresponding to the plurality of sample image frames.
[0027] In some embodiments, the computing module is further configured to:
[0028] For each sample image frame, obtaining an immediate reward of environmental feedback after executing the sample action data under the sample state data;
[0029] Acquire the updated sample state data and the updated sample action data of each sample image frame at the next moment, and calculate the updated state value at the next moment based on the updated sample state data and the updated sample action data;
[0030] Obtaining preset adjustment parameters, and adjusting the updated state value according to the adjustment parameters to obtain a target state value;
[0031] Obtaining target value data according to the sum of the instant reward and the target state value;
[0032] Obtaining a first difference based on a difference between the expected cumulative reward of the action and the target value data;
[0033] A value network sub-loss is constructed based on the average of multiple first differences corresponding to multiple sample image frames.
[0034] In some embodiments, the computing module is further configured to:
[0035] For each sample image frame, obtaining an update action probability distribution generated by the strategy decoder for the update sample state data;
[0036] Performing action sampling according to the next action probability distribution to obtain updated sampled action data, and determining an updated action probability of the updated sampled action data;
[0037] Obtaining a preset entropy parameter, and obtaining a first product based on the product of the update action probability and the preset entropy parameter;
[0038] Calculating the expected cumulative reward of the corresponding update action according to the update sample action data and the update sample state data through the value network included in the policy decoder;
[0039] Obtaining a second difference based on a difference between the expected cumulative reward of the update action and the first product;
[0040] An updated state value at the next moment is obtained based on an average of a plurality of second difference values corresponding to a plurality of sample image frames.
[0041] In some embodiments, the computing module is further configured to:
[0042] Obtaining a preset entropy parameter, and obtaining a second product according to the product of the preset entropy parameter and the action probability for each sample image frame;
[0043] Obtaining a preset target action expected cumulative reward, and obtaining a third difference based on a difference between the second product and the target action expected cumulative reward;
[0044] A strategy sub-loss is constructed based on the average of multiple third differences corresponding to multiple sample image frames.
[0045] In some embodiments, the acquisition module is further configured to:
[0046] Acquire sample image frames, wherein the sample image frames include a first image frame at a current moment, a second image frame at a second moment, and a third image frame at a third moment, wherein the second moment is subsequent to the current moment, and the third moment is subsequent to the second moment;
[0047] Acquire sample motion data corresponding to the first image frame and reference sample motion data corresponding to the second image frame;
[0048] Obtain a preset semantic hint, and input the semantic hint, the first image frame, the second image frame, the third image frame, the sample action data, and the reference sample action data into a first visual large language model to obtain semantic category information corresponding to the sample image frame.
[0049] Correspondingly, the third aspect of the embodiments of the present application proposes a computer device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the model training method based on visual reinforcement learning of any one of the embodiments of the first aspect of the present application.
[0050] Correspondingly, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the model training method based on visual reinforcement learning of any one of the embodiments of the first aspect of the present application.
[0051] The embodiment of the present application obtains a sample image frame and corresponding semantic category information, and inputs the semantic category information into the initial first convolution kernel of the visual large language model to extract text features, thereby obtaining the first convolution kernel parameter of the first feature convolution kernel after extracting the text features, and inputs the sample image frame into the first feature convolution kernel to obtain a first feature heat map; processes the sample image frame through the initial second feature convolution kernel included in the visual encoder of the preset visual reinforcement learning model to obtain the second convolution kernel parameter of the second feature convolution kernel after processing the image and the second feature heat map; constructs a first distillation loss based on the difference between the first convolution kernel parameter and the second convolution kernel parameter, and constructs a first distillation loss based on the difference between the first feature heat map and the second feature heat map. Second distillation loss; through the preset self-supervised model, the state and reward of the sample action data and sample state data corresponding to the second feature heat map are predicted at the next moment to obtain the prediction result of each sample image frame, and the self-supervised loss is constructed according to the multiple prediction results corresponding to the multiple sample image frames; through the policy decoder of the preset visual reinforcement learning model, the action value and action sampling calculation are performed based on the sample action data and sample state data to obtain the calculation result of each sample image frame, and the target policy loss is constructed according to the multiple calculation results corresponding to the multiple sample image frames; based on the first distillation loss, the second distillation loss, the self-supervised loss and the target policy loss, the parameters of the preset visual reinforcement learning model are adjusted to obtain the target visual reinforcement learning model. In this way, the large visual language model can be innovatively applied to the field of reinforcement learning for efficient auxiliary training. By combining the representation ability of the large visual language model, the semantic information and reasoning ability of the large visual language model are distilled into a lightweight preset visual reinforcement learning model, so that the preset visual reinforcement learning model can learn a deeper visual representation with more semantic relevance and discrimination, rather than relying solely on the shallow features of the task reward. This improves the preset visual reinforcement learning model's ability to understand complex scenes and the sample training efficiency, reduces the ineffective exploration of the preset visual reinforcement learning model in the environment, and thus improves the training efficiency and the decision-making performance of the model. In addition, by combining the self-supervision loss and the target strategy loss, the decision-making ability of the policy decoder is further optimized to ensure that the visual reinforcement learning model has both semantic generalization and environmental adaptability. In summary, this application can improve the efficiency of model training and improve the decision-making performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 Schematic diagram of the architecture of a model training system based on visual reinforcement learning provided in an embodiment of the present application;
[0053] Figure 2 This is a flowchart of a model training method based on visual reinforcement learning provided in an embodiment of the present application;
[0054] Figure 3This is an overall flow chart of the model training method based on visual reinforcement learning provided in an embodiment of the present application;
[0055] Figure 4 Schematic diagram of the functional modules of the model training device based on visual reinforcement learning provided in an embodiment of the present application;
[0056] Figure 5 This is a schematic diagram of the hardware structure of the computer device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0057] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0058] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0060] Vision-based reinforcement learning (VRL) is an AI technology that learns decision-making policies from high-dimensional visual observations (such as raw pixel inputs). It is widely used in scenarios such as autonomous driving, embodied intelligence, and autonomous mobile robots. Reinforcement learning models use a visual encoder to compress raw images into a low-dimensional latent space representation and a policy decoder to output control actions based on this low-dimensional latent space representation.
[0061] Related technologies typically use a temporal difference loss approach to train models based on visual reinforcement learning. Specifically, the difference between the estimated value of the current state and the actual observed immediate reward plus the estimated value of the next state is used as a loss signal to drive the model's updates to the visual encoder and policy decoder. However, this optimization approach causes the visual encoder to learn shallow features that overly rely on task rewards, failing to capture high-level semantic associations between objects and actions in the scene. This reduces the model's ability to understand complex scenes, which in turn reduces the efficiency of model training and the model's decision-making performance.
[0062] Based on this, the embodiments of the present application provide a model training method, device, equipment and medium based on visual reinforcement learning, which can improve the efficiency of model training and improve the decision-making performance of the model.
[0063] The model training method, device, equipment and medium based on visual reinforcement learning provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the model training system based on visual reinforcement learning in the embodiments of the present application is described.
[0064] Please refer to Figure 1 In some embodiments, an embodiment of the present application provides a model training system based on visual reinforcement learning, including a terminal 11 and a server 12.
[0065] For example, the terminal 11 may be a high-performance embedded computing device, an edge computing device, or a mobile device with strong computing capabilities, such as an autonomous mobile robot, an on-board computing unit, or an embedded component of an embodied intelligent system. The server 12 may be a cloud computing platform or a high-performance computing cluster, such as a graphics processing unit (GPU) server cluster, a cloud computing virtual machine instance, or a distributed training platform.
[0066] In some embodiments, during the training phase, terminal 11 may be responsible for collecting raw image data and performing necessary preprocessing (e.g., extracting a three-frame observation data sequence), interacting with the environment to generate sample data, and then sending the processed sample data to server 12, where it utilizes a large visual language model (e.g., the CLIP series model) deployed on server 12 to generate high-level semantic information and self-supervisory signals. Server 12 then performs complex computational tasks, including representation distillation, self-supervised learning, and reinforcement learning optimization, to adjust the parameters of the preset visual reinforcement learning model to obtain the target visual reinforcement learning model. The optimized model parameters are fed back to terminal 11 by server 12, allowing the preset visual reinforcement learning model (e.g., a combination of a visual encoder and a policy decoder) in terminal 11 to be solidified and updated, thereby enabling continuous evolution and enhanced adaptability when independently executing action decisions in the subsequent application phase.
[0067] The model training method based on visual reinforcement learning in the embodiments of the present application can be illustrated by the following examples.
[0068] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0069] In the embodiment of the present application, the model training device based on visual reinforcement learning will be described from the perspective of the model training device based on visual reinforcement learning. The model training device based on visual reinforcement learning can be integrated into a computer device. Figure 2 , Figure 2 This is a flowchart of the steps of the model training method based on visual reinforcement learning provided in an embodiment of the present application. In the embodiment of the present application, the model training device based on visual reinforcement learning is specifically integrated on a terminal or server as an example. When the processor on the terminal or server executes the program instructions corresponding to the model training method based on visual reinforcement learning, the specific process is as follows:
[0070] Step 101: obtain a sample image frame and corresponding semantic category information, and input the semantic category information into the initial first convolution kernel of the visual large language model to extract text features, obtain the first convolution kernel parameter of the first feature convolution kernel after extracting text features, and input the sample image frame into the first feature convolution kernel to obtain a first feature heat map.
[0071] In some embodiments, in order to provide a more accurate and semantically relevant visual representation, the sample image frame and its corresponding semantic category information can be input into the initial first convolution kernel of the Large Vision-Language Model (LVLM) to perform text feature extraction and further generate the first convolution kernel parameters and the first feature heat map to generate a text-guided convolution kernel and feature heat map as self-supervisory signals, thereby enabling the preset visual reinforcement learning model to learn valuable decision-making information from high-dimensional visual observations.
[0072] The sample image frame can be obtained from the experience pool, and the preset visual reinforcement learning model can collect three consecutive frames of three-color (Red Green Blue, RGB) image observation sequences in the visual reinforcement learning task, which is recorded as ,in, Represents the image frame at the current moment, which is used to provide spatiotemporal visual information to support action reasoning and semantic analysis.
[0073] Among them, the semantic category information can be a high-level semantic label generated by the reasoning process of the visual language model, denoted as .
[0074] The visual language model can be a multimodal model based on the CLIP series (such as CLIP-S4), which can be used to jointly encode features of text and images. It should be noted that the visual language model here is different from the visual language model that generates semantic category information. The model that generates semantic category information is primarily an inference model.
[0075] Among them, the initial first convolution kernel can be the text-guided convolution kernel weight directly generated by the visual large language model based on semantic category information, which can serve as a feature detector to capture visual patterns (such as edges or textures) that are strongly related to semantic labels. Its parameters are output by the text encoder of the visual large language model (based on the CLIP series).
[0076] The first feature convolution kernel can be a convolution kernel derived from inputting semantic category information. This helps the model better capture image features that match the input text description, thereby achieving more accurate semantic association and feature representation. For example, the first feature convolution kernel can be abstracted into a transferable visual filter (such as a vehicle detector), allowing the lightweight model to learn feature patterns that are strongly relevant to the task.
[0077] The first convolution kernel parameters can be the weight matrix parameters of the first feature convolution kernel, specifically a set of values for the convolution kernel in the spatial dimension (such as 3x3 or 5x5) and the channel dimension. The first convolution kernel parameters characterize the model's feature extraction capability. Specifically, for input semantic category information, the visual large language model can use its internal mechanism to map it to a set of convolution kernel parameters suitable for feature extraction. This mapped set of convolution kernel parameters is the first convolution kernel parameters.
[0078] Among them, the first feature heat map can be a feature map generated after the sample image frame is input into the first feature convolution kernel, which shows the spatial area in the sample image frame that is most relevant to the semantic category information (such as the position of "left lane vehicle").
[0079] For example, the semantic category information corresponding to the sample image frame (such as "vehicle in the left lane") can be input into the text encoder of the visual large language model, and the initial first convolution kernel can be updated according to the text semantics to generate a task-adapted first feature convolution kernel, whose parameters encode visual patterns that are strongly associated with semantics (such as vehicle edge texture).
[0080] Furthermore, the original sample image frame can be processed using the first feature convolution kernel that has been optimized with semantic information. Specifically, each sample image frame (i.e. ) is input into the first feature convolution kernel to generate a first feature heatmap. The first feature heatmap shows the regions of the sample image frame that are most important for a particular semantic category. For example, in the scenario of detecting a pedestrian crossing the road, the first feature heatmap might highlight the location of the pedestrian and the part of the road they are about to cross.
[0081] Through the above method, it is convenient to provide high-level semantic self-supervision signals to the visual encoder of the preset visual reinforcement learning model, so that the preset visual reinforcement learning model can focus more on task-related areas (such as key objects in traffic scenes) during the feature extraction process, thereby enhancing the discrimination and robustness of the representation, and establishing an explainable optimization path for the subsequent distillation training of the preset visual reinforcement learning model.
[0082] In some embodiments, to provide more accurate and semantically relevant contextual information, three spatiotemporally continuous image frames, corresponding action data, and preset semantic cues may be obtained and input into a first large visual language model, also known as an inference large model, to drive reliable inference of scene semantics. For example, "obtaining sample image frames and corresponding semantic category information" in step 101 may include:
[0083] (101.1) Acquire sample image frames, wherein the sample image frames include a first image frame at a current moment, a second image frame at a second moment, and a third image frame at a third moment, wherein the second moment is subsequent to the current moment, and the third moment is subsequent to the second moment;
[0084] (101.2) Obtaining sample motion data corresponding to the first image frame and reference sample motion data corresponding to the second image frame;
[0085] (101.3) Obtaining a preset semantic cue, and inputting the semantic cue, the first image frame, the second image frame, the third image frame, the sample action data, and the reference sample action data into a first visual language model to obtain semantic category information corresponding to the sample image frame.
[0086] Among them, the first image frame can be the RGB image observation collected at the current moment, that is, the original pixel data obtained at time step t, reflecting the immediate environmental state, which can be used to provide the latest visual scene information as the core input of the first visual large language model reasoning.
[0087] Among them, the second image frame can be an RGB image observation collected at the second moment (i.e., the moment after the current moment), which is used to provide time continuity information and support the first visual large language model to analyze scene changes caused by actions.
[0088] The third image frame may be an RGB image observation collected at a third moment (i.e., a moment after the second moment), and is used to capture the dynamic evolution of the scene.
[0089] Among them, the sample action data can be the execution action of the preset visual reinforcement learning model corresponding to the first image frame, specifically the discrete or continuous control signal (such as steering angle) output by the policy decoder of the preset visual reinforcement learning model at the current moment, which serves as the action context input for inference by the first visual large language model.
[0090] Among them, the reference sample action data can be the execution action of the preset visual reinforcement learning model corresponding to the second image frame, which can be historical action data, and together with the sample action data constitute an action sequence, which is used by the first visual large language model to infer the causal impact of the action on the scene.
[0091] Semantic prompts can be a set of preset text rules, including task definitions, common sense rules, and thought chain guidance (such as "analyze the location of traffic objects"). They can also be fixed prompt templates that drive the First Vision Large Language Model through contextual learning, ensuring the standardization and repeatability of semantic reasoning.
[0092] The first visual large language model may be a visual large language model for reasoning semantic categories, which may be an inference engine based on multimodal pre-training (such as GPT-4 or LLaMA series).
[0093] In some embodiments, the dynamic evolution of the environment can be captured by capturing three consecutive frames of images to form a time series. , to solve the motion trajectory. Among them, the first image frame is the observation at the current moment (such as the real-time camera image of an autonomous vehicle); the second image frame is the observation at the previous moment (historical state benchmark); the third image frame is the observation of the first two moments (providing acceleration information of state change).
[0094] Furthermore, the temporal coupling data (i.e., sample action data) of the preset visual reinforcement learning model can be introduced. and reference sample action data ) to distinguish between environmental changes and active intervention. Specifically, the sample action data can be the action performed by the preset visual reinforcement learning model at the current moment (such as turning the steering wheel 30 degrees), and the reference sample action data can be the action at the previous moment (such as the braking force in the previous frame).
[0095] By putting the sample image frame When input together with sample action data, the first visual large language model can accurately identify causal changes caused by actions (such as vehicle position offset after turning) and autonomous environmental changes (such as the movement of other vehicles).
[0096] Furthermore, to enhance the first visual large language model's ability to understand scenes, semantic cues can be provided as input. Specifically, semantic cues can be textual information describing task objectives or scene characteristics, such as "avoid pedestrians" or "find a parking space." These semantic cues are first set and then fed into the pre-trained first visual large language model along with previously acquired sample image frames and sample action data (or the sample image frames and sample action data can be integrated into the semantic cues). The first visual large language model leverages its powerful cross-modal reasoning capabilities to analyze these inputs and output information about the sample image frames, extracting semantic category information from this output. For example, for an image frame containing a pedestrian, the output semantic category might be "pedestrian area," while for a scene with a traffic light, it might be "traffic light status." Depending on the actual situation, the output content can vary. This not only improves the model's understanding of complex scenes but also enhances the accuracy and adaptability of its decision-making.
[0097] Specifically, semantic cues need to be input into the first visual language model at the beginning so that the model can understand the task to be processed. Subsequently, only three consecutive frames of images (sample image frames) and corresponding sample action data need to be input. In this way, the model can reason according to the thinking chain part in the semantic cues. Finally, it is only necessary to extract semantic category information from the output content of the model.
[0098] Through the above method, high-level semantic category information related to the task can be effectively extracted from the sample image frames, that is, the visual features related to the decision can be accurately captured, which solves the semantic loss problem of reinforcement learning in complex scenarios (such as unseen traffic environments), and enables the subsequent distillation stage to obtain high-quality self-supervisory signals.
[0099] Step 102: Process the sample image frame using the initial second feature convolution kernel included in the visual encoder of the preset visual reinforcement learning model to obtain the second convolution kernel parameters and the second feature heat map of the second feature convolution kernel after processing the image.
[0100] In some embodiments, in order to extract richer and more dynamic visual features, the input image frame can be processed by the visual encoder of a vision-based reinforcement learning (VRL) agent (i.e., a preset vision reinforcement learning model) to generate task-related feature representations and provide an optimization basis for subsequent representation distillation.
[0101] The visual encoder can be a preset image processing module of a preset visual reinforcement learning model (reinforcement learning agent), which can convert sample image frames (three consecutive frames) into ) is compressed into a latent space representation and a spatial feature map. The visual encoder consists of convolutional layers and fully connected layers, which process the input to generate a second feature heat map and a one-dimensional feature latent variable obtained after the second feature heat map passes through the multilayer perceptron (MLP) layer. .
[0102] Among them, the initial second feature convolution kernel can be the initial weight set of the convolution layer in the visual encoder, which can be a parameter matrix of spatial dimension (such as 3x3 or 5x5) and channel dimension, used to extract low-level features (such as edges or textures) from the sample image frame.
[0103] The second feature convolution kernel may be a convolution kernel obtained by updating after the initial second feature convolution kernel processes the sample image frame.
[0104] Among them, the second convolution kernel parameter can be a specific weight value of the second feature convolution kernel, including a spatial filter coefficient and a channel scaling factor.
[0105] Among them, the second feature heat map can be a feature map corresponding to the current time t generated after the sample image frame is processed by the second feature convolution kernel, which can highlight the area in the sample image frame that is related to the task semantics (such as the position of the "left lane vehicle").
[0106] In some embodiments, sample image frames (multiple consecutive frames) can be fed as input to a visual encoder of a pre-set visual reinforcement learning model. In the visual encoder, a convolution operation is performed on the input sample image frame using an initial second-feature convolution kernel. Based on the result of the convolution operation, the weight parameters of the initial second-feature convolution kernel are adjusted to better capture key information in the image. After the adjustment is completed, the second-feature convolution kernel is obtained, and the updated parameter values of the second-feature convolution kernel, i.e., the second convolution kernel parameters, can be obtained.
[0107] Furthermore, after the sample image frame is processed by the convolution kernel of the visual encoder, a second feature heat map can be generated. After that, the second feature heat map is passed through the multilayer perceptron (MLP) layer to obtain a one-dimensional feature latent variable. , that is, the sample state data corresponding to the sample image frame.
[0108] For example, let's say you want to train a pre-set visual reinforcement learning model for a self-driving car to recognize pedestrians and avoid collisions. First, you can input a series of road scene image frames containing pedestrians (for example, three consecutive frames: ) is fed to the visual encoder. Each frame is convolved with the initial second feature convolution kernel preset in the visual encoder to preliminarily detect pedestrian features in the image. The second convolution kernel parameters and second feature heat map are then obtained for the processed image. For example, the second feature special case map can display areas in the sample image frame that are considered important by the preset visual reinforcement learning model, facilitating subsequent understanding and improvement of the model's decision-making process.
[0109] By obtaining the second convolution kernel parameters and the second feature heat map, the representation quality and generalization ability of the current preset visual reinforcement learning model can be accurately known, which facilitates subsequent adjustments.
[0110] Step 103: construct a first distillation loss based on the difference between the first convolution kernel parameters and the second convolution kernel parameters, and construct a second distillation loss based on the difference between the first feature heat map and the second feature heat map.
[0111] In some embodiments, in order to transfer the semantic representation capability of the large visual language model to the visual encoder of the preset visual reinforcement learning model, a first distillation loss can be constructed to force the convolution kernel parameters of the preset visual reinforcement learning model to be aligned with the text-guided convolution kernel (i.e., the first convolution kernel parameters) generated by the large visual language model; and a second distillation loss can be constructed to force the feature heat map of the preset visual reinforcement learning model to be aligned with the semantic heat map generated by the large visual language model, thereby improving the spatial semantic consistency of the preset visual reinforcement learning model and providing an interpretable and robust representation basis for reinforcement learning decisions.
[0112] Among them, the first distillation loss can be a loss function constructed based on the distance between the first convolution kernel parameters and the second convolution kernel parameters (such as L2 distance). Minimizing the first distillation loss can force the underlying feature extractor of the preset visual reinforcement learning model to learn filtering patterns that are strongly related to semantic categories (such as edge detection), thereby realizing the transfer of semantic knowledge of the large visual language model to the lightweight preset visual language model, and realizing efficient training of the preset visual language model.
[0113] The second distillation loss can be a loss function constructed based on the distance (e.g., L2 distance) between the first and second feature heatmaps. Minimizing the second distillation loss minimizes the spatial distribution difference between the feature maps of the preset visual reinforcement learning model and the visual large language model. This forces the feature maps of the preset visual reinforcement learning model to align with the first feature heatmap after channel aggregation, thereby improving the semantic spatial consistency of the representation.
[0114] In some embodiments, the first distillation loss It is constructed based on the difference between the first convolution kernel parameters and the second convolution kernel parameters. Specifically, By comparing the visual language model with the semantic category information The first feature convolution kernel and multiple second feature convolution kernels corresponding to multiple sample image frames in the corresponding preset visual reinforcement learning model The mean The specific calculation process is as follows:
[0115] ;
[0116] Where t represents the time step, is the set of time steps, that is, the total number of all sample image frames; Indicates the number of output channels, c indicates the channel; Represents the second feature convolution kernel of the k-th sample image frame.
[0117] Finally, by calculating the difference in the second norm (also known as the Euclidean distance) between the two and summing them up, we can obtain the first distillation loss. By minimizing the first distillation loss, we can force the second feature convolution kernel of the preset visual reinforcement learning model to approach the first feature convolution kernel of the visual large language model, thereby improving the feature extraction performance of the preset visual reinforcement learning model.
[0118] In some embodiments, the second distillation loss is constructed based on the difference between the first feature heat map and the second feature heat map. Specifically, by comparing the semantic category information in the visual language model The calculation is done by taking the first feature heat map and the average of multiple second feature heat maps corresponding to multiple sample image frames in the preset visual reinforcement learning model. The specific calculation process is as follows:
[0119] ;
[0120] in, represents the second distillation loss (heat map distillation loss), which is used to measure the difference between the visual large language model and the preset visual reinforcement learning model in the feature heat map; t represents the time step, is the set of time steps, that is, the total number of all sample image frames; Indicates the number of output channels, c indicates the channel; Represents the sample image frame in the visual language model Related first feature heatmap; Represents the second feature heatmap after passing through the visual encoder of the preset visual reinforcement learning model.
[0121] By minimizing the second distillation loss, the feature heat map generated by the preset visual reinforcement learning model can be made closer to the heat map of the visual large language model, thereby improving the performance of the preset visual reinforcement learning model in feature localization and semantic understanding.
[0122] This approach effectively transfers the powerful semantic understanding and feature extraction capabilities of the large visual language model to the pre-built visual reinforcement learning model. This not only improves the pre-built visual reinforcement learning model's ability to capture and understand key features, but also enhances its decision-making accuracy and robustness in complex tasks. This significantly improves the pre-built visual reinforcement learning model's overall performance and adaptability, enabling the system to perform tasks more accurately in diverse and dynamic environments.
[0123] In step 104, a preset self-supervised model is used to predict the state and reward of the sample action data and sample state data corresponding to the second feature heat map at the next moment, to obtain the prediction result of each sample image frame, and a self-supervised loss is constructed based on the multiple prediction results corresponding to the multiple sample image frames.
[0124] In some embodiments, in order to further optimize the representation quality of the preset visual reinforcement learning model, a preset self-supervised model (such as a state transition predictor and a reward predictor) can be used to perform dynamic environment modeling on the distilled visual representation, forcing the latent variables to contain temporal continuity and task value association information, so as to solve the problem that the pure distillation method may ignore the dynamic changes of the environment, and improve the robustness and decision-making guidance of the representation in complex scenarios (such as embodied intelligent systems).
[0125] Among them, the preset self-supervised model may include a state transition predictor and a reward predictor in a self-supervised learning (SSL) task, whose function is to predict the predicted state data and predicted reward data at the next moment based on the current hidden state (i.e., sample state data) and sample action data.
[0126] The sample action data can be information about actions performed by a preset visual reinforcement learning model (e.g., a robot or autonomous mobile device) under specific conditions. The sample action data can be obtained from a sample training set randomly selected from an experience pool based on the second thermal signature (or sample image frames).
[0127] Among them, the sample state data can be a one-dimensional latent variable after the second feature heat map is compressed by the MLP layer, which is used to represent the semantic compression representation of the preset visual reinforcement learning model in the current environment state.
[0128] Among them, the prediction results can be two predicted values output by the preset self-supervisory model, specifically predicted state data and predicted reward data.
[0129] Among them, the self-supervised loss can be two loss functions constructed based on the difference between the predicted results and the true values, specifically the first self-supervised loss (state transfer loss) and the second self-supervised loss (reward loss). These two losses will be introduced in detail below.
[0130] In some implementations, to enhance the visual reinforcement learning model's ability to predict future states and rewards, a state transition predictor within a pre-defined self-supervised model can be used. This model receives current sample state data and sample action data as input, generates a prediction for the next state, and outputs the predicted state data. Training this model allows the pre-defined visual reinforcement learning model to better understand the dynamics of the environment—specifically, how the environment will change after a specific action is taken—facilitating subsequent long-term planning and decision-making.
[0131] Furthermore, a reward predictor in the pre-set self-supervised model receives the current sample state data and sample action data as input, outputs a prediction of the reward at the next moment, and outputs the predicted reward data. By training this model, the pre-set visual reinforcement learning model can more accurately predict the immediate feedback from its actions, thereby optimizing its action strategy to maximize cumulative rewards. This allows the pre-set visual reinforcement learning model to quickly learn effective behavior patterns in complex environments.
[0132] Specifically, for each input sample image frame, the preset self-supervised model generates a corresponding prediction result (including predicted state data and predicted reward data). By comparing these predictions with the actual results (the actual target sample state data and target sample reward data, obtained from the experience pool), the error for each sample image frame can be calculated. The errors for all sample image frames are then aggregated and averaged to construct a self-supervised loss, which captures the degree of discrepancy between the results predicted by the preset visual reinforcement learning model and the actual situation. By minimizing the self-supervised loss, the preset visual reinforcement learning model's ability to predict future states and rewards can be gradually improved, thereby enhancing decision quality.
[0133] By constructing a self-supervised loss, we can introduce additional self-supervisory signals, which facilitates the subsequent optimization of the representation capabilities of the preset visual reinforcement learning model, enabling the model to more accurately understand the dynamics of the environment and predict future states and rewards.
[0134] In some embodiments, in order to improve the ability of the preset visual reinforcement learning model to predict future environmental dynamics and enhance its strategy optimization efficiency and decision-making accuracy in complex task scenarios, the preset visual reinforcement learning model can be trained by constructing a self-supervised loss, thereby achieving a more efficient and stable learning process and stronger generalization performance. Exemplarily, the self-supervised loss includes a first self-supervised loss and a second self-supervised loss. Step 104 may include:
[0135] (104.1) Using a preset self-supervised model, for each sample image frame, predict the state and reward at the next moment for the sample action data and sample state data corresponding to the second feature heat map, and obtain the predicted state data and predicted reward data for each sample image frame at the next moment;
[0136] (104.2) Obtain target sample state data and target sample reward data at the next moment, and construct a first self-supervised loss based on the difference between the predicted state data and the target sample state data of the multiple sample image frames, and construct a second self-supervised loss based on the difference between the predicted reward data and the target sample reward data of the multiple sample image frames.
[0137] The predicted state data may be a state estimation value at the next time t+1 of the current time t output by a state transition predictor in a preset self-supervisory model.
[0138] The predicted reward data may be an immediate reward estimate output by a reward predictor in a preset self-supervised model.
[0139] The target sample state data may be the actual hidden state observation value at the next moment t+1 after the current moment t.
[0140] The target sample reward data may be the actual reward value at the next moment.
[0141] The first self-supervised loss may be a loss function constructed based on the L2 distance between the predicted state data and the target sample state data.
[0142] Among them, the second self-supervised loss can be a loss function constructed based on the L2 distance between the predicted reward data and the target sample reward data.
[0143] For example, it can be preset that the self-supervisory model can receive sample action data for each sample image frame. and sample status data As input, the state transition predictor outputs the predicted state data for the next moment , and the predicted reward data obtained by predicting the next moment through the reward predictor output .
[0144] Furthermore, the actual target sample state data at the next moment obtained when the preset visual reinforcement learning model interacts with the environment based on the sample action data and sample state data can be obtained. , in order to construct the first self-supervisory loss The specific process is as follows:
[0145] ;
[0146] Furthermore, the actual target sample reward data at the next moment obtained when the preset visual reinforcement learning model interacts with the environment based on the sample action data and sample state data can be obtained. , in order to construct the second self-supervisory loss The specific process is as follows:
[0147] ;
[0148] Through the above steps, the preset self-supervised model can be effectively used to predict the state and reward at the next moment, so as to optimize the prediction ability of the preset visual reinforcement learning model by constructing the first self-supervised loss and the second self-supervised loss, improve the preset visual reinforcement learning model's understanding of future environmental dynamics, enhance the accuracy and stability of its decision-making strategy, and thus achieve better performance in complex tasks.
[0149] In step 105, a strategy decoder of a preset visual reinforcement learning model is used to perform action value and action sampling calculations based on sample action data and sample state data to obtain calculation results for each sample image frame, and a target strategy loss is constructed based on multiple calculation results corresponding to multiple sample image frames.
[0150] In some embodiments, in order to improve the decision-making accuracy and efficiency of the preset visual reinforcement learning model in complex scenarios (such as autonomous driving), a policy decoder and a value network (or action-value function) can be used to perform policy evaluation and action distribution sampling calculations on sample state data and sample action data to evaluate the quality of the current policy and explore possible better action choices, ensuring that the preset visual reinforcement learning model can make decisions in complex environments that are both exploratory and maximize long-term cumulative rewards.
[0151] Among them, the calculation results can be two types of output values generated during the strategy decoding process, specifically the expected cumulative reward of the action and the action probability.
[0152] Among them, the target strategy loss can include two types of losses, specifically value network sub-loss and strategy sub-loss.
[0153] In some embodiments, a policy decoder in a pre-set visual reinforcement learning model can be used to calculate action values and action sampling based on sample action data and sample state data. Specifically, the action value and state value can be evaluated using the Bellman equation to calculate the value network sub-loss and policy sub-loss. The value network sub-loss can be used to measure the difference between the current action value and the reward and state value at the next moment; while the policy sub-loss combines the difference in action probability distribution and action value to ensure that the action selected for each sample image frame has both high expected rewards and conforms to the optimal policy, thereby improving the overall decision quality.
[0154] By constructing a target policy loss, the model not only improves its accurate evaluation of the current policy's performance but also promotes an effective balance between exploration and exploitation, ensuring that the pre-set visual reinforcement learning model can learn the optimal policy that maximizes long-term rewards. Furthermore, by optimizing the target policy loss, the pre-set visual reinforcement learning model's adaptability and decision-making efficiency in complex and dynamic environments can be enhanced, enabling it to make more accurate and efficient decisions in unknown environments.
[0155] In some embodiments, in order to achieve accurate evaluation of the action value of the preset visual reinforcement learning model and efficient strategy learning, on the one hand, the expected cumulative reward of the action can be calculated based on the sample state and action data, and a value network sub-loss can be constructed to improve the model's ability to predict future returns; on the other hand, the action probability distribution can be generated by the policy decoder and action sampling can be performed, and then the strategy sub-loss can be constructed based on the sampling results to optimize the exploration and utilization balance of the strategy itself. In summary, constructing a target strategy loss can improve the decision-making accuracy, strategy stability and long-term cumulative reward performance of the visual reinforcement learning model in complex task scenarios. For example, the target strategy loss can include a value network sub-loss and a strategy sub-loss, and step 105 can also include:
[0156] (105.1) Using the value network included in the policy decoder of the pre-set visual reinforcement learning model, for each sample image frame, the action value is calculated based on the sample action data and sample state data to obtain the corresponding expected cumulative reward for the action;
[0157] (105.2) Construct a value network sub-loss based on the expected cumulative rewards of multiple actions corresponding to multiple sample image frames;
[0158] (105.3) For each sample image frame, obtain the action probability distribution generated by the strategy decoder based on the sample state data, perform action sampling based on the action probability distribution, obtain sampled action data, and calculate the action probability based on the sampled action data;
[0159] (105.4) Construct a policy sub-loss based on multiple action probabilities corresponding to multiple sample image frames.
[0160] Among them, the value network can be the action value function in the policy decoder, whose function is to evaluate the future expected cumulative reward of executing sample action data under sample state data.
[0161] The expected cumulative reward of an action can be the state-action value output by the value network, which is used to represent the expected value of the discounted cumulative reward of executing sample action data starting from time t.
[0162] Among them, the value network sub-loss can be a loss function that minimizes the difference between the expected cumulative reward of the action and the temporal difference target.
[0163] The action probability distribution may be a conditional probability distribution output by a policy decoder, which is used to represent the probability density of optional actions under sample state data (such as a Gaussian distribution of steering angles).
[0164] The sampled action data may be a specific action value randomly sampled from the action probability distribution (eg, a sampled steering angle of 30°).
[0165] The action probability can be the probability estimate of the sampled action by the policy decoder, which can be used to optimize the action selection tendency in the policy gradient calculation.
[0166] Among them, the policy sub-loss can be a loss function that maximizes the action value while retaining the policy randomness.
[0167] In some embodiments, for a given sample state data and sample action data , which can be calculated through the value network Next execution The expected cumulative reward for future actions that can be obtained Assume there is a self-driving car task, the agent needs to make decisions based on the current traffic conditions (sample state data ) decide whether to turn (sample action data The value network predicts the expected cumulative reward of the action that this decision may bring in the future (such as successfully reaching the destination without collision) based on the input state and action information. If the predicted result is positive, it means that taking this action will bring a higher expected cumulative reward. Furthermore, the instant reward corresponding to each sample image frame can be obtained. The updated state value at the next moment and the expected cumulative reward of the action are used to construct the value network sub-loss. This process will be further expanded in the subsequent examples. The formula is as follows:
[0168] ;
[0169] in, represents the expected cumulative reward of the action, represents sample action data, Represents sample status data, Indicates immediate reward; represents the adjustment parameter, Represents the updated state value. By minimizing the value network sub-loss, the model can simultaneously optimize the prediction of immediate rewards and long-term cumulative rewards, improving overall performance and decision-making efficiency.
[0170] Furthermore, the policy decoder can be used to determine the current sample state data. Generate a probability distribution of an action, randomly extract a sample action data as the actual action, record the action probability of the action, and obtain . Then, get the preset entropy parameters , based on the product of the preset entropy parameter and the action probability, the second product is obtained. Further, a preset target action expected cumulative reward can be obtained , the third mean of each sample image frame is obtained by the difference between the second product and the expected cumulative reward of the target action.
[0171] Therefore, by taking the expectation of multiple third differences corresponding to multiple sample image frames (all sample image frames contained in the training data set corresponding to the current task stage), the strategy sub-loss can be obtained. The specific formula is as follows:
[0172] ;
[0173] in, represents the preset entropy parameter, Represents sampling action data, Represents sample status data, Represents the expected cumulative reward for the target action. By constructing a policy sub-loss, the pre-set visual reinforcement learning model can make better long-term plans even in the face of uncertainty, while maintaining a certain level of exploration capability and avoiding falling into local optimality.
[0174] The value network evaluates the action value of each sample image frame to obtain the expected cumulative reward for the action. This is then used to construct a value network sub-loss based on the expected cumulative reward for multiple samples, effectively improving the model's accuracy in predicting future rewards. Simultaneously, the policy decoder generates an action probability distribution based on state data, and performs action sampling and probability calculations, enabling the pre-set visual reinforcement learning model to maintain a balance between exploration and utilization. Finally, the policy sub-loss is constructed from the action probabilities of multiple samples to guide the direction of policy optimization. In summary, constructing a value network sub-loss and a policy sub-loss enables the coordinated optimization of the value function and the policy function, improving the decision-making stability, long-term cumulative reward performance, and policy exploration efficiency of the visual reinforcement learning model in complex environments, enhancing the model's generalization ability and training convergence.
[0175] In some embodiments, in order to make the expected cumulative rewards of actions calculated by the value network approach the actual long-term benefits, that is, to improve the accuracy of the value network's value assessment, the difference between the expected cumulative rewards of actions of the current strategy and the target value data can be evaluated by calculating the value network sub-loss, so as to guide the preset visual reinforcement learning model to optimize its ability to predict future rewards, thereby improving decision quality and ensuring that the preset visual reinforcement learning model can effectively learn and adapt in complex environments to maximize long-term cumulative rewards. Exemplarily, (105.2) may include:
[0176] (105.2.1) For each sample image frame, obtain the immediate reward of the environment feedback after executing the sample action data under the sample state data;
[0177] (105.2.2) Obtain the updated sample state data and updated sample action data for each sample image frame at the next moment, and calculate the updated state value at the next moment based on the updated sample state data and updated sample action data;
[0178] (105.2.3) Obtain preset adjustment parameters and adjust the updated state value according to the adjustment parameters to obtain the target state value;
[0179] (105.2.4) Obtain target value data based on the sum of the immediate reward and the target state value;
[0180] (105.2.5) obtaining a first difference value based on the difference between the expected cumulative reward of the action and the target value data;
[0181] (105.2.6) Construct a value network sub-loss based on the mean of multiple first differences corresponding to multiple sample image frames.
[0182] Among them, the immediate reward can be a scalar value of the environment feedback, which is expressed in the sample state data Execute sample action data The immediate benefits obtained after the game (such as positive rewards for successfully avoiding obstacles).
[0183] The updated sample state data can be the sample action data executed by the preset visual reinforcement learning model. After that, the environment enters the next state .
[0184] The updated sample action data can be the action that the preset visual reinforcement learning model performs under the updated sample state data at the next moment. .
[0185] The updated state value can be the estimated value of the updated sample state data at the next moment. .
[0186] Among them, the adjustment parameter can be the discount factor , the specific value can be set according to the actual situation.
[0187] Among them, the target state value can be the discounted future value , which represents the expected long-term return starting from t+1.
[0188] Among them, the target value data can be a time series difference target .
[0189] The first difference may be a value obtained by subtracting the expected cumulative reward of the action from the target value data.
[0190] In some embodiments, for each sample image frame, a preset visual reinforcement learning model can be used to Execute sample action data , execute sample action data After that, you can get the immediate reward based on the immediate feedback from the environment. ( 、 、 can be directly obtained from the training data set selected by the experience pool).
[0191] Furthermore, we can obtain sample image frames from the training data set and execute After that, the environment enters the state of the next moment , and update the sample status data Under the preset visual reinforcement learning model, the new action is taken , thus, the updated state value at the next moment can be calculated .
[0192] Furthermore, the preset adjustment parameters can be obtained , and multiply it by the adjustment parameter and the updated state value , get the target state value to balance the immediate benefits and future value, and avoid short-sighted decisions (such as crashing for immediate rewards). Then, according to the sum of the immediate reward and the target state value , get the target value data.
[0193] Furthermore, for each sample image frame, the difference between the expected cumulative reward of the action and the target value data can be calculated to obtain the first difference Finally, the first differences corresponding to all sample image frames are added together and divided by the total number of all sample image frames to obtain the value network sub-loss. This can accurately evaluate the prediction error of the current value network, effectively improve the stability of the value network, and ensure that the model is optimized overall.
[0194] Specifically, calculate the value network sub-loss The formula is as follows:
[0195] ;
[0196] in, represents the expected cumulative reward of the action, represents sample action data, Represents sample status data, Indicates immediate reward; represents the adjustment parameter, Represents the updated state value. By minimizing the value network sub-loss, the model can simultaneously optimize the prediction of immediate rewards and long-term cumulative rewards, improving overall performance and decision-making efficiency.
[0197] By constructing a value network sub-loss that reflects the prediction error of the value network, the accuracy of the visual reinforcement learning model's estimation of the expected cumulative reward of an action can be effectively improved, enabling it to more accurately evaluate the long-term benefits of different actions in complex environments, significantly enhancing the ability of the preset visual reinforcement learning model to make high-quality decisions in unknown environments.
[0198] In some embodiments, to improve the preset visual reinforcement learning model's ability to predict future environmental dynamics and improve the accuracy of strategy optimization, the updated state value at the next moment can be accurately calculated based on the updated sample state data and the updated sample action data. This ensures that the model maximizes future cumulative rewards while maintaining strategy diversity and robustness, thereby achieving an exploration-exploitation balance for the reinforcement learning strategy. For example, "calculating the updated state value at the next moment based on the updated sample state data and the updated sample action data" in (105.2.2) may include:
[0199] (105.2.2.1) For each sample image frame, obtain the updated action probability distribution generated by the policy decoder for the updated sample state data;
[0200] (105.2.2.2) Sampling actions according to the next action probability distribution to obtain updated sampled action data, and determining the updated action probability of the updated sampled action data;
[0201] (105.2.2.3) Obtain a preset entropy parameter and obtain a first product based on the product of the updated action probability and the preset entropy parameter;
[0202] (105.2.2.4) Calculate the expected cumulative reward for the corresponding update action based on the update sample action data and the update sample state data using the value network contained in the policy decoder;
[0203] (105.2.2.5) obtaining a second difference based on the difference between the expected cumulative reward of the updated action and the first product;
[0204] (105.2.2.6) Obtain an updated state value at the next moment based on an average of multiple second difference values corresponding to multiple sample image frames.
[0205] The updated action probability distribution can be the conditional probability output by the policy decoder in the next state (i.e., updated sample state data), that is, the probability density of the optional action (such as the Gaussian distribution of the steering angle).
[0206] The updated sampled action data may be a specific action randomly sampled from the updated action probability distribution (eg, the sampled steering angle is -15°).
[0207] The updated action probability may be a probability estimate of the updated sampled action data by the policy decoder.
[0208] The preset entropy parameter may be a temperature coefficient, which is used to control the exploration intensity. For example, the preset entropy parameter may be 0.2, 0.3, and so on.
[0209] Among them, the expected cumulative reward of the update action can be the value network's long-term value estimate of the next state-action pair (that is, the updated sample state data and the updated sample action data).
[0210] The second difference may be a value obtained by subtracting the expected cumulative reward of the updated action from the first product.
[0211] For example, for each sample image frame, the policy decoder can update the sample state data according to Generate updated action probability distribution After that, the policy decoder can perform action sampling according to the updated action probability distribution to obtain the updated sampled action data . Then, get the preset entropy parameters , and calculate The first product is obtained. Furthermore, the expected cumulative reward of the update action can be directly calculated through the value network based on the updated sample state data and updated sample action data of the sample image frame obtained from the experience pool. Furthermore, the second difference can be obtained based on the difference between the expected cumulative reward of the updated action and the first product. Furthermore, the updated state value at the next moment can be obtained based on the average of multiple second difference values corresponding to multiple sample image frames. From this, we can determine that the formula for calculating the updated state value is as follows:
[0212] ;
[0213] in, Indicates the number of sample image frames.
[0214] Through the above methods, the quality of the current strategy can be measured to provide a basis for subsequent strategy adjustments. This ensures that the preset visual reinforcement learning model can maximize future rewards in the subsequent learning process while maintaining sufficient exploratory power to discover potentially better strategies. This helps to improve the adaptability and efficiency of the preset visual reinforcement learning model in complex environments.
[0215] In some implementations, to balance exploration and exploitation in reinforcement learning and promote more efficient learning and policy convergence, policy entropy can be calculated for each sample image frame to maintain policy randomness. Subsequently, by constructing a policy sub-loss, the policy decoder is driven to output a high-value action distribution and target action, thereby maximizing long-term rewards while retaining policy diversity, addressing overfitting issues in complex scenarios (such as unseen traffic obstacles), and improving decision generalization. For example, (105.4) may include:
[0216] (105.4.1) Obtain a preset entropy parameter, and for each sample image frame, obtain a second product based on the product of the preset entropy parameter and the action probability;
[0217] (105.4.2) Obtaining a preset target action expected cumulative reward, and obtaining a third difference based on the difference between the second product and the target action expected cumulative reward;
[0218] (105.4.3) Construct a strategy sub-loss based on the mean of multiple third differences corresponding to multiple sample image frames.
[0219] The second product can be the product of the preset entropy parameter and the action probability. It serves as a regularization term for the policy sub-loss and can be used to control the exploration intensity.
[0220] Among them, the expected cumulative reward of the target action can be the benchmark action value in the strategy optimization target, which represents the long-term reward estimate of the sample action data under the sample state data, to provide an optimization direction for the policy gradient and drive the strategy to shift toward the high-value action distribution.
[0221] Among them, the third difference can be the value obtained by subtracting the second product from the expected cumulative reward of the target action. By minimizing the third difference through gradient ascent, the policy network can simultaneously improve the action value expectation and policy randomness.
[0222] Specifically, we can first use the preset entropy parameter and action probability The product of , calculates the second product to measure the exploration of the strategy; then, by comparing the second product with the preset target action expected cumulative reward The third difference is obtained, which reflects the performance gap between the current strategy and the ideal strategy. Finally, a strategy sub-loss is constructed based on the average of multiple third differences corresponding to multiple sample image frames to guide the strategy optimization process, ensuring that the preset visual reinforcement learning model can maintain appropriate exploratory power while maximizing long-term cumulative rewards, thereby improving decision-making quality and adaptability in complex and uncertain environments.
[0223] For example, the formula for calculating the strategy loss is as follows:
[0224] ;
[0225] in, Represents the preset entropy parameter, which is used to balance the exploratory and deterministic nature of the strategy. It encourages the preset visual reinforcement learning model to maintain a certain degree of exploratory nature during the decision-making process and avoid premature convergence to the local optimal solution. Indicates that the policy decoder is in the sample state data The action data obtained by random sampling is adopted. Indicates the sample status data Next execution The probability of represents the expected cumulative reward of the target action.
[0226] By constructing a policy sub-loss, not only can the preset visual reinforcement learning model be optimized towards increasing long-term cumulative rewards, but it also ensures sufficient policy exploration to prevent premature entrapment of local optimal solutions. This effectively improves the adaptability and decision-making efficiency of the preset visual reinforcement learning model in complex and dynamic environments, enhances the model's learning stability and generalization capabilities, and enables the preset visual reinforcement learning model to make better decisions in unknown environments.
[0227] Step 106: Based on the first distillation loss, the second distillation loss, the self-supervision loss, and the target strategy loss, the parameters of the preset visual reinforcement learning model are adjusted to obtain a target visual reinforcement learning model.
[0228] In some embodiments, in order to build a reinforcement learning model with stronger generalization ability, higher training efficiency and better decision-making performance, the preset visual reinforcement learning model can be end-to-end parameter adjustment and optimization by integrating the first distillation loss, the second distillation loss, the self-supervision loss and the target policy loss, so that the model can learn from high-level semantic information, and also enhance the model's ability to understand the dynamic changes of the future environment, and achieve the optimal balance between exploration and utilization in the decision-making process, thereby obtaining an efficient and robust target visual reinforcement learning model.
[0229] Among them, the target visual reinforcement learning model can be the final model after the first distillation loss, the second distillation loss, the self-supervision loss and the target strategy loss are jointly optimized, including a visual encoder and a strategy decoder.
[0230] In some embodiments, after the preset visual reinforcement learning training is completed and the target visual reinforcement model is obtained, there is no need to use the visual large language model and the self-supervised model to assist in processing data.
[0231] In some embodiments, the parameters of the preset visual reinforcement learning model can be iteratively adjusted so that the above-mentioned first distillation loss, second distillation loss, self-supervision loss and target strategy loss are all minimized, that is, the model weights are gradually updated through optimization algorithms such as gradient descent until convergence, and finally an efficient and robust target visual reinforcement learning model is obtained.
[0232] For example, the trained target visual reinforcement learning model can be applied to fields such as autonomous driving, intelligent robot navigation, robotic arm control in industrial automation, autonomous flight and mission execution of drones, and collaborative management of devices in smart home systems. It has great potential in improving the level of intelligence and enhancing environmental adaptability.
[0233] Please refer to Figure 3 In some embodiments, combined Figure 3 This section introduces the overall process of this application.
[0234] For example, a sample image frame, corresponding sample action data, and predefined semantic cues can be input into a first large visual language model (inference model). This model generates an output through thought chain reasoning, and extracts the required semantic category information from the output. Simultaneously, the large visual language model (guidance model) can generate the first text-guided convolution kernel parameters and first feature heatmap based on the semantic category information, serving as the distilled self-supervisory signal.
[0235] Furthermore, the sample image frame can be Input visual encoder, output second thermal feature map And sample state information, also known as hidden state Based on the self-supervisory signal provided by the large visual language model, the first distillation loss (convolution kernel parameter alignment) and the second distillation loss (feature heat map alignment) are calculated to drive the visual encoder of the preset visual reinforcement learning model to learn high-level semantic association features.
[0236] Furthermore, the sample status data can be With sample action data Combined with the input of a preset self-supervised model (with trainable parameters), the state and reward at the next moment are predicted to construct a self-supervised loss and enhance the temporal continuity of the representation.
[0237] Furthermore, the sample status data can be The policy decoder is input, and the expected cumulative reward of each action is calculated through the value network, which constructs the value network sub-loss. After that, the action probability distribution is generated and sampled to obtain sampled actions, which are combined with the entropy parameter to construct the policy sub-loss.
[0238] Finally, the first distillation loss, second distillation loss, self-supervision loss, value network sub-loss, and policy sub-loss are jointly minimized to update the visual encoder and policy decoder parameters in the pre-set visual reinforcement learning model. After training is complete, the large visual language model and self-supervision branch are removed, achieving lightweight real-time decision-making.
[0239] Experimental results show that compared to traditional visual reinforcement learning methods combined with self-supervised learning (such as DtQ and DeepMDP), this application achieves higher episode rewards (Episode return) with the same number of training steps, has the fastest convergence speed, and improves sample training efficiency by over 40% (the reward exceeds the baseline when the number of training steps is reduced by half). At the same time, it solves the problem of violent fluctuations in the reward curve of other methods in complex scenarios.
[0240] In addition, the semantic perception ability of the visual reinforcement learning model of this application has been significantly enhanced, and the decision-making is more accurate. Specifically, the feature heat map generated by this application can accurately focus on the key areas corresponding to the task (such as "pedestrians on the right" and "vehicles in the left lane", etc.). For example, in an autonomous driving scenario, the error in capturing the position and movement trend of traffic participants by the thermal feature map is reduced by about 35%, enabling the policy decoder to output fine-grained actions based on highly discriminative semantic representations (such as the accuracy of obstacle avoidance steering angles is increased to ±2°), significantly improving the decision-making robustness of the target visual reinforcement learning model in complex dynamic environments.
[0241] Furthermore, experiments visualizing latent space distributions show that the state representations learned by this application exhibit clear cluster separation by semantic category (e.g., pedestrians, vehicles, and bicycles) in the latent space. This highly discriminative representation structure enables more efficient policy learning: in tests, the decision error rate for unseen scenarios was reduced by 42%, and the amount of interaction data required for policy convergence was only one-third of that required by traditional methods.
[0242] The embodiment of the present application obtains a sample image frame and corresponding semantic category information, and inputs the semantic category information into the initial first convolution kernel of the visual large language model to extract text features, thereby obtaining the first convolution kernel parameter of the first feature convolution kernel after extracting the text features, and inputs the sample image frame into the first feature convolution kernel to obtain a first feature heat map; processes the sample image frame through the initial second feature convolution kernel included in the visual encoder of the preset visual reinforcement learning model to obtain the second convolution kernel parameter of the second feature convolution kernel after processing the image and the second feature heat map; constructs a first distillation loss based on the difference between the first convolution kernel parameter and the second convolution kernel parameter, and constructs a first distillation loss based on the difference between the first feature heat map and the second feature heat map. Second distillation loss; through the preset self-supervised model, the state and reward of the sample action data and sample state data corresponding to the second feature heat map are predicted at the next moment to obtain the prediction result of each sample image frame, and the self-supervised loss is constructed according to the multiple prediction results corresponding to the multiple sample image frames; through the policy decoder of the preset visual reinforcement learning model, the action value and action sampling calculation are performed based on the sample action data and sample state data to obtain the calculation result of each sample image frame, and the target policy loss is constructed according to the multiple calculation results corresponding to the multiple sample image frames; based on the first distillation loss, the second distillation loss, the self-supervised loss and the target policy loss, the parameters of the preset visual reinforcement learning model are adjusted to obtain the target visual reinforcement learning model. In this way, by combining the representational capabilities of the large visual language model, the semantic information and reasoning capabilities of the large visual language model can be distilled into a lightweight preset visual reinforcement learning model, so that the preset visual reinforcement learning model can learn a deeper visual representation with more semantic relevance and discrimination, rather than relying solely on the shallow features of the task reward. This improves the preset visual reinforcement learning model's ability to understand complex scenes and sample efficiency, reduces the ineffective exploration of the preset visual reinforcement learning model in the environment, and thus improves the training efficiency and decision-making performance of the model. In addition, by combining self-supervision loss and target strategy loss, the decision-making ability of the policy decoder is further optimized to ensure that the visual reinforcement learning model has both semantic generalization and environmental adaptability. In summary, this application can improve the efficiency of model training and improve the decision-making performance of the model.
[0243] See also Figure 4 The embodiment of the present application further provides a model training device based on visual reinforcement learning, which can implement the above-mentioned model training method based on visual reinforcement learning. The model training device based on visual reinforcement learning includes:
[0244] An acquisition module 41 is configured to acquire a sample image frame and corresponding semantic category information, input the semantic category information into an initial first convolution kernel of a visual large language model to extract text features, obtain first convolution kernel parameters of the first feature convolution kernel after extracting text features, and input the sample image frame into the first feature convolution kernel to obtain a first feature heat map;
[0245] a processing module 42 configured to process the sample image frame using an initial second feature convolution kernel included in a visual encoder of a preset visual reinforcement learning model to obtain second convolution kernel parameters and a second feature heat map of the second feature convolution kernel after processing the image;
[0246] A construction module 43 is configured to construct a first distillation loss based on a difference between the first convolution kernel parameters and the second convolution kernel parameters, and to construct a second distillation loss based on a difference between the first feature heat map and the second feature heat map;
[0247] A prediction module 44 is configured to predict the state and reward at the next moment for the sample action data and sample state data corresponding to the second feature heat map using a preset self-supervised model, obtain a prediction result for each sample image frame, and construct a self-supervised loss based on multiple prediction results corresponding to multiple sample image frames;
[0248] A calculation module 45 is configured to calculate action value and action sampling based on the sample action data and the sample state data using a policy decoder of a preset visual reinforcement learning model, obtain a calculation result for each sample image frame, and construct a target policy loss based on multiple calculation results corresponding to multiple sample image frames;
[0249] The adjustment module 46 is used to adjust the parameters of the preset visual reinforcement learning model based on the first distillation loss, the second distillation loss, the self-supervision loss and the target strategy loss to obtain a target visual reinforcement learning model.
[0250] The specific implementation of the model training device based on visual reinforcement learning is basically the same as the specific embodiment of the model training method based on visual reinforcement learning described above, and will not be repeated here. On the premise of meeting the requirements of the embodiments of this application, the model training device based on visual reinforcement learning can also be provided with other functional modules to implement the model training method based on visual reinforcement learning in the above embodiment.
[0251] The present application also provides a computer device comprising a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned model training method based on visual reinforcement learning. The computer device can be any intelligent terminal, including a tablet computer and an in-vehicle computer.
[0252] See also Figure 5 , Figure 5 The hardware structure of a computer device according to another embodiment is shown. The computer device includes:
[0253] The processor 51 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.
[0254] The memory 52 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 52 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 52 and is called by the processor 51 to execute the model training method based on visual reinforcement learning in the embodiments of this application.
[0255] Input / output interface 53, used to implement information input and output;
[0256] Communication interface 54, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);
[0257] bus 55 , which transmits information between the various components of the device (e.g., processor 51 , memory 52 , input / output interface 53 , and communication interface 54 );
[0258] The processor 51 , the memory 52 , the input / output interface 53 and the communication interface 54 are connected to each other in communication within the device via a bus 55 .
[0259] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned model training method based on visual reinforcement learning.
[0260] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0261] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0262] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0263] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0264] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.
[0265] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0266] It should be understood that in the present application, "at least one (item)" and "several" are one or more, and "multiple" is two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions is any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0267] In the several embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0268] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0269] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0270] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0271] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.
Claims
1. A model training method based on visual reinforcement learning, characterized in that: The method comprises: Obtaining a sample image frame and corresponding semantic category information, inputting the semantic category information into an initial first convolution kernel of a visual large language model to extract text features, obtaining first convolution kernel parameters of the first feature convolution kernel after extracting text features, and inputting the sample image frame into the first feature convolution kernel to obtain a first feature heat map; Processing the sample image frame using an initial second feature convolution kernel included in a visual encoder of a preset visual reinforcement learning model to obtain second convolution kernel parameters and a second feature heat map of the second feature convolution kernel after processing the image; Constructing a first distillation loss based on a difference between the first convolution kernel parameters and the second convolution kernel parameters, and constructing a second distillation loss based on a difference between the first feature heat map and the second feature heat map; Predicting the state and reward at the next moment for the sample action data and sample state data corresponding to the second feature heat map using a preset self-supervised model, obtaining a prediction result for each sample image frame, and constructing a self-supervised loss based on multiple prediction results corresponding to multiple sample image frames; Performing action value and action sampling calculations based on the sample action data and the sample state data using a policy decoder of the preset visual reinforcement learning model to obtain calculation results for each sample image frame, and constructing a target policy loss based on multiple calculation results corresponding to multiple sample image frames; Based on the first distillation loss, the second distillation loss, the self-supervision loss, and the target strategy loss, the parameters of the preset visual reinforcement learning model are adjusted to obtain a target visual reinforcement learning model.
2. The model training method based on visual reinforcement learning according to claim 1, characterized in that: The self-supervised loss includes a first self-supervised loss and a second self-supervised loss. The state and reward at the next moment are predicted for the sample action data and sample state data corresponding to the second feature heat map by a preset self-supervised model to obtain a prediction result for each sample image frame, and a self-supervised loss is constructed based on multiple prediction results corresponding to multiple sample image frames, including: By using a preset self-supervisory model, for each sample image frame, the state and reward at the next moment are predicted for the sample action data and sample state data corresponding to the second feature heat map, thereby obtaining the predicted state data and predicted reward data for each sample image frame at the next moment; Obtain the target sample state data and target sample reward data at the next moment, and construct a first self-supervised loss based on the difference between the predicted state data and the target sample state data of multiple sample image frames, and construct a second self-supervised loss based on the difference between the predicted reward data and the target sample reward data of the multiple sample image frames.
3. The model training method based on visual reinforcement learning according to claim 1, characterized in that The target policy loss includes a value network sub-loss and a policy sub-loss. The policy decoder of the preset visual reinforcement learning model performs action value and action sampling calculations based on the sample action data and the sample state data to obtain a calculation result for each sample image frame, and constructs a target policy loss based on multiple calculation results corresponding to multiple sample image frames, including: The value network included in the policy decoder of the preset visual reinforcement learning model is used to calculate the action value for each sample image frame based on the sample action data and the sample state data to obtain the corresponding expected cumulative reward for the action; Construct a value network sub-loss based on the expected cumulative rewards of multiple actions corresponding to multiple sample image frames; For each sample image frame, obtaining an action probability distribution generated by the strategy decoder according to the sample state data, performing action sampling according to the action probability distribution to obtain sampled action data, and calculating an action probability according to the sampled action data; A strategy sub-loss is constructed according to a plurality of action probabilities corresponding to the plurality of sample image frames.
4. The model training method based on visual reinforcement learning according to claim 3, characterized in that: The constructing of a value network sub-loss according to the expected cumulative rewards of multiple actions corresponding to multiple sample image frames includes: For each sample image frame, obtaining an immediate reward of environmental feedback after executing the sample action data under the sample state data; Acquire the updated sample state data and the updated sample action data of each sample image frame at the next moment, and calculate the updated state value at the next moment based on the updated sample state data and the updated sample action data; Obtaining preset adjustment parameters, and adjusting the updated state value according to the adjustment parameters to obtain a target state value; Obtaining target value data according to the sum of the instant reward and the target state value; Obtaining a first difference based on a difference between the expected cumulative reward of the action and the target value data; A value network sub-loss is constructed based on the average of multiple first differences corresponding to multiple sample image frames.
5. The model training method based on visual reinforcement learning according to claim 4, characterized in that: The calculating of the update state value at the next moment based on the update sample state data and the update sample action data includes: For each sample image frame, obtaining an update action probability distribution generated by the strategy decoder for the update sample state data; Performing action sampling according to the next action probability distribution to obtain updated sampled action data, and determining an updated action probability of the updated sampled action data; Obtaining a preset entropy parameter, and obtaining a first product based on the product of the update action probability and the preset entropy parameter; Calculating the expected cumulative reward of the corresponding update action according to the update sample action data and the update sample state data through the value network included in the policy decoder; Obtaining a second difference based on a difference between the expected cumulative reward of the update action and the first product; An updated state value at the next moment is obtained based on an average of a plurality of second difference values corresponding to a plurality of sample image frames.
6. The model training method based on visual reinforcement learning according to claim 3, characterized in that: The constructing a strategy sub-loss according to the multiple action probabilities corresponding to the multiple sample image frames includes: Obtaining a preset entropy parameter, and obtaining a second product according to the product of the preset entropy parameter and the action probability for each sample image frame; Obtaining a preset target action expected cumulative reward, and obtaining a third difference based on a difference between the second product and the target action expected cumulative reward; A strategy sub-loss is constructed based on the average of multiple third differences corresponding to multiple sample image frames.
7. The model training method based on visual reinforcement learning according to claim 1, characterized in that: The obtaining of the sample image frame and the corresponding semantic category information includes: Acquire sample image frames, wherein the sample image frames include a first image frame at a current moment, a second image frame at a second moment, and a third image frame at a third moment, wherein the second moment is subsequent to the current moment, and the third moment is subsequent to the second moment; Acquire sample motion data corresponding to the first image frame and reference sample motion data corresponding to the second image frame; Obtain a preset semantic hint, and input the semantic hint, the first image frame, the second image frame, the third image frame, the sample action data, and the reference sample action data into a first visual large language model to obtain semantic category information corresponding to the sample image frame.
8. A model training device based on visual reinforcement learning, characterized in that: The device comprises: an acquisition module, configured to acquire a sample image frame and corresponding semantic category information, input the semantic category information into an initial first convolution kernel of a visual large language model to extract text features, obtain first convolution kernel parameters of the first feature convolution kernel after extracting text features, and input the sample image frame into the first feature convolution kernel to obtain a first feature heat map; a processing module, configured to process the sample image frame using an initial second feature convolution kernel included in a visual encoder of a preset visual reinforcement learning model to obtain second convolution kernel parameters and a second feature heat map of the second feature convolution kernel after processing the image; A construction module, configured to construct a first distillation loss based on a difference between the first convolution kernel parameters and the second convolution kernel parameters, and to construct a second distillation loss based on a difference between the first feature heat map and the second feature heat map; A prediction module, configured to predict the state and reward at the next moment for the sample action data and sample state data corresponding to the second feature heat map using a preset self-supervised model, obtain a prediction result for each sample image frame, and construct a self-supervised loss based on multiple prediction results corresponding to multiple sample image frames; a calculation module, configured to perform action value and action sampling calculations based on the sample action data and the sample state data using a policy decoder of the preset visual reinforcement learning model, obtain a calculation result for each sample image frame, and construct a target policy loss based on multiple calculation results corresponding to multiple sample image frames; An adjustment module is used to adjust the parameters of the preset visual reinforcement learning model based on the first distillation loss, the second distillation loss, the self-supervision loss and the target strategy loss to obtain a target visual reinforcement learning model.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the model training method based on visual reinforcement learning as described in any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the model training method based on visual reinforcement learning according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Training method, using method, device and equipment of multi-modal pre-training model
CN116756574A
Knowledge-driven scene priors for semantic audio-visual embodied navigation
US20250022296A1