Training Method, Device, Electronic Device, Storage Medium and Computer Program Product for Single-Agent Reinforcement Learning Model Based on Visual Representation

Through a single agent reinforcement learning model based on visual representation, using Transformer architecture and action and reward prediction learning, the problem of low sample efficiency of visual reinforcement learning in complex continuous control tasks is solved, and performance improvement and stable training is achieved.

CN119580029BActive Publication Date: 2025-07-22INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411601987.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-11
Publication Date
2025-07-22
Estimated Expiration
2044-11-11

AI Technical Summary

Technical Problem

Existing visual reinforcement learning has low sample efficiency and poor control performance in complex continuous control tasks, especially when processing high-dimensional visual images, it is difficult to quickly learn effective strategies.

Method used

A single agent reinforcement learning model based on visual representation is adopted, including an online state encoder, action encoder, reinforcement learning network and auxiliary task network. The state and action representation are learned through the state prediction model, and the time series information is processed using the Transformer architecture, combining action and reward prediction learning to optimize the model training process.

Benefits of technology

The performance and sample efficiency of single agents in complex continuous control tasks are improved, model crash problems are avoided, and the stability and convergence speed of model training are improved through asymmetric projection networks and layer normalization components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119580029B_ABST
    Figure CN119580029B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a training method, device, electronic device, storage medium, and computer program product for a single-agent reinforcement learning model based on visual representation. The single-agent reinforcement learning model includes an online state encoder, an action encoder, a reinforcement learning network, and an auxiliary task network. The auxiliary task network includes a state prediction model. According to the state information and action information of the target agent in the current time period based on the observation image of the target agent, and the reward information in the current time period, starting from the perspective of visual representation through the auxiliary task network, the state representation and action representation of the target agent are learned. The reinforcement learning network selects the best decision-making action for the target agent. Moreover, by making full use of the temporal information of time periods in reinforcement learning, the performance and sample efficiency of a single agent in complex continuous control tasks with images as state inputs can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of reinforcement learning, and in particular, to a training method, device, electronic device, storage medium, and computer program product for a single-agent reinforcement learning model based on visual representation. Background Art

[0002] With the significant improvement in computing power and the continuous emergence of algorithm innovations, reinforcement learning (RL) has made remarkable progress in many fields, including games, robotics, autonomous driving, etc. Despite the brilliant achievements, RL still faces a core challenge in practical applications, that is, how to improve sample efficiency.

[0003] The agent needs to have the ability to quickly learn effective strategies from limited interactions, especially when dealing with high-dimensional visual images and performing complex continuous control tasks. This challenge becomes particularly prominent. The high dimensionality, redundancy, and diversity of visual signals bring a series of challenges to the application of vision-based reinforcement learning. For example, vision-based reinforcement learning generally faces problems such as low sample efficiency and poor control performance. Summary of the Invention

[0004] The training method, device, electronic device, storage medium, and computer program product for a single-agent reinforcement learning model based on visual representation provided by the exemplary embodiments of the present disclosure can at least solve the above technical problems and other technical problems not mentioned above.

[0005] According to one aspect of the present disclosure, there is provided a training method for a single-agent reinforcement learning model based on visual representation. The single-agent reinforcement learning model based on visual representation includes an online state encoder, an action encoder, a reinforcement learning network, and an auxiliary task network. The auxiliary task network includes a state prediction model. The training method for the single-agent reinforcement learning model based on visual representation includes: obtaining state information, action information, and reward information of a target agent in a current time period, where the current time period consists of a preset plurality of consecutive moments including the current moment, and the state information and the action information are obtained based on an observation image of the target agent; inputting the state information into the online state encoder to obtain a state feature; inputting the action information into the action encoder to obtain an action feature; inputting the state feature, the action feature, and the reward information into the state prediction model to obtain a state prediction feature of the target agent in a next time period, where the next time period consists of a preset plurality of consecutive moments including the next moment; calculating a state prediction loss based on a difference between the state prediction feature and a corresponding true value; inputting the state feature and the action feature into the reinforcement learning network to calculate a reinforcement learning loss; and training the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss and the state prediction loss.

[0006] Optionally, the inputting the state feature, the action feature, and the reward information into the state prediction model to obtain a state prediction feature of the target agent in a next time period includes: inputting the state feature, the action feature, and the reward information into the state prediction model, and averaging an output of the state prediction model to obtain an initial state prediction feature of the target agent in the next time period; and inputting the initial state prediction feature into an online projection network to obtain the state prediction feature.

[0007] Optionally, the online projection network includes an online projection head and an online prediction head; wherein, inputting the initial state prediction feature into the online projection network to obtain the state prediction feature includes: inputting the initial state prediction feature into the online projection head to obtain first projection data; inputting the first projection data into the online prediction head to obtain the state prediction feature; wherein, calculating the state prediction loss based on the difference between the state prediction feature and the corresponding ground truth includes: obtaining the state information of the target agent in the next time period; inputting the state information in the next time period into the target state encoder to obtain a second state prediction feature, wherein the parameters of the target state encoder are obtained by exponential moving average based on the current parameters of the online state encoder and a preset decay rate; inputting the second state prediction feature into the target projection network to obtain the ground truth corresponding to the state prediction feature, wherein the target projection network includes a target projection head, and the parameters of the target projection head are obtained by exponential moving average based on the current parameters of the online projection head; calculating the state prediction loss based on the difference between the state prediction feature and the corresponding ground truth.

[0008] Optionally, the auxiliary task network further includes an action prediction model; wherein, the training method of the single-agent reinforcement learning model based on visual representation further includes: obtaining the state information of the target agent at the current moment and the state information at the next moment; inputting the state information at the current moment and the state information at the next moment into the online state encoder respectively to obtain the state feature at the current moment and the state feature at the next moment; inputting the state feature at the current moment and the state feature at the next moment into the action prediction model to obtain the action prediction feature at the current moment; calculating the action prediction loss based on the difference between the action prediction feature and the corresponding ground truth; wherein, training the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss and the state prediction loss includes: training the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss, the state prediction loss and the action prediction loss.

[0009] Optionally, the auxiliary task network further includes a reward prediction model; wherein, the training method of the single-agent reinforcement learning model based on visual representation further includes: obtaining the state information of the target agent at the current moment and the action information of the current time period; inputting the action information of the current time period into an action encoder to obtain the action feature of the current time period; inputting the state feature of the current moment and the action feature of the current time period into the reward prediction model to obtain the reward prediction feature at the last moment of the current time period; calculating a reward prediction loss based on the difference between the reward prediction feature and the corresponding true value; wherein, training the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss, the state prediction loss, and the action prediction loss includes: training the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss, the state prediction loss, the action prediction loss, and the reward prediction loss.

[0010] Optionally, inputting the state feature, the action feature, and the reward information into the state prediction model includes: inputting the state feature, the action feature, the reward information, and spatial position information into the state prediction model, wherein the state feature, the action feature, and the reward information at the same moment share the same spatial position information.

[0011] Optionally, the state prediction model is a Transformer model composed of a plurality of preset identical blocks, each block includes a multi-head self-attention layer and a feed-forward network layer composed of a multi-layer perceptron, a layer normalization component is included before each layer, and a residual connection is included after each layer.

[0012] Optionally, the method further includes: after obtaining the state information of the target agent in the current time period, performing a random masking operation on a preset proportion of the data in the state information to obtain updated state information.

[0013] According to another aspect of the present disclosure, there is also provided a training device for a single-agent reinforcement learning model based on visual representation. The single-agent reinforcement learning model based on visual representation includes an online state encoder, an action encoder, a reinforcement learning network, and an auxiliary task network. The auxiliary task network includes a state prediction model. The training device for the single-agent reinforcement learning model based on visual representation includes:

[0014] An information acquisition unit, configured to: acquire the state information, action information, and reward information of the target agent in the current time period, where the current time period consists of a preset number of consecutive moments including the current moment, and the state information and the action information are obtained based on the observation image of the target agent; a state representation unit, configured to: input the state information into the online state encoder to obtain state features; an action representation unit, configured to: input the action information into the action encoder to obtain action features; a state prediction unit, configured to: input the state features, the action features, and the reward information into the state prediction model to obtain the state prediction features of the target agent in the next time period, where the next time period consists of a preset number of consecutive moments including the next moment; a state loss calculation unit, configured to: calculate a state prediction loss based on the difference between the state prediction features and the corresponding true values; a reinforcement learning loss calculation unit, configured to: input the state features and the action features into the reinforcement learning network to calculate a reinforcement learning loss; a model training unit, configured to: train the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss and the state prediction loss.

[0015] According to another aspect of the embodiments of the present disclosure, there is also provided an electronic device, including: at least one processor; at least one memory storing computer-executable instructions, where when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method of the single-agent reinforcement learning model based on visual representation as described in any one of the above.

[0016] According to another aspect of the embodiments of the present disclosure, there is also provided a computer-readable storage medium storing instructions, which when run by at least one processor, cause the at least one processor to execute the training method of the single-agent reinforcement learning model based on visual representation as described in any one of the above.

[0017] According to another aspect of the embodiments of the present disclosure, there is also provided a system including at least one computing device and at least one storage device storing instructions, where when the instructions are run by the at least one computing device, the at least one computing device is caused to execute the training method of the single-agent reinforcement learning model based on visual representation as described in any one of the above.

[0018] According to another aspect of the embodiments of the present disclosure, there is also provided a computer program product, including computer programs / instructions, which when executed by a processor, implement the training method of the single-agent reinforcement learning model based on visual representation as described in any one of the above.

[0019] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0020] According to the training method, device, electronic device, storage medium, and computer program product of the single-agent reinforcement learning model based on visual representation of the present disclosure, for a target agent, it is possible to learn the state representation and action representation of the target agent from the perspective of visual representation through an auxiliary task network, select the best decision-making action for the target agent through a reinforcement learning network, and fully utilize the temporal information of time periods in reinforcement learning, so as to improve the performance and sample efficiency of a single agent in a challenging complex continuous control task with images as state inputs.

[0021] In addition, by adopting an asymmetric projection network architecture, the problem of model collapse in the self-supervised learning process can be avoided.

[0022] In addition, introducing action prediction learning as an additional learning constraint can enhance the contribution of state representation in future action prediction.

[0023] In addition, adding reward prediction learning additionally to constrain state and action representations can promote the agent to better understand the possible consequences of its actions.

[0024] In addition, the transformer architecture enables the simultaneous processing of information of a time series of the target agent to fully utilize the temporal information of time periods in reinforcement learning; a layer normalization component is equipped in front of each layer, which can improve the stability in the model training process and accelerate the convergence speed; a residual connection is added to the output of each layer, which can promote the effective training of deeper networks and help prevent the problem of gradient disappearance or explosion, ensuring the smooth flow of information in the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.

[0026] Figure 1 Showing a schematic diagram of the composition of a single-agent reinforcement learning model based on visual representation in an exemplary embodiment of the present disclosure;

[0027] Figure 2 Showing a flowchart of a training method of a single-agent reinforcement learning model based on visual representation in an exemplary embodiment of the present disclosure;

[0028] Figure 3 Showing a schematic diagram of a data processing flow related to a state prediction model in an exemplary embodiment of the present disclosure;

[0029] Figure 4Schematic diagram of data processing flow related to an action prediction model in an exemplary embodiment of the present disclosure;

[0030] Figure 5 Schematic diagram of data processing flow related to a reward prediction model in an exemplary embodiment of the present disclosure;

[0031] Figure 6 Schematic diagram of the training process by a single-agent reinforcement learning model based on visual representation in an exemplary embodiment of the present disclosure;

[0032] Figure 7 Normalized return curve diagram during the training of single-agent reinforcement learning based on visual representation in an exemplary embodiment of the present disclosure;

[0033] Figure 8 Block diagram of a training device for a single-agent reinforcement learning model based on visual representation in an exemplary embodiment of the present disclosure;

[0034] Figure 9 Block diagram of an electronic device in an exemplary embodiment of the present disclosure. Detailed implementation manners

[0035] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0037] It should be noted here that "at least one of several items" in the present disclosure all represents three parallel situations including "any one of the several items", "any combination of multiple items of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example, "performing at least one of step one and step two" means the following three parallel situations: (1) performing step one; (2) performing step two; (3) performing step one and step two.

[0038] In the field of visual reinforcement learning, although research based on state representation has made some progress in improving sample efficiency, most of these methods focus on environments with discrete action spaces such as Atari games or simple tasks in DMControl, often neglecting applications in challenging complex continuous control tasks. For example, in humanoid control tasks, it is difficult to achieve an ideal performance level relying solely on state representation. In addition, action representation also has important value and significance for the policy training of reinforcement learning. An effective action representation can promote reinforcement learning to better understand the dynamics of the environment. However, existing research has not explored action representation sufficiently and has not fully utilized the temporal information in reinforcement learning, resulting in waste of resources.

[0039] To solve the above problems, the present disclosure provides a training method, apparatus, electronic device, storage medium, and computer program product for a single-agent reinforcement learning model based on visual representation. For a target agent, it can learn the state representation and action representation of the target agent from the perspective of visual representation through an auxiliary task network, select the best decision-making action for the target agent through a reinforcement learning network, and fully utilize the temporal information in reinforcement learning, thereby achieving performance and sample efficiency improvement of a single agent in challenging complex continuous control tasks with images as state inputs.

[0040] Next, the training method, apparatus, electronic device, storage medium, and computer program product of the single-agent reinforcement learning model based on visual representation of the present disclosure will be specifically described with reference to Figures 1 to 8 Specifically describe the training method, apparatus, electronic device, storage medium, and computer program product of the single-agent reinforcement learning model based on visual representation of the present disclosure.

[0041] First, the composition of the single-agent reinforcement learning model based on visual representation will be introduced.

[0042] Figure 1 The schematic diagram of the composition of the single-agent reinforcement learning model based on visual representation in an exemplary embodiment of the present disclosure is shown.

[0043] Refer to Figure 1 , the single-agent reinforcement learning model based on visual representation includes but is not limited to an online state encoder, an action encoder, a reinforcement learning network, and an auxiliary task network, and the auxiliary task network includes but is not limited to a state prediction model.

[0044] According to an exemplary embodiment of the present disclosure, the online state encoder f for state feature extraction θ can be constructed by a convolutional neural network, where θ represents the parameters of the online state encoder f θ of.

[0045] The (online) action encoder g for action feature extraction α can be constructed by a convolutional neural network, where α represents the action encoder g αparameters.

[0046] According to an exemplary embodiment of the present disclosure, the state prediction model φ can be a Transformer model composed of a preset plurality of identical blocks. For example, a Transformer model composed of a preset plurality (e.g., L, and the value of L can be determined by specific reinforcement learning tasks and hardware resources) of identical blocks is adopted. Each block can internally contain two main stacked layers: a multi-head self-attention layer and a feed-forward network layer composed of a multi-layer perceptron. To improve the stability during model training and accelerate the convergence speed, a layer normalization component can be equipped in front of each layer. In addition, to facilitate the effective training of deeper networks, a residual connection can be added to the output of each layer, which helps prevent the problem of gradient vanishing or explosion and ensures the smooth flow of information in the network.

[0047] The Transformer architecture adopted by the state prediction model in the exemplary embodiment of the present disclosure can process the information of the empirical trajectory of a time series (i.e., the current time period composed of a preset plurality of consecutive moments including the current moment, e.g., K moments, and the value of K can be determined by specific reinforcement learning tasks and hardware resources) by sequential data input, and explicitly prompt the state prediction model to capture the context information around the state.

[0048] The online state encoder can process the state information to obtain state features, and the action encoder can process the action information to obtain action features; the state prediction model included in the auxiliary task network of the auxiliary task network can process the state features, action features, and reward information to predict the next state, and by minimizing the difference between the predicted features and the true features, prompt the online state encoder and the action encoder to learn effective state representations and action representations respectively; the reinforcement learning network can process the state features and action features to obtain the next action selection of the agent and the evaluation of the quality of the action selection strategy.

[0049] Figure 2 A flowchart showing the training method of the single-agent reinforcement learning model based on visual representation in the exemplary embodiment of the present disclosure.

[0050] Refer to Figure 2 , in step 201, obtain the state information, action information, and reward information of the target agent in the current time period, where the current time period is composed of a preset plurality of consecutive moments including the current moment, and the state information and action information are obtained based on the observed images of the target agent.

[0051] According to an exemplary embodiment of the present disclosure, from the perspective of visual representation, the performance and sample efficiency of visual reinforcement learning can be improved by learning effective state representations and action representations.

[0052] The current moment can be represented by t, and the state information at the current moment can be expressed as s t , and the action information at the current moment can be represented by a l , and the reward information at the current moment can be represented by r t .

[0053] The state information and action information of the target agent can be extracted from the observed image of the target agent.

[0054] The current time period can be composed of K consecutive moments including the current moment. Specifically, the current time period can be the time period experienced by the experience trajectory of length K including the current moment, which can be represented by t:t+K-1.

[0055] In each training iteration, a batch of experience trajectories of length K {s t , a t , r t , …, s t+K-1 , a t+K-1 , r t+L-1} can be obtained as the state information, action information, and reward information of the target agent in the current time period, that is, the samples, also known as the sequence data for state prediction.

[0056] According to the exemplary embodiments of the present disclosure, for a batch of randomly sampled samples, first, data augmentation techniques such as Random Shift can be used for preprocessing. Then, this batch of data can be used to simultaneously perform contrastive representation learning based on state-action-reward and learning of reinforcement learning policies. In addition, prediction learning of actions and rewards can also be performed. Specifically, in each training iteration, for a batch of experience trajectories of length K {s t , a t , r t , …, s t+K-1 , a t+K-1 , r t+L-1}, data augmentation techniques can be applied to all states in the batch: s t:t+K-1 ←Aug(s t:t+K-1 ), to obtain a batch of augmented samples of length K {s t , s t+1 , …, s t+K-2}.

[0057] Next, the process of randomly masking the state information is introduced.

[0058] To achieve more effective state prediction learning, for the sequence data {s t , at , r t , …, s t+K-1 , a t+K-1 , r t+K-1} for data preprocessing.

[0059] According to an exemplary embodiment of the present disclosure, after obtaining the state information of the target agent in the current time period, a random masking operation can be performed on a preset proportion of the data in the state information to obtain the updated state information.

[0060] For a batch of augmented sample sequences {s t , s t+1 , …, s t+K-2} of length K, 50% of the states in the sequence can be randomly selected for masking.

[0061] Specifically, a masking sequence M = {M t , M t+1 , …, M t+K-2} corresponding to {s t , s t+1 , …, s t+K-2} can be first defined. For each M t+i ∈ [M], there is a 50% probability that M l+i = 0 and a 50% probability that M t+i = 1.

[0062] If M t+i = 1, the corresponding state in {s t , s t+1 , …, s t+K-2} remains unchanged;

[0063] If M t+i = 0, the corresponding state in {s t , s t+1 , …, s t+K-2} performs the following masking operation:

[0064]

[0065] In step 202, the state information is input into the online state encoder to obtain the state features.

[0066] The online encoder can map the augmented sample S t to its corresponding state features:

[0067] According to an exemplary embodiment of the present disclosure, for convenience of representation, the states in the sequence that have undergone random masking can be represented as S mask , and its corresponding state features are represented as Zmask For a batch of augmented sample sequences {s t , s t+1 , …, s t+K-2} of length K, after the random masking operation, the batch of samples can be represented as {s t , s mask , …, s t+K-2}, and the state feature sequence obtained by passing it through the online state encoder can be represented as: {z t , z mask , …, z t+K-2}.

[0068] In step 203, the action information is input into the action encoder to obtain the action feature.

[0069] The action encoder can map the action a t to its corresponding action feature: where α represents the parameter of the online action encoder g α .

[0070] These state features and action features can then be used in the policy network π and the state prediction model φ.

[0071] According to an exemplary embodiment of the present disclosure, based on the state feature sequence {z t , z mask , …, z t+K-2}, the action feature sequence {u t , u t+1 , …, u t+K-2}, and the reward sequence {r t , …, r t+K-2} after the random masking operation, they can be respectively used as state tokens, action tokens, and reward tokens (i.e., state information, action information, and reward information) and input into the Transformer-based state prediction model.

[0072] In step 204, the state feature, action feature, and reward information are input into the state prediction model to obtain the state prediction feature of the target agent in the next time period, where the next time period consists of a preset number of consecutive moments including the next moment. The goal of the state prediction model is to learn the state feature of the next moment based on the state, action, and reward information at each moment.

[0073] According to an exemplary embodiment of the present disclosure, the next time period can consist of K consecutive moments including the next moment. Specifically, the next time period can be the time period experienced by an empirical trajectory of length K including the next moment, which can be represented as t+1:t+K.

[0074] In step 205, a state prediction loss is calculated based on the difference between the state prediction feature and the corresponding true value.

[0075] Next, first refer to Figure 3 the data processing process related to the state prediction model is introduced.

[0076] Figure 3 The schematic diagram of the data processing flow related to the state prediction model in the exemplary embodiment of the present disclosure is shown.

[0077] Refer to Figure 3 , according to the exemplary embodiment of the present disclosure, the state prediction model is used to construct a state prediction task. The structural design of the state prediction model φ based on Transformer allows it to receive three types of input tokens: one type is state tokens, the feature sequence {z t , …, z t+k-2} extracted from the state sequence, and these features are generated by the online state encoder; another type is action tokens, the feature sequence {u t , u t+1 , …, u t+K-2} extracted from the action sequence, and they are generated by the online action encoder; the last type is reward tokens {r t , …, r t+K-2}.

[0078] In order to explicitly simulate strong local relationships, the input tokens are grouped according to "state - action - reward", and each element in the group has a strong causal relationship with other elements. Specifically, the state feature Z t , action feature u t and reward r t at each time step are divided into a group. Based on the image state, action, and reward information feedback from the environment of the target agent at the current moment, state prediction sequence data {s t , a t , r t , …, s t+K-1 , a t+K-1 , r t+K-1} can be generated.

[0079] According to the exemplary embodiment of the present disclosure, the state feature, action feature, reward information, and spatial position information can be input into the state prediction model, where the state feature, action feature, and reward information at the same moment share the same spatial position information.

[0080] The relative position embedding {p t , p t+1 , …, p t+k-2}, spatial position information is added to all input tokens to better capture the temporal relationship in the sequence. Refer to Figure 2 , at the same time step (moment), the state tokens, action tokens, and reward tokens can share the same position embedding to ensure that the model can correctly understand the temporal relationship between the state and the action.

[0081] The output of the state prediction model φ can be expressed as:

[0082] According to an exemplary embodiment of the present disclosure, the state feature, action feature, and reward information are input into the state prediction model, and the average value of the output of the state prediction model is calculated to obtain the initial state prediction feature of the target agent in the next time period; the initial state prediction feature is input into the online projection network to obtain the state prediction feature.

[0083] The state feature z at each time step t , the action feature u t and the reward r t are divided into a group. At the output end of the state prediction model, the outputs corresponding to this group are averaged to generate a feature prediction of the state at the next moment. In this way, the state prediction model can learn the initial state prediction feature of the predicted next time period, that is, the state representation, based on a large number of "state-action-reward" combinations Through the feature extraction and dimensionality reduction of the online projection network, the final state prediction feature can be obtained.

[0084] To avoid the collapse problem of the model during the self-supervised learning process, an asymmetric projection network architecture can be adopted. For example, the online projection network can include, but is not limited to, an online projection head and an online prediction head.

[0085] According to an exemplary embodiment of the present disclosure, the initial state prediction feature can be input into the online projection head to obtain the first projection data; the first projection data is input into the online prediction head to obtain the state prediction feature.

[0086] Specifically, for the initial state prediction feature of the next time period output by the state prediction model, that is, the prediction feature of the target state the online projection network can adopt an online projection head and an online prediction head to process to obtain the final state prediction feature:

[0087] In this way, the model can learn the predicted state representation based on a large number of "state-action-reward" combinations After feature extraction and dimensionality reduction by the online projection network, the final predicted features are obtained. Among them, m1 and m2 are the parameters of the online projection head and the online prediction head respectively.

[0088] Next, the calculation process of the state prediction loss is introduced below.

[0089] According to an exemplary embodiment of the present disclosure, an additional target state encoder can be introduced to calculate the features of the target state, that is, the true value of the state prediction features, rather than directly using the online encoder. The target state encoder parameters are not updated by gradient descent, but are updated by the exponential moving average technique according to the parameters of the online encoder and a decay rate τ ∈ [0, 1). The update rule can be expressed as:

[0090] According to an exemplary embodiment of the present disclosure, the state information of the target agent in the next time period can be obtained; and the state information in the next time period is input into the target state encoder to obtain the second state prediction features, where the parameters of the target state encoder are obtained by exponential moving average based on the current parameters of the online state encoder and a preset decay rate; the second state prediction features are input into the target projection network to obtain the corresponding true value of the state prediction features, where the target projection network includes a target projection head, and the parameters of the target projection head are obtained by exponential moving average based on the current parameters of the online projection head; based on the difference between the state prediction features and the corresponding true value, the state prediction loss is calculated.

[0091] For the target state s t+1:t+K-1 , only the target projection head can be used to obtain the final features, that is, the true value of the state prediction features: Among them, are the parameters of the target projection head. During the training process, the parameters of the target projection head do not participate in gradient update, but are inherited from the parameters m1 of the online projection head through the exponential moving average strategy, that is, gradient truncation.

[0092] To achieve representation learning, state prediction learning can be optimized using contrastive learning. Specifically, the mutual information between the state prediction features, that is, the predicted state features and the corresponding true value, that is, the true target state features can be maximized, so as to achieve the simultaneous learning of state-action representations.

[0093] The loss of contrastive learning can be defined as:

[0094]

[0095] wherein represents the state feature of the anchor sample, represents the state feature of the positive sample,

[0096] represents the state feature of any sample in the sequence, including positive and negative samples. W is a learnable parameter that provides a computational space for calculating the similarity metric between q and k.

[0097] The objective of contrastive learning is to make the state feature q of the anchor sample i and the state feature k of the positive sample + similar, while being different from the state feature k of other negative samples in the sequence i \{k +}.

[0098] Returning to reference Figure 2 , in step 206, the state feature and the action feature are input into the reinforcement learning network to calculate the reinforcement learning loss.

[0099] According to an exemplary embodiment of the present disclosure, DrQ-v2 can be selected as the benchmark method for reinforcement learning. DrQ-v2 is a model-agnostic reinforcement learning algorithm that can directly learn from image states based on the off-policy Actor-Critic method and data augmentation. The Actor (policy network) is responsible for selecting appropriate actions according to the current state, while the Critic (value network) is responsible for evaluating the value function of the state and the action.

[0100] The reinforcement learning loss, that is can be divided into two parts. Among them, the loss of the Critic network based on the n-step return can be expressed as:

[0101]

[0102] wherein, a t = π(z t ) + ∈, represents the state feature at time t, represents the Q value estimated by the k-th Q network, represents the execution of the expectation operation on (s , a t , r t , s t ) sampled from the experience pool t+1 ;

[0103]

[0104] wherein, a t+n = π(z t+n ) + ∈, Represents the state characteristics at time t, The Q-value estimated by the k-th target Q-network, where γ is the discount factor.

[0105] The loss of the Actor network can be expressed as:

[0106]

[0107] where a t = π(z t ) + ∈, σ is the exploration noise, c is the clipped value, represents the expectation operation on S sampled from the experience pool t .

[0108] In step 207, the single-agent reinforcement learning model based on visual representation is trained based on the reinforcement learning loss and the state prediction loss.

[0109] According to an exemplary embodiment of the present disclosure, the auxiliary task network and the reinforcement learning network can be trained synchronously and share the online state encoder and the action encoder. The parameters of the online state encoder, the action encoder, and the state prediction model can be optimized using gradient descent.

[0110] According to an exemplary embodiment of the present disclosure, in order to enhance the contribution of the state representation to future action prediction, action prediction learning can be introduced as an additional learning constraint. Next, the operations related to action prediction learning are introduced.

[0111] According to an exemplary embodiment of the present disclosure, the auxiliary task network may further include, but is not limited to, an action prediction model. The action prediction model is used to construct an action prediction task.

[0112] Specifically, an action prediction model composed of a fully connected network can be introduced where p1 are the parameters of the action prediction model.

[0113] The action prediction model may take as input the current state s t and the next state s t+1 , and may output the predicted features of the current action

[0114] ​As an example, the state information of the target agent at the current moment and the state information at the next moment can be obtained; the state information at the current moment and the state information at the next moment are respectively input into the online state encoder to obtain the state feature at the current moment and the state feature at the next moment; the state feature at the current moment and the state feature at the next moment are input into the action prediction model to obtain the action prediction feature at the current moment; based on the difference between the action prediction feature and the corresponding true value, the action prediction loss is calculated.

[0115] Figure 4 FIG. shows a schematic diagram of a data processing flow related to an action prediction model in an exemplary embodiment of the present disclosure.

[0116] According to an exemplary embodiment of the present disclosure, based on the image state and action information of the target agent at the current moment, an action prediction sequence data {s t , a t , …, s t+K-1 , a t+K-1} can be generated. The state information is used to generate predicted action features, and the action information is used to generate true action features.

[0117] Based on the image state s t at each moment in the action prediction sequence data, the state feature can be extracted through the online state encoder Referring to Figure 4 , for any two adjacent moments of states s t+i and s t+i+1 , the state features Z t+i and Z t+i+1 obtained by passing them through the online state encoder can be input into the action prediction model to obtain the action prediction feature at the t + i moment, that is, the prediction of the action feature

[0118] Based on the action a t at each moment in the action prediction sequence data, the corresponding true value of the action prediction feature, that is, the true action feature u Figure 3 , can be obtained through the action encoder in t+i .

[0119] Based on the predicted action feature and the true action feature u t+i , the difference between the two can be minimized, and the parameters of the online state encoder and the action prediction model are optimized using gradient descent according to the action prediction loss function .

[0120] According to an exemplary embodiment of the present disclosure, a single-agent reinforcement learning model based on visual representation can also be trained based on reinforcement learning loss, state prediction loss, and action prediction loss, that is, the parameters of the online state encoder, action encoder, state prediction model, and action prediction model are adjusted.

[0121] In addition, in order to facilitate the agent to better understand the possible consequences of its actions, a reward prediction learning can be additionally added to constrain the state and action representations. Therefore, according to an exemplary embodiment of the present disclosure, the auxiliary task network may also include, but is not limited to, a reward prediction model. The reward prediction model is used to construct a reward prediction task.

[0122] Specifically, a reward prediction model composed of a fully connected network can be introduced where p2 are the parameters of the reward prediction model.

[0123] Reward prediction model The input of the t can be the state S at the current moment t:t+n-1 and an action sequence a of length n

[0124] and the output can be the predicted reward at the (t + n - 1)-th future moment

[0125] Figure 5 shows a schematic diagram of the data processing flow related to the reward prediction model in an exemplary embodiment of the present disclosure.

[0126] According to an exemplary embodiment of the present disclosure, based on the image state information, action information, and reward information of the target agent at the current moment, a reward prediction sequence data {s t , a t , r t , …, s t+K-1 , a t+K-1 , r t+K-1} can be generated.

[0127] Based on the image state at each moment in the reward prediction sequence data, state features can be extracted through the online state encoder Based on the action at each moment in the reward prediction sequence data, action features can be extracted through the action encoder Referring to Figure 5, for the state s at any given time t+i and an action sequence a of length n t+i:t+i+n-1 , the state feature z obtained by passing them through the online state encoder t+i and the action feature u obtained by passing through the action encoder t+i:t+i+n-1 can be input into the reward prediction model to obtain the reward prediction feature for time t + i + n - 1, i.e., the predicted reward

[0128] Based on the predicted reward and the corresponding true reward r t+i+n-1 , the difference between the two can be minimized, and according to the reward prediction loss function the parameters of the online state encoder, action encoder, and reward prediction model are optimized using gradient descent for training.

[0129] The action prediction model and the reward prediction model can guide the learning of state representation and action representation by constructing additional constraint tasks and minimizing the gap between the predicted feature and the true target feature.

[0130] According to an exemplary embodiment of the present disclosure, the single-agent reinforcement learning model based on visual representation can also be trained based on the reinforcement learning loss, state prediction loss, action prediction loss, and reward prediction loss, i.e., the parameters of the online state encoder, action encoder, state prediction model, action prediction model, and reward prediction model are adjusted.

[0131] The auxiliary task network and the reinforcement learning network can be trained synchronously, and based on the state prediction loss action prediction loss reward prediction loss and reinforcement learning loss the total loss is calculated, and the total loss can be expressed as: where λ1, λ2, and λ3 respectively control the contributions of the three prediction losses to the total loss.

[0132] According to an exemplary embodiment of the present disclosure, the specific implementation process of the single-agent reinforcement learning method based on visual representation may include but is not limited to the following steps:

[0133] Obtain the data of the target agent at the current moment, where the data of the target agent at the current moment includes the image state, the executed action, and the reward information feedback by the environment of the target agent at the current moment;

[0134] Based on the data of the target agent corresponding to the current moment, learn the state representation and action representation through the auxiliary task network, and select the best decision-making action through the reinforcement learning network.

[0135] Figure 6 A schematic flowchart of training through a single-agent reinforcement learning model based on visual representations in an exemplary embodiment of the present disclosure is shown.

[0136] Referring to Figure 6 , according to an exemplary embodiment of the present disclosure, the auxiliary task network may include, but is not limited to, a state prediction model, an action prediction model, and a reward prediction model to perform state prediction learning, action prediction learning, and reward prediction learning. The three prediction models can minimize the difference between the predicted features and the true features to prompt the online state encoder and the action encoder to learn effective state representations and action representations respectively; the deep reinforcement learning (DRL) network may include, but is not limited to, a policy network and a value network. The input of the policy network can be the state representation output by the online state encoder, and the output can be the action selection of the agent. The input of the value network can be the state representation output by the online state encoder and the action representation output by the action encoder, and the output can be the Q value (state-action value), which is used to evaluate the quality of the current policy. The auxiliary task network and the reinforcement learning network can be trained synchronously and share the online state encoder and the action encoder.

[0137] Only as an example, for a complete training process of a single-agent reinforcement learning based on visual representations, it may specifically include the following steps S1-S6:

[0138] Step S1, initialize the parameters of the online state encoder f θ , the target state encoder , the action encoder g, the state prediction model φ, the online projection head G m , the target projection head , the online prediction head H, the action prediction model h, the reward prediction model l, and the policy network π, and initialize the memory space of the experience pool.

[0139] Step S2, the agent interacts with the environment, and stores the generated trajectory in the experience pool cache.

[0140] Step S3, randomly sample a batch of samples from the experience pool, and perform data augmentation and random masking on the states in the batch of samples.

[0141] Step S4, perform state prediction tasks, action prediction tasks, and reward prediction tasks, and update the network parameters of each prediction model according to their respective optimization objectives; perform reinforcement learning policy training, and update the network parameters of the reinforcement learning network according to the optimization objective.

[0142] Step S5, determine whether the training termination condition is reached, for example, whether the preset number of training times is reached, otherwise repeat steps S2-S4.

[0143] Step S6: When the training ends, obtain the parameters of the agent network (single-agent reinforcement learning model based on visual representation) updated for the last time as the training result.

[0144] In this embodiment, the unit of the preset number of training steps is time steps. The training can stop after sampling the preset number of time steps. After the training ends, the parameters of the agent network can be retained to use the policy network of reinforcement learning to select actions for the agent in subsequent executions.

[0145] Use the method implemented in this embodiment to control the agent to conduct experimental verification on 6 challenging complex control tasks of DMControl (Deepmind Control). The normalized return curve during the training process is as Figure 7 shown.

[0146] Figure 7 Show the normalized return curve graph during the training process of the single-agent reinforcement learning based on visual representation in the exemplary embodiment of the present disclosure.

[0147] Refer to Figure 7 , (a)Cheetah Run, (b)Acrobot Swingup, (c)Quadruped Walk, (d)Quadruped Run, (e)Reacher Hard, (f)Hopper Hop are the environment names of DeepMind Control, and each environment corresponds to a different single-agent reinforcement learning task. In the single-agent reinforcement learning based on visual representation (Transformer-based State-Action-Reward Prediction Representation Learning, TSAR) of this embodiment, the learning time for each task is 2M or 3M time steps. Compared with the baseline algorithm (DrQ-v2), the method of this embodiment can not only show excellent convergence performance, but also has high sample efficiency.

[0148] According to the exemplary embodiment of the present disclosure, a single-agent reinforcement learning method based on visual representation is provided, which can solve the technical problems of low sample efficiency and poor control performance in visual reinforcement learning caused by the high dimensionality, redundancy, and diversity of visual signals in the related art.

[0149] Figure 8 Show the block diagram of the training device of the single-agent reinforcement learning model based on visual representation in the exemplary embodiment of the present disclosure.

[0150] The single-agent reinforcement learning model based on visual representation includes an online state encoder, an action encoder, a reinforcement learning network, and an auxiliary task network, and the auxiliary task network includes a state prediction model.

[0151] Referring Figure 8 , an exemplary embodiment of the present disclosure further provides a training device 800 for a single-agent reinforcement learning model based on visual representation, which may include, but is not limited to, an information acquisition unit 801, a state representation unit 802, an action representation unit 803, a state prediction unit 804, a state loss calculation unit 805, a reinforcement learning loss calculation unit 806, and a model training unit 807.

[0152] The information acquisition unit 801 can acquire the state information, action information, and reward information of the target agent in the current time period, where the current time period consists of a preset number of consecutive moments including the current moment, and the state information and action information are obtained based on the observation images of the target agent.

[0153] The state representation unit 802 can input the state information into the online state encoder to obtain state features.

[0154] The action representation unit 803 can input the action information into the action encoder to obtain action features.

[0155] The state prediction unit 804 can input the state features, action features, and reward information into the state prediction model to obtain the state prediction features of the target agent in the next time period, where the next time period consists of a preset number of consecutive moments including the next moment.

[0156] The state loss calculation unit 805 can calculate the state prediction loss based on the difference between the state prediction features and the corresponding true values.

[0157] The reinforcement learning loss calculation unit 806 can input the state features and action features into the reinforcement learning network to calculate the reinforcement learning loss.

[0158] The model training unit 807 can train the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss and the state prediction loss.

[0159] It can be understood that in the exemplary embodiments of the training apparatus 800 of the above-mentioned single-agent reinforcement learning model based on visual representation, the specific implementation process is substantially the same as that of the exemplary embodiments of the above-mentioned training method of the single-agent reinforcement learning model based on visual representation, and will not be elaborated here. The training apparatus 800 of the single-agent reinforcement learning model based on visual representation can be respectively configured as software, hardware, firmware, or any combination of the above items for performing specific functions. For example, these apparatuses can correspond to dedicated integrated circuits, or can correspond to pure software codes, or can also correspond to modules combining software and hardware. In addition, one or more functions implemented by these apparatuses can also be uniformly executed by components in a physical entity device (such as a processor, a client, or a server, etc.).

[0160] Figure 9 The block diagram of an electronic device showing an exemplary embodiment of the present disclosure is shown.

[0161] Referring to Figure 9 , the electronic device 900 includes at least one memory 901 and at least one processor 902. A set of computer-executable instructions is stored in the at least one memory 901. When the set of computer-executable instructions is executed by the at least one processor 902, the training method of the single-agent reinforcement learning model based on visual representation according to the exemplary embodiment of the present disclosure is executed.

[0162] As an example, the electronic device 900 can be a PC computer, a tablet device, a personal digital assistant, a smart phone, or other devices capable of executing the above instruction set. Here, the electronic device 900 does not have to be a single electronic device, and can also be any aggregate of devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device 900 can also be a part of an integrated control system or a system manager, or can be configured as a portable electronic device that is interconnected locally or remotely (for example, via wireless transmission).

[0163] In the electronic device 900, the processor 902 can include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. As an example but not a limitation, the processor can also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0164] The processor 902 can run the instructions or codes stored in the memory 901, where the memory 901 can also store data. The instructions and data can also be sent and received via a network interface device through a network, where the network interface device can adopt any known transmission protocol.

[0165] The memory 901 can be integrated with the processor 902. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor, etc. In addition, the memory 901 can include independent devices such as external disk drives, storage arrays, or other storage devices that can be used by any database system. The memory 901 and the processor 902 can be operatively coupled or can communicate with each other, for example, through I / O ports, network connections, etc., such that the processor 902 can read files stored in the memory.

[0166] In addition, the electronic device 900 can also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device 900 can be connected to each other via a bus and / or a network.

[0167] According to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions can also be provided, wherein when the instructions are run by at least one computing device, the at least one computing device is caused to execute the above-mentioned method for training a single-agent reinforcement learning model based on visual representations.

[0168] Examples of the computer-readable storage medium herein include: read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memory, hard disk drive (HDD), solid state drive (SSD), cartridge memory (such as multimedia card, secure digital (SD) card or extreme digital (XD) card), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid state disk and any other device configured to store a computer program and any associated data, data files and data structures in a non-transitory manner and provide the computer program and any associated data, data files and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer program in the above computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer program and any associated data, data files and data structures are distributed on a networked computer system such that the computer program and any associated data, data files and data structures are stored, accessed and executed in a distributed manner by one or more processors or computers. It should be noted that the instructions can also be used to perform additional steps other than the above steps or perform more specific processing when performing the above steps. The content of these additional steps and further processing has been mentioned during the description of the related methods, so it will not be repeated here to avoid redundancy.

[0169] Another embodiment of the present disclosure relates to a system including at least one computing device and at least one storage device storing instructions, wherein when the instructions are run by the at least one computing device, the at least one computing device is caused to execute the above-mentioned method for training a single-agent reinforcement learning model based on visual representation.

[0170] It should be noted that the system according to the exemplary embodiment of the present disclosure can fully rely on the running of a computer program or instructions to implement corresponding functions, that is, each unit corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a dedicated software package (such as a lib library) to implement corresponding functions.

[0171] On the other hand, when the above system is implemented in software, firmware, middleware, or microcode, the program code or code segment for performing the corresponding operations can be stored in a computer-readable medium such as a storage medium, so that at least one processor or at least one computing device can execute the corresponding operations by reading and running the corresponding program code or code segment.

[0172] According to an exemplary embodiment of the present disclosure, the storage device can be integrated with the computing device. For example, RAM or flash memory can be arranged within an integrated circuit microprocessor or the like. In addition, the storage device can include a separate device, such as an external disk drive, a storage array, or other storage devices that can be used by any database system. The storage device and the computing device can be operatively coupled or can communicate with each other, for example, through an I / O port, a network connection, etc., so that the computing device can read the instructions stored in the storage device.

[0173] Another embodiment of the present disclosure relates to a computer program product, including a computer program / instructions, which when executed by a processor, implement the training method of the single-agent reinforcement learning model based on visual representation described in any one of the above.

[0174] According to the training method, device, electronic device, storage medium, and computer program product of the single-agent reinforcement learning model based on visual representation provided by the present disclosure, for a target agent, it is possible to learn the state representation and action representation of the target agent from the perspective of visual representation through an auxiliary task network, select the best decision-making action for the target agent through a reinforcement learning network, and fully utilize the temporal information in reinforcement learning, so as to improve the performance and sample efficiency of the single-agent in a challenging complex continuous control task with images as state inputs.

[0175] In addition, by adopting an asymmetric projection network architecture, the problem of model collapse during the self-supervised learning process can be avoided.

[0176] In addition, introducing action prediction learning as an additional learning constraint can enhance the contribution of the state representation to future action prediction.

[0177] In addition, adding reward prediction learning additionally to constrain the state and action representations can promote the agent to better understand the possible consequences of its actions.

[0178] In addition, the transformer architecture enables the simultaneous processing of the information of a time series of target agents to make full use of the temporal information in reinforcement learning; a layer normalization component is equipped in front of each layer, which can improve the stability during the model training process and accelerate the convergence speed; a residual connection is added to the output of each layer, which can facilitate the effective training of deeper networks and help prevent the problems of gradient disappearance or explosion, ensuring the smooth flow of information in the network.

[0179] The above describes the exemplary embodiments of the present disclosure. It should be understood that the above description is only exemplary and not exhaustive, and the present disclosure is not limited to the disclosed exemplary embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the scope of the claims.

Claims

1. A training method for a single-agent reinforcement learning model based on visual representation, characterized in that The single-agent reinforcement learning model based on visual representation includes an online state encoder, an action encoder, a reinforcement learning network, and an auxiliary task network. The auxiliary task network includes a state prediction model. The training method of the single-agent reinforcement learning model based on visual representation includes: Obtain the state information, action information, and reward information of the target agent in the current time period. Herein, the current time period consists of a preset number of consecutive moments including the current moment, and the state information and the action information are obtained based on the observed images of the target agent; Input the state information into the online state encoder to obtain state features; Input the action information into the action encoder to obtain action features; Input the state features, the action features, and the reward information into the state prediction model to obtain the state prediction features of the target agent in the next time period. Herein, the next time period consists of a preset number of consecutive moments including the next moment; Calculate the state prediction loss based on the difference between the state prediction features and the corresponding true values; Input the state features and the action features into the reinforcement learning network to calculate the reinforcement learning loss; Train the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss and the state prediction loss; Wherein, the auxiliary task network further includes an action prediction model; Wherein, the training method of the single-agent reinforcement learning model based on visual representation further includes: Obtain the state information of the target agent at the current moment and the state information at the next moment; Input the state information at the current moment and the state information at the next moment into the online state encoder respectively to obtain the state features at the current moment and the state features at the next moment; Input the state features at the current moment and the state features at the next moment into the action prediction model to obtain the action prediction features at the current moment; Calculate the action prediction loss based on the difference between the action prediction features and the corresponding true values; Wherein, the training of the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss and the state prediction loss includes: Train the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss, the state prediction loss, and the action prediction loss.

2. The training method of the single-agent reinforcement learning model based on visual representation according to claim 1, wherein, The inputting the state features, the action features, and the reward information into the state prediction model to obtain the state prediction features of the target agent in the next time period includes: Input the state features, the action features, and the reward information into the state prediction model, and take the average of the output of the state prediction model to obtain the initial state prediction features of the target agent in the next time period; Input the initial state prediction features into the online projection network to obtain the state prediction features.

3. The training method of the single-agent reinforcement learning model based on visual representation according to claim 2, wherein, The online projection network includes an online projection head and an online prediction head; Wherein, the inputting the initial state prediction features into the online projection network to obtain the state prediction features includes: Input the initial state prediction feature into the online projection head to obtain first projection data; Input the first projection data into the online prediction head to obtain the state prediction feature; Among them, calculating the state prediction loss based on the difference between the state prediction feature and the corresponding true value includes: Obtain the state information of the target agent in the next time period; Input the state information of the next time period into the target state encoder to obtain a second state prediction feature, where the parameters of the target state encoder are obtained by exponential moving average based on the current parameters of the online state encoder and a preset decay rate; Input the second state prediction feature into the target projection network to obtain the corresponding true value of the state prediction feature, where the target projection network includes a target projection head, and the parameters of the target projection head are obtained by exponential moving average based on the current parameters of the online projection head; Calculate the state prediction loss based on the difference between the state prediction feature and the corresponding true value.

4. The training method of the single-agent reinforcement learning model based on visual representation according to claim 1, wherein The auxiliary task network further includes a reward prediction model; Among them, the training method of the single-agent reinforcement learning model based on visual representation further includes: Obtain the state information of the target agent at the current moment and the action information in the current time period; Input the action information in the current time period into the action encoder to obtain the action feature in the current time period; Input the state feature at the current moment and the action feature in the current time period into the reward prediction model to obtain the reward prediction feature at the last moment in the current time period; Calculate the reward prediction loss based on the difference between the reward prediction feature and the corresponding true value; Among them, training the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss, the state prediction loss, and the action prediction loss includes: Train the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss, the state prediction loss, the action prediction loss, and the reward prediction loss.

5. The training method of the single-agent reinforcement learning model based on visual representation according to claim 1, characterized in that, The inputting the state feature, the action feature, and the reward information into the state prediction model includes: Input the state feature, the action feature, the reward information, and the spatial position information into the state prediction model, where the state feature, the action feature, and the reward information at the same moment share the same spatial position information.

6. The training method of the single-agent reinforcement learning model based on visual representation according to claim 1, characterized in that, The state prediction model is a Transformer model composed of a preset number of identical blocks, each block includes a multi-head self-attention layer and a feed-forward network layer composed of a multi-layer perceptron, and each layer includes a layer normalization component before and a residual connection after.

7. The training method of the single-agent reinforcement learning model based on visual representation according to claim 1, wherein The method further includes: After obtaining the state information of the target agent in the current time period, perform a random masking operation on a preset proportion of the data in the state information to obtain updated state information.

8. A training device for a single-agent reinforcement learning model based on visual representation, characterized in that, The single-agent reinforcement learning model based on visual representation includes an online state encoder, an action encoder, a reinforcement learning network, and an auxiliary task network. The auxiliary task network includes a state prediction model. The training device for the single-agent reinforcement learning model based on visual representation includes: An information acquisition unit, configured to: acquire the state information, action information, and reward information of the target agent in the current time period, where the current time period consists of a preset number of consecutive moments including the current moment, and the state information and the action information are obtained based on the observed images of the target agent; A state representation unit, configured to: input the state information into the online state encoder to obtain state features; An action representation unit, configured to: input the action information into the action encoder to obtain action features; A state prediction unit, configured to: input the state features, the action features, and the reward information into the state prediction model to obtain the state prediction features of the target agent in the next time period, where the next time period consists of a preset number of consecutive moments including the next moment; A state loss calculation unit, configured to: calculate a state prediction loss based on the difference between the state prediction features and the corresponding true values; A reinforcement learning loss calculation unit, configured to: input the state features and the action features into the reinforcement learning network to calculate a reinforcement learning loss; A model training unit, configured to: train the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss and the state prediction loss; wherein the auxiliary task network further includes an action prediction model; The training device further includes: A state information acquisition unit, configured to: acquire the state information of the target agent at the current moment and the state information at the next moment; A state feature acquisition unit, configured to: input the state information at the current moment and the state information at the next moment into the online state encoder respectively to obtain the state features at the current moment and the state features at the next moment; An action prediction unit, configured to: input the state features at the current moment and the state features at the next moment into the action prediction model to obtain the action prediction features at the current moment; An action prediction loss calculation unit, configured to: calculate an action prediction loss based on the difference between the action prediction features and the corresponding true values; The model training unit is specifically configured to: Train the single-agent reinforcement learning model based on visual representation based on the reinforcement learning loss, the state prediction loss, and the action prediction loss.

9. An electronic device, characterized in that, Includes: At least one processor; At least one memory storing computer-executable instructions, wherein when the computer-executable instructions are run by the at least one processor, the at least one processor is caused to execute the training method of the single-agent reinforcement learning model based on visual representation according to any one of claims 1-7.

10. A computer-readable storage medium for storing instructions, characterized in that, When the instructions are run by at least one processor, causing the at least one processor to execute the training method of the vision-representation-based single-agent reinforcement learning model according to any one of claims 1-7.

11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, implementing the training method of the vision-representation-based single-agent reinforcement learning model according to any one of claims 1-7.

Citation Information

Patent Citations

  • Reinforcement learning model training method and device

    CN117669650A

  • Reinforcement learning sequence decision-making method, system, equipment and medium

    CN117972588A