Strategy control method based on self-supervised reinforcement learning
By introducing dynamic mask ratio and hierarchical reward prediction in reinforcement learning, the problems of low sample utilization and unstable strategy training under high-dimensional visual input are solved, and more efficient sample utilization and stronger long-term planning capabilities are achieved.
Patent Information
- Application Number
- CN202510444863.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-24
AI Technical Summary
Existing reinforcement learning methods face problems such as low sample utilization, instability in strategy training, and poor performance in complex tasks when dealing with high-dimensional visual inputs.
A strategy control method based on self-supervised reinforcement learning is proposed, which dynamically calculates the mask ratio, mask reconstruction enhancement state characterization, hierarchical modeling reward signals, and jointly optimizes short-term strategies and long-term planning.
It significantly improves the efficiency of reinforcement learning samples based on image input, improves the long-term planning and generalization capabilities of the agent, and is suitable for complex reinforcement learning tasks of high-dimensional visual input.
Smart Images

Figure CN120198749A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of reinforcement learning, and particularly relates to a policy control method based on self-supervised reinforcement learning. Background Art
[0002] In reinforcement learning tasks, agents based on pixel inputs face challenges such as high-dimensional state spaces, low sample utilization rates, and unstable policy training. Existing research mainly improves sample efficiency and enhances the generalization ability of agents through methods such as self-supervised learning, data augmentation, or world models.
[0003] 1. Self-supervised reinforcement learning method:
[0004] Contrastive Unsupervised Reinforcement Learning (CURL) optimizes visual representations through contrastive learning. It uses different augmented views of the same observation as positive samples and different observations as negative samples, maximizing the similarity between positive samples and minimizing the similarity between negative samples to improve the sample efficiency of reinforcement learning based on pixel inputs. CURL uses contrastive loss to cluster similar visual states, enabling the agent to learn more discriminative visual features, thus showing higher sample utilization rates in robot control and game agent training.
[0005] 2. Data augmentation reinforcement learning method:
[0006] Data-regularized Q-learning (DrQ) expands sample diversity through data augmentation such as image translation and cropping, and combines Q-function regularization to improve the robustness of the policy, which is applicable to real-time control tasks with high-dimensional visual inputs. DrQ randomly crops and translates the input images during the reinforcement learning training process to improve the agent's adaptability to environmental changes, thereby increasing the sample utilization rate and reducing training variance.
[0007] 3. World model method:
[0008] Masked World Models (MWM) improve the sample efficiency of visual control by reconstructing masked convolutional features and combining reward prediction tasks to decouple visual representations and dynamic modeling. MWM uses a self-supervised learning approach to model the environment, enabling the agent to learn the latent dynamic structure of the environment from limited samples. At the same time, it uses a masking mechanism to prevent the model from overfitting local features and improve the generalization ability.
[0009] CURL mainly focuses on visual feature extraction and does not explicitly model the environmental dynamics and reward signals, resulting in limited performance in complex tasks (such as long-term planning tasks). Moreover, it relies on manually designed data augmentation strategies, making it difficult to model temporal dynamic relationships and possibly optimizing task-irrelevant features, leading to low sample efficiency.
[0010] DrQ data augmentation improves robustness through image translation and regularization. However, the limitations of the augmentation strategy and noise interference restrict its application in complex tasks, especially in scenarios with low sample budgets where the performance significantly deteriorates.
[0011] MWM enhances the representation learning ability through a masking mechanism, focusing on visual feature modeling and environmental dynamics prediction. However, it does not deeply optimize the reward modeling for different time scales, limiting its performance in sparse or delayed reward tasks. In addition, its masking ratio is fixed, making it difficult to adapt to the learning needs at different stages of the training process and also difficult to adjust according to the differences in task complexity, such as simple control tasks and long-term planning tasks, thus affecting the generalization ability of the model. Summary of the Invention
[0012] In view of the deficiencies of the prior art, this application proposes a policy control method based on self-supervised reinforcement learning. By dynamically calculating the masking ratio, enhancing the state representation through mask reconstruction, hierarchically modeling the reward signal, and jointly optimizing the short-term policy and long-term planning, the sample efficiency of reinforcement learning based on image input is significantly improved.
[0013] In a first aspect, this application proposes a policy control method based on self-supervised reinforcement learning, including:
[0014] Obtain an image sequence to be observed;
[0015] Input the image sequence to be observed into a pre-trained policy control model for self-supervised reinforcement learning to obtain a control policy; in the training process of the pre-trained policy control model for self-supervised reinforcement learning, the training data is processed using a dynamic spatio-temporal mask, the masked training data is reconstructed, and the reconstructed training data is predicted through immediate reward prediction and long-term return prediction to obtain the pre-trained policy control model for self-supervised reinforcement learning.
[0016] The processing of the training data using a dynamic spatio-temporal mask includes:
[0017] Obtain training data, where the training data is a historical sequence of consecutive observed images;
[0018] Perform data augmentation operations on the historical sequence of consecutive observed images;
[0019] Calculate the current masking ratio according to the current training step and the total training steps;
[0020] Generate a spatio-temporal cubic occlusion region for the enhanced observed image sequence according to the current masking ratio, and randomly occlude the selected regions in the spatio-temporal cubic occlusion region to obtain the masked training data.
[0021] The calculation formula of the current masking ratio is as follows:
[0022]
[0023] where step is the current training step, Steps is the total number of training steps, η start is the initial masking ratio, and η end is the final masking ratio.
[0024] The generating a spatio-temporal cubic occlusion region for the enhanced observed image sequence according to the current masking ratio, and randomly occluding the selected regions in the spatio-temporal cubic occlusion region to obtain the masked training data includes:
[0025] Stack K enhanced observed image sequences into a cube with a shape of K×H×W, where H×W is the spatial size of the enhanced observed image, H is the height of the enhanced observed image, and W is the width of the enhanced observed image;
[0026] Divide the cube into regular and non-overlapping M sub-cubes, and the sub-cubes are evenly distributed in the cube;
[0027] According to the current masking ratio, randomly select sub-cubes as the spatio-temporal cubic occlusion region for masking according to a uniform distribution to obtain a masked observed image sequence, and the masked observed image sequence is the masked training data.
[0028] The reconstructing the masked training data includes:
[0029] Use an online encoder to encode the masked training data to obtain a latent state;
[0030] Use a momentum encoder to encode the historical consecutive observed image sequences to obtain a reconstruction target state;
[0031] Use a Transformer decoder to reconstruct the latent state through global self-attention to obtain a predicted latent state.
[0032] The predicting the reconstructed training data through immediate reward prediction and long-term return prediction includes:
[0033] Input the latent state into the immediate reward prediction head to obtain a first scalar, and use the first scalar as the immediate reward for the next step of training;
[0034] Input the latent state into the long-term return prediction head to obtain a second scalar, and use the second scalar as the return reward for N-step training.
[0035] The immediate reward prediction head is constructed using a multi-layer perceptron structure, including: a first fully connected layer and a second fully connected layer. The first fully connected layer uses the ReLU activation function to increase non-linearity and uses layer normalization to stabilize training. The second fully connected layer directly maps to a first scalar for output.
[0036] The long-term return prediction head is constructed using a multi-layer perceptron structure, including: a third fully connected layer, a fourth fully connected layer, and a fifth fully connected layer. The third fully connected layer uses the GELU activation function to increase non-linearity and uses layer normalization to stabilize training. The fourth fully connected layer uses the ReLU activation function to provide additional expressive power. The fifth fully connected layer directly maps to a second scalar for output.
[0037] The training process of the pre-trained self-supervised reinforcement learning policy control model further includes:
[0038] Calculate the reconstruction loss based on the cosine similarity between the predicted latent state and the reconstructed target state.
[0039] Calculate the immediate reward loss based on the rewards actually obtained in each step of training.
[0040] Calculate the N-step training reward loss based on the total rewards actually obtained in N-step training.
[0041] Calculate the combined loss based on the immediate reward loss, the N-step training reward loss, and the balance coefficient.
[0042] Perform a weighted sum of the reconstruction loss and the combined loss to obtain the total loss function.
[0043] Update the self-supervised reinforcement learning policy control model using the total loss function to obtain the pre-trained self-supervised reinforcement learning policy control model.
[0044] In a second aspect, the present application proposes an electronic device, including: one or more processors, and a memory. The memory is used to store instructions. When the instructions are executed by the one or more processors, the one or more processors execute the described policy control method based on self-supervised reinforcement learning.
[0045] In a third aspect, the present application proposes a computer-readable storage medium, which stores executable instructions. When the instructions are executed, the processor executes the described policy control method based on self-supervised reinforcement learning.
[0046] In a fourth aspect, the present application proposes a computer program product, including a computer program or instruction, which, when executed by a processor, implements the above-described method for policy control based on self-supervised reinforcement learning.
[0047] Beneficial effects:
[0048] The present application proposes a method for policy control based on self-supervised reinforcement learning, which combines dynamic spatio-temporal mask latent alignment (DSMLA) with hierarchical reward prediction (HRP). Among them, DSMLA uses a masked autoencoder mechanism to learn more stable state representations, effectively alleviating the problem of feature confusion caused by high-dimensional visual inputs; at the same time, the dynamic mask ratio strategy adaptively adjusts the masking intensity according to the training progress, adopting a lower ratio in the initial stage of training to quickly capture basic visual features, and gradually increasing the ratio in the later stage to strengthen the modeling of long-range spatio-temporal dependencies, so that the agent can learn effective policies faster. HRP, on the other hand, refines the reward modeling by jointly optimizing the immediate reward and the long-term return, significantly enhancing the long-term planning ability of the agent, enabling reinforcement learning to find effective policies in a shorter time. Compared with existing methods, the method of the present application improves the policy performance while reducing the computational overhead, and is applicable to complex reinforcement learning tasks with high-dimensional visual inputs, such as Atari games and DMControl tasks. Description of the drawings
[0049] Figure 1 Flowchart of a method for policy control based on self-supervised reinforcement learning according to an embodiment of the present application;
[0050] Figure 2 Schematic diagram of the training process of a policy control model for self-supervised reinforcement learning according to an embodiment of the present application;
[0051] Figure 3 Schematic diagram of the training process of a policy control model for self-supervised reinforcement learning according to an embodiment of the present application;
[0052] Figure 4 Schematic diagram of the spatio-temporal mask strategy according to an embodiment of the present application;
[0053] Figure 5 Schematic diagram of hierarchical reward prediction according to an embodiment of the present application. Detailed implementation manners
[0054] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present application.
[0055] Embodiment 1:
[0056] This embodiment proposes a method for policy control based on self-supervised reinforcement learning, as Figure 1 shown, including:
[0057] Step S1: Obtain the image sequence to be observed;
[0058] Step S2: Input the image sequence to be observed into the pre-trained policy control model of self-supervised reinforcement learning to obtain a control policy; for the pre-trained policy control model of self-supervised reinforcement learning, the training process includes: processing the training data with a dynamic spatio-temporal mask, reconstructing the masked training data, and predicting the reconstructed training data through immediate reward prediction and long-term return prediction to obtain the pre-trained policy control model of self-supervised reinforcement learning.
[0059] In this embodiment, inputting the image sequence to be observed into the pre-trained policy control model of self-supervised reinforcement learning obtains a series of control actions. The pre-trained policy control model of self-supervised reinforcement learning is the key to obtaining accurate control actions. In this embodiment, through dynamic spatio-temporal mask potential alignment, the environmental dynamics are implicitly modeled in the latent space, reducing redundant calculations and enhancing the expression ability of spatio-temporal continuity features, overcoming the over-reliance of traditional mask methods on pixels or convolutional features. At the same time, a mask ratio that adaptively adjusts with the training progress is adopted, providing more observation information in the early stage of training to accelerate representation learning, and enhancing the mask constraint in the later stage to improve the model's context reasoning ability and reduce the risk of overfitting. In this embodiment, by designing hierarchical reward prediction, jointly optimizing short-term (1-step) and long-term (N-step) reward signals, explicitly modeling multi-granularity temporal dependencies, enhancing long-term planning ability, and making up for the limitations caused by single-scale reward prediction in existing methods. And combining mask self-supervision and hierarchical reward prediction: reducing the dependence on artificial data augmentation, improving the sample efficiency and generalization ability of visual reinforcement learning in complex control tasks. The pre-trained policy control model of self-supervised reinforcement learning established by the method of this embodiment can adaptively adjust the mask policy at different training stages, effectively balancing the learning rate and stability, and improving the performance in reward-sparse or long-term planning tasks.
[0060] Regarding the pre-trained policy control model of self-supervised reinforcement learning, the specific training process is described in detail as follows, as Figure 2 、 Figure 3 shown:
[0061] (1) Mask module:
[0062] In step S2, the processing of the training data with a dynamic spatio-temporal mask includes:
[0063] Step S2.1: Obtain training data, where the training data is a historical continuous observation image sequence;
[0064] In this embodiment, obtain a historical continuous observation image sequence {o t ,ot+1 ,…,o t+K-1} as the input data for training.
[0065] Step S2.2: Perform data augmentation operations on the historical consecutive observation image sequence;
[0066] In this embodiment, first, apply conventional data augmentation operations to the input image, such as random cropping, color perturbation, etc., to improve the generalization ability of the model. Subsequently, according to the current masking ratio η t Generate a spatio-temporal cube masking region and randomly mask the selected region, thereby guiding the model to learn more comprehensive global context information and enhancing its modeling ability for spatio-temporal dynamic features.
[0067] Step S2.3: Calculate the current masking ratio according to the current training step and the total training steps;
[0068] The calculation formula for the current masking ratio is as follows:
[0069]
[0070] where step is the current training step, Steps is the total training steps, and η start is the initial masking ratio (0.2 in this embodiment), and η end is the final masking ratio (0.7 in this embodiment).
[0071] Step S2.4: Generate a spatio-temporal cube masking region for the enhanced observation image sequence according to the current masking ratio, and randomly mask the selected region in the spatio-temporal cube masking region to obtain the masked training data, including:
[0072] Step S2.4.1: Stack K enhanced observation image sequences into a cube with a shape of K×H×W, where H×W is the spatial size of the enhanced observation image, H is the height of the enhanced observation image, and W is the width of the enhanced observation image;
[0073] Step S2.4.2: Divide the cube into regular and non-overlapping M sub-cubes, and the sub-cubes are evenly distributed in the cube;
[0074] Step S2.4.3: Randomly select a part (where the value range of the current masking ratio η t is greater than 0 and less than 1) of the sub-cubes as the spatio-temporal cube masking region for masking according to the uniform distribution, and obtain a masked observation image sequence, and the masked observation image sequence is the masked training data
[0075] In this embodiment, the masker performs structured spatio-temporal masking on the input visual observation sequence according to the masking ratio η obtained from the dynamic masking ratio scheduler. t , and performs structured spatio-temporal masking on the input visual observation sequence. As Figure 4 shown, the specific process is as follows: First, take out the observation sequence {o t , o t+1 , …, o t+K-1} of K consecutive time steps from the experience replay pool. Each observation contains n RGB frames. Stack these observations into a cube with a shape of K×H×W (actually K×H×W×D, where D = 3n is the number of channels), and H×W is the spatial size of the image; then divide this large cube into regular and non-overlapping small cubes, each with a shape of k×h×w, and these small cubes are evenly distributed in the original large cube; subsequently, randomly select a part of the small cubes for masking according to the masking ratio η t according to a uniform distribution to obtain a masked observation sequence, and input it into the data augmenter.
[0076] (2) Reconstruction module:
[0077] In step S2, the reconstruction of the masked training data includes:
[0078] Step S2.5: Use an online encoder to encode the masked training data to obtain a latent state
[0079] Step S2.6: Use a momentum encoder to encode the historical continuous observation image sequence to obtain a reconstruction target state
[0080] Step S2.7: Use a Transformer decoder to reconstruct the latent state through global self-attention to obtain a predicted latent state
[0081] In this embodiment, an online encoder is used to encode the masked observation sequence into a latent state, and at the same time, a momentum encoder is used to encode the original observation sequence to form a reconstruction target. Then, based on the decoder of the Transformer architecture, the masked state is reconstructed through the global self-attention mechanism, and the predicted latent state is output to prompt the model to capture visual features and environmental dynamics.
[0082] (3) Reward prediction module:
[0083] In step S2, the prediction of the reconstructed training data through immediate reward prediction and long-term return prediction includes:
[0084] Step S2.8: Input the latent state into the immediate reward prediction head to obtain a first scalar, and use the first scalar as the immediate reward for the next step of training.
[0085] The immediate reward prediction head is constructed using a multi-layer perceptron (MLP, a network structure composed of multiple fully connected layers), which contains two fully connected layers. After the first fully connected layer, the ReLU activation function is applied to increase non-linearity, and layer normalization is applied to stabilize the training. The second fully connected layer directly maps to a scalar output, representing the predicted immediate reward value.
[0086] Step S2.9: Input the latent state into the long-term return prediction head to obtain a second scalar, and use the second scalar as the return reward for N-step training.
[0087] The long-term return prediction head is constructed using a multi-layer perceptron (MLP) structure, which contains three fully connected layers. After the third fully connected layer, the GELU activation function is applied to increase non-linearity, and layer normalization (LayerNormalization) is applied to stabilize the training. After the fourth fully connected layer, the ReLU activation function is applied to provide additional expressive power. The fifth fully connected layer directly maps to a scalar output, representing the predicted N-step discounted return value.
[0088] In this embodiment, based on the latent state output by the decoder, two independent prediction heads are added:
[0089] Immediate reward prediction head: Processes the latent state at the current time step and outputs a scalar as the prediction of the immediate reward for the next step.
[0090] Long-term return prediction head: Processes the latent state at the current time step and outputs a scalar as the prediction of the future N-step discounted return, where the N-step discounted return The calculation formula is:
[0091]
[0092] where γ is the discount factor, r t+i is the actual reward obtained in each future step, i is the i-th step of training, N is the total number of training steps, is the return reward for N-step training.
[0093] In this embodiment, according to the actually obtained immediate reward r t and the cumulative reward for the future N steps calculate the loss of reward prediction.
[0094] These two prediction heads learn immediate rewards and long-term returns respectively through supervision signals, thus achieving hierarchical modeling of the temporal structure of rewards.
[0095] (4) Total loss calculation and network update:
[0096] In step S2, for the pre-trained self-supervised reinforcement learning policy control model, the training process further includes, as Figure 5 shown:
[0097] Step S2.10: Calculate the reconstruction loss according to the cosine similarity between the predicted latent state and the reconstructed target state;
[0098] In this embodiment, the cosine similarity between the predicted latent state and the target latent state is used to calculate the DSMLA reconstruction loss L DSMLA
[0099]
[0100] where L DSMLA is the reconstruction loss, is the predicted latent state, is the target latent state, and K is the total length of the historical continuous observation image sequence;
[0101] Step S2.11: Calculate the immediate reward loss according to the rewards actually obtained in each step of training;
[0102] Step S2.12: Calculate the N-step training reward loss according to the total rewards actually obtained in N-step training;
[0103] In this embodiment, the mean squared error is used to calculate the HRP reward prediction loss L HRP :
[0104]
[0105] where L immediate is the immediate reward loss, L N-step is the N-step training reward loss, and MSE() is the abbreviation of Mean Squared Error, which is a commonly used loss function for measuring the difference between predicted values and true values. In this embodiment, MSE is used to calculate the error between the predicted reward and the actual reward. is the predicted N-step discounted return value of the model, that is, the scalar output by the long-term return prediction head. This is compared with the actual N-step discounted return to calculate the loss of long-term return prediction.
[0106] Step S2.13: Calculate the combined loss according to the immediate reward loss, the N-step training reward loss, and the balance coefficient;
[0107] Combined with the balance coefficient α, the HRP loss is combined as follows:
[0108] L HRP = α · (L immediate + L N-step )
[0109] where L HRP is the combined loss.
[0110] Step S2.14: Perform weighted summation on the reconstruction loss and the combined loss to obtain the total loss function;
[0111] In this embodiment, the main loss function of the joint reinforcement learning, and the final total loss is calculated as:
[0112] L total = L RL + λ1L DSMLA + λ2L HRP
[0113] where L total is the total loss function.
[0114] Step S2.15: Update the policy control model of the self-supervised reinforcement learning using the total loss function to obtain the pre-trained policy control model of the self-supervised reinforcement learning.
[0115] In this embodiment, perform backpropagation update on the online encoder, decoder, and prediction head according to the total loss to obtain the final policy control model of the self-supervised reinforcement learning.
[0116] Table 1: Comparison table of simulation results:
[0117]
[0118] As shown in Table 1, it presents the performance comparison between the method of this embodiment and the prior art in a typical Atari game environment. From the data, it can be seen that the policy control method (DSMLA+HRP) based on self-supervised reinforcement learning proposed in this embodiment has achieved optimal or near-optimal performance in most games. In the "Alien" and "Ms Pacman" games, the methods of this embodiment reached high scores of 860.1 and 1337.2 respectively, significantly exceeding other comparison algorithms. Especially in the "Pong" game, the method of this embodiment is the only algorithm that achieved a positive score (2.2), indicating its superiority in complex control tasks. In the "Seaquest" environment, the score of this method is 557.4, which is about 2.6% higher than the closest SPR method, and 83.8% and 79.1% higher than DrQ and CURL respectively. It is worth noting that in the "Breakout" game, although the SPR method has a slight advantage, the method of this embodiment still maintains a high level of performance, far superior to the random baseline, CURL, and DrQ methods. These results fully verify the effectiveness of the technical solution combining dynamic spatio-temporal masking and hierarchical reward prediction in improving the performance of reinforcement learning. Among them, Random is the random action policy, DrQ (Data-regularized Q-learning) is data-augmented reinforcement learning, CURL (Contrastive Unsupervised Reinforcement Learning) is contrastive self-supervised reinforcement learning, SPR (Self-Predictive Representations) is self-predictive representation, and Ours is the method of this embodiment (DSMLA+HRP).
[0119] A policy control method based on self-supervised reinforcement learning proposed in this embodiment adopts the DSMLA mechanism. By masking and reconstructing the latent features of the input data, the model can learn more robust visual representations under limited data conditions, thereby improving the sample utilization efficiency of reinforcement learning and the generalization ability of the agent in high-dimensional visual tasks. Different from the traditional fixed-ratio masking method, the masking strategy proposed in the present invention can adaptively adjust the masking intensity according to the training progress and task characteristics, solving the problem of mismatch between the masking intensity and the learning stage. Specifically, this mechanism helps the model quickly capture basic visual features in the early stage of training, while strengthening the reasoning ability for long-range spatio-temporal dependencies in the later stage, thus significantly improving the sample efficiency and policy stability under complex tasks.
[0120] Meanwhile, in this embodiment, through a hierarchical structure of immediate reward prediction (1-step) and long-term return prediction (N-step), refined reward modeling is performed, enabling the agent to optimize both short-term policies and long-term plans simultaneously, and significantly enhancing the decision-making ability in sparse reward tasks. In addition, by jointly optimizing DSMLA and HRP as auxiliary tasks with the reinforcement learning loss function, the agent can improve the reward prediction effect while learning the state representation, effectively enhancing the learning signal, thereby reducing the sample requirement and improving the training efficiency.
[0121] In summary, through the above technological innovations, this embodiment significantly improves the sample utilization efficiency and decision-making ability of the reinforcement learning algorithm based on image input in complex tasks, and has broad application prospects.
[0122] Embodiment 2:
[0123] This embodiment provides an electronic device, including: one or more processors, and a memory for storing instructions, which when executed by the one or more processors, cause the one or more processors to execute the above-mentioned policy control method based on self-supervised reinforcement learning.
[0124] The electronic device can be a mobile phone, a computer, a tablet computer, etc., including a memory and a processor, with a computer program stored on the memory, and when the computer program is executed by the processor, it implements a policy control method based on self-supervised reinforcement learning as described in the embodiment. It can be understood that the electronic device may further include an input / output (I / O) interface and a communication component.
[0125] Among them, the processor is used to execute all or part of the steps in the above-mentioned policy control method based on self-supervised reinforcement learning. The memory is used to store various types of data, which may include, for example, instructions of any application program or method in the electronic device, as well as data related to the application program.
[0126] The processor may be implemented by an application specific integrated circuit (ASIC), a digital signal processor (DSP), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic components, and is used to execute the above-mentioned policy control method based on self-supervised reinforcement learning.
[0127] Embodiment 3:
[0128] This embodiment provides a computer-readable storage medium storing executable instructions, which, when executed and implemented in the form of software functional units and sold or used as independent products, can be stored in a computer-readable storage medium.
[0129] This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of a policy control method based on self-supervised reinforcement learning described in various embodiments of this application.
[0130] The foregoing storage medium includes: flash memory, hard disk, multimedia card, card-type memory (such as SD (Secure Digital Memory Card, secure digital memory card) or DX (abbreviation for Memory Data Register, MDR, memory data register) memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, server, APP (abbreviation for Application, application software) application store, and other various media that can store program check codes. A computer program is stored thereon, and when the computer program is executed by a processor, it can implement each step of the foregoing policy control method based on self-supervised reinforcement learning.
[0131] Embodiment 4:
[0132] This embodiment provides a computer program product including a computer program or instructions, which, when executed by a processor, implement the foregoing policy control method based on self-supervised reinforcement learning.
[0133] Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a computer program product.
[0134] The various embodiments in this application are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0135] The protection scope of this application is not limited to the above embodiments. Obviously, those skilled in the art can make various modifications and deformations to the present disclosure without departing from the scope and spirit of the present disclosure. If these modifications and deformations fall within the scope of the claims of the present disclosure and their equivalent technologies, the intention of the present disclosure also includes these modifications and deformations.
Claims
1. A strategy control method based on self-supervised reinforcement learning, characterized in that: include: Obtain the image sequence to be observed; Inputting the image sequence to be observed into a pre-trained self-supervised reinforcement learning strategy control model to obtain a control strategy; The pre-trained self-supervised reinforcement learning policy control model has a training process comprising: processing the training data using a dynamic spatiotemporal mask, reconstructing the masked training data, predicting the reconstructed training data through immediate reward prediction and long-term reward prediction, and obtaining the pre-trained self-supervised reinforcement learning policy control model.
2. A strategy control method based on self-supervised reinforcement learning according to claim 1, characterized in that: The processing of the training data using a dynamic spatiotemporal mask includes: Acquire training data, wherein the training data is a historical continuous observation image sequence; Performing a data enhancement operation on the historical continuous observation image sequence; Calculate the current mask ratio based on the current training steps and total training steps; According to the current mask ratio, a spatiotemporal cubic masking region is generated for the enhanced observed image sequence, and a selected region in the spatiotemporal cubic masking region is randomly masked to obtain masked training data.
3. A strategy control method based on self-supervised reinforcement learning according to claim 2, characterized in that: The current mask ratio is calculated as follows: Among them, step is the current training step number, Steps is the total training step number, η start is the initial mask ratio, η end is the final mask ratio.
4. A strategy control method based on self-supervised reinforcement learning according to claim 2, characterized in that: The method generates a spatiotemporal cubic masking region for the enhanced observed image sequence according to the current mask ratio, and randomly masks a selected region in the spatiotemporal cubic masking region to obtain masked training data, including: Stack the K enhanced observation image sequences into a cube with a shape of K×H×W, where H×W is the spatial size of the enhanced observation image, H is the height of the enhanced observation image, and W is the width of the enhanced observation image; Splitting the cube into M regular and non-overlapping sub-cubes, wherein the sub-cubes are evenly distributed in the cube; According to the current mask ratio, a sub-cube is randomly selected as the space-time cube masking area according to uniform distribution for masking, so as to obtain a masked observation image sequence, and the masked observation image sequence is the masked training data.
5. The strategy control method based on self-supervised reinforcement learning according to claim 1, characterized in that: The step of reconstructing the masked training data includes: Use an online encoder to encode the masked training data to obtain the latent state; The momentum encoder is used to encode the historical continuous observation image sequence to obtain the reconstructed target state; The Transformer decoder is used to reconstruct the latent state through global self-attention to obtain the predicted latent state.
6. A strategy control method based on self-supervised reinforcement learning according to claim 1, characterized in that: The reconstructed training data is predicted through immediate reward prediction and long-term reward prediction, including: Input the potential state into the immediate reward prediction head to obtain a first scalar, and use the first scalar as the immediate reward for the next step of training; the immediate reward prediction head is constructed using a multi-layer perceptron structure, including: a first fully connected layer and a second fully connected layer, the first fully connected layer uses a ReLU activation function to increase nonlinearity, and uses application layer normalization to stabilize training, and the second fully connected layer is directly mapped to the first scalar for output; The latent state is input into the long-term reward prediction head to obtain a second scalar, and the second scalar is used as the reward for N-step training. The long-term reward prediction head is constructed with a multi-layer perceptron structure, including: a third fully connected layer, a fourth fully connected layer and a fifth fully connected layer. The third fully connected layer uses a GELU activation function to increase nonlinearity and uses application layer normalization to stabilize training. The fourth fully connected layer uses a ReLU activation function to provide additional expression capabilities. The fifth fully connected layer is directly mapped to the second scalar for output.
7. The strategy control method based on self-supervised reinforcement learning according to claim 1, characterized in that: The pre-trained self-supervised reinforcement learning strategy control model, the training process also includes: Compute the reconstruction loss based on the cosine similarity between the predicted latent state and the reconstructed target state; Calculate the immediate reward loss based on the actual rewards obtained at each step of training; Calculate the reward loss for N-step training based on the total reward actually obtained from N-step training; Calculate the combined loss based on the immediate reward loss, N-step training reward loss, and the balance coefficient; Performing weighted summation on the reconstruction loss and the combined loss to obtain a total loss function; The total loss function is used to update the policy control model of self-supervised reinforcement learning to obtain a pre-trained policy control model of self-supervised reinforcement learning.
8. An electronic device, characterized in that: include: One or more processors, and a memory, wherein the memory is used to store instructions, and when the instructions are executed by the one or more processors, the one or more processors execute a policy control method based on self-supervised reinforcement learning as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that: It stores executable instructions, which, when executed, enable the processor to execute a strategy control method based on self-supervised reinforcement learning as described in any one of claims 1 to 7.
10. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, a strategy control method based on self-supervised reinforcement learning as described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Reinforced learning method based on image time-frequency domain enhancement and dynamic mask generation network
CN120525747A
Reinforcement learning method based on image time-frequency domain enhancement and dynamic mask generation network
CN120525747B
Deep learning-based electroencephalogram slow wave real-time feedback transcranial electrical stimulation method and system
CN121731667A