Path coverage method based on retrieval enhancement mechanism
By introducing external memory bank and similar trajectory retrieval mechanism, Decision Transformer and pre-trained language model are used to solve the problems of low sample utilization and difficult model convergence in coverage path planning, and efficient path planning and decision-making capabilities are achieved.
Patent Information
- Application Number
- CN202510730405.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-03
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-03
AI Technical Summary
The existing coverage path planning method based on deep reinforcement learning has problems such as low sample utilization, sparse environmental rewards and long-round tasks, which lead to difficult model convergence, especially in sparse reward environments.
Using a path coverage method based on the search enhancement mechanism, by building external memory and similar trajectory retrieval mechanisms, using the Decision Transformer architecture and pre-trained language model, we quickly build path planning strategies, reduce redundant paths, and improve data utilization and model decision-making capabilities in sparse reward environments.
It effectively reduces computing resource consumption, improves the model's decision-making ability and coverage path planning performance in sparse reward environments, improves coverage ability in complex environments, and realizes earlier model convergence and efficient path planning.
Smart Images

Figure CN120258279A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep reinforcement learning and coverage path planning, and particularly relates to a path coverage method based on a retrieval enhancement mechanism. Background Art
[0002] Coverage path planning technology has a wide range of application scenarios in many production and life fields. In the service industry: autonomous cleaning, weed removal, snow shoveling, etc. In agriculture: pesticide spraying, automatic land reclamation, automatic farming, etc. In environmental exploration tasks: map surveying, interstellar exploration, oil and other mineral surveys. In similar application fields, the goal of the coverage path planning algorithm is to efficiently and comprehensively plan a path so that the mobile robot can effectively and without dead angles cover the specified area to perform tasks. An efficient coverage path planning algorithm can improve the operation efficiency and reduce labor costs and resource waste. With the development and innovation of artificial intelligence and robot technology, the coverage path planning algorithm is also constantly optimized to adapt to more complex environments and diverse application requirements. However, the traditional reinforcement learning-based methods still have some limitations: First, the sample utilization rate is low. Existing agents based on Deep Reinforcement Learning (DRL) often require a large amount of interaction data to learn a suitable strategy, and the data sampling cost in reinforcement learning is high, especially in practical applications, it is difficult to collect training data on a large scale.
[0003] Second, the environmental rewards are sparse. In the coverage path planning task, the reward signal is usually sparse, that is, the agent can only obtain effective feedback after completing a large-scale coverage, resulting in low learning efficiency and easy to fall into local optima.
[0004] Finally, long-round tasks lead to difficult model convergence. Many reinforcement learning tasks require the agent to perform thousands of steps of interaction to complete the task, and most traditional reinforcement learning methods (such as DQN, PPO, SAC) are difficult to effectively handle ultra-long time-step tasks. In addition, although Transformer, as a powerful sequence modeling tool, can process long sequence data, its computational cost is high, and its application in reinforcement learning tasks is limited to a certain extent. Summary of the Invention
[0005] In view of the above situation, the main purpose of the present invention is to propose a path coverage method based on a retrieval enhancement mechanism to solve the above technical problems.
[0006] The present invention proposes a path coverage method based on a retrieval enhancement mechanism, and the method includes the following steps: Step 1: Obtain the pre-collected dataset, construct an external memory bank based on the pre-collected dataset, construct a policy model based on the Decision Transformer architecture, regard the robot as an agent, and use the policy model to predict the action distribution of the agent to obtain the action distribution predicted by the policy model at the current moment; Step 2: Based on the pre-collected dataset, sequentially process the input trajectory through the FH mechanism and the pre-trained language model to obtain a set of key-value pairs, and save the set of key-value pairs to the external memory bank; Step 3: Obtain the current input sub-trajectory, retrieve similar sub-trajectories from the set of key-value pairs based on the current input sub-trajectory to obtain a set of retrieved similar sub-trajectories; Step 4: Based on the set of retrieved similar sub-trajectories, re-weight the similar sub-trajectories to obtain a sorted set of re-weighted sub-trajectories; Step 5: Based on the current input sub-trajectory and the sorted set of re-weighted sub-trajectories, adjust the action distribution predicted by the policy model at the current moment to obtain the finally predicted action distribution of the policy model, and use the finally predicted action distribution of the policy model to optimize the policy model to obtain an optimized policy model; Use the optimized policy model to confirm the recommendation rate of the agent to execute each action.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By introducing an external memory bank and adding a similar trajectory retrieval mechanism, the present invention quickly constructs a path planning strategy in the initial stage of training the model, effectively reduces the effect of redundant coverage path planning, significantly reduces the computing resources consumed by training, and improves the decision-making ability of the model in a sparse reward environment; 2. By introducing an external memory bank and adding a similar trajectory retrieval mechanism, the present invention learns from the retrieved similar paths, realizes the pertinence of training samples, further removes redundant paths, enables the coverage path planning model to converge earlier under limited computing resources, improves the performance of coverage path planning, and enhances the coverage ability of the model in a complex environment; 3. By designing and improving the external memory mechanism, the present invention can store the trajectory data that has existed in the historical trajectory, significantly reduces the model's requirement for the amount of training data, queries the data highly relevant to the current task in the stored historical data when performing the coverage path planning task, improves the utilization rate of data, effectively alleviates the problem of the model's dependence on the context length, and enables the robot to complete the coverage path planning task more efficiently.
[0008] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. Brief Description of the Drawings
[0009] Figure 1 It is a flowchart of the steps of a path coverage method based on a retrieval enhancement mechanism proposed by the present invention; Figure 2 It is a schematic diagram of the method framework of a path coverage method based on a retrieval enhancement mechanism proposed by the present invention Figure 1 ; Figure 3 It is a schematic diagram of the method framework of a path coverage method based on a retrieval enhancement mechanism proposed by the present invention Figure 2 . Detailed Embodiment
[0010] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0011] These and other aspects of the embodiments of the present invention will be clear with reference to the following description and drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited by this.
[0012] Please refer to Figure 1 , this embodiment provides a path coverage method based on a retrieval enhancement mechanism, and the method includes the following steps: Step 1, obtain a pre-collected data set, construct an external memory bank based on the pre-collected data set, construct a policy model based on the Decision Transformer architecture, regard the robot as an agent, and use the policy model to predict the action distribution of the agent to obtain the action distribution predicted by the policy model at the current moment.
[0013] Please refer to Figure 2 , in Step 1, when obtaining a pre-collected data set, constructing an external memory bank based on the pre-collected data set, and constructing a policy model based on the Decision Transformer architecture, the following relational expressions exist in the corresponding process: ; Among them, represents the pre-collected data set, represents the th trajectory, represents the index of the trajectory, represents the total number of trajectories, respectively represent the 0th group to the The states, actions, rewards, and returns of the group, represent the length of the trajectory; Regarding the robot as an agent, use the policy model to predict the action distribution of the agent, and obtain the action distribution predicted by the policy model at the current moment. The following relationship exists in the corresponding process: ; Among them, represents the action at the current moment, represents the reinforcement learning policy at the current moment, represents the current moment, represents the state sequence, represents the return sequence, represents the action sequence, represents the reward sequence, represents the context length and .
[0014] It should be noted that the context length is used to define the sub-trajectory range. In Figure 2 , represents the actions at the previous moments of the current moment; The nearest neighbor retrieval algorithm is a similarity-based parameter-free machine learning method. Its core idea follows the inductive principle of "proximity samples have similar properties". Given a query vector and a dataset, the goal of the nearest neighbor retrieval algorithm is to find the samples of the required number that are closest to the query vector in the dataset, and form these samples into a neighborhood set.
[0015] Step 2: Based on the pre-collected dataset, sequentially process the input trajectory through the FH mechanism and the pre-trained language model to obtain a set of key-value pairs, and save the set of key-value pairs to an external memory bank.
[0016] In Step 2, based on the pre-collected dataset, sequentially process the input trajectory through the FH mechanism and the pre-trained language model to obtain a set of key-value pairs, and save the set of key-value pairs to an external memory bank. Specifically, it includes the following sub-steps: Based on the pre-collected dataset, use the FH mechanism to transform all sub-trajectories from time to to obtain the transformed all sub-trajectories from time to . The following relationship exists in the corresponding process: ; Among them, represents all the sub-trajectories from the transformed time to , Indicates that it has been processed by the FH mechanism, Indicates from the moment To All sub-trajectories of Indicates the embedding matrix, Indicates transpose, Indicates the scaling factor, Indicates the random matrix, Indicates that it has been processed by the softmax activation function; Using the pre-trained language model to extract features from all sub-trajectories of the converted from the moment To All sub-trajectories, obtaining the embedding vectors of all sub-trajectories from the moment To There is the following relational expression in the corresponding process: ; Among them, Indicates the embedding vectors of all sub-trajectories from the moment To All sub-trajectories of Indicates that it has been processed by the FH mechanism and the pre-trained language model in sequence, Indicates that feature extraction has been performed by the pre-trained language model; Based on the embedding vectors of all sub-trajectories from the moment To And all sub-trajectories from the moment To All sub-trajectories are constructed to obtain a key-value pair set, and the key-value pair set is saved to the external memory bank. There is the following relational expression in the corresponding process: ; Among them, Indicates the key-value pair set, Indicates the key set, Indicates the value set, Indicates from the moment To All sub-trajectories of
[0017] Step 3: Obtain the current input sub-trajectory, and retrieve similar sub-trajectories from the key-value pair set based on the current input sub-trajectory to obtain a set of retrieved similar sub-trajectories.
[0018] In step 3, obtain the current input sub-trajectory, and retrieve similar sub-trajectories from the key-value pair set based on the current input sub-trajectory to obtain a set of retrieved similar sub-trajectories, which specifically includes the following sub-steps: Obtain the current input sub-trajectory, and process the current input sub-trajectory through the FH mechanism and the pre-trained language model in sequence to obtain a query vector. There is the following relational expression in the corresponding process: ; Among them, represents the query vector, represents the current input sub-trajectory; Based on the key-value pair set, calculate the similarity between the key in each key-value pair and the query vector to obtain a set of retrieved similar sub-trajectories. There are the following relational expressions in the corresponding process: ; Among them, represents the set of retrieved similar sub-trajectories, represents after similarity calculation, represents the key, that is, the embedding vector of each sub-trajectory in the external memory index library.
[0019] Step 4: Based on the set of retrieved similar sub-trajectories, re-weight the similar sub-trajectories to obtain a sorted set of re-weighted sub-trajectories.
[0020] Please refer to Figure 3 , in Step 4, based on the set of retrieved similar sub-trajectories, re-weight the similar sub-trajectories to obtain a sorted set of re-weighted sub-trajectories, which specifically includes the following sub-steps: Based on the key-value pair set, calculate the cosine similarity between the query vector and the key to obtain a correlation score. There are the following relational expressions in the corresponding process: ; Among them, represents the correlation score, represents after cosine similarity calculation; In the training stage, calculate the similarity between the current input sub-trajectory and the previous relevant sub-trajectories to obtain the similarity between the input trajectory and the similar task. There are the following relational expressions in the corresponding process: ; Among them, represents the similarity between the input trajectory and the similar task, represents the previous relevant sub-trajectories, represents after processing by the task index function, represents after processing by the indicator function, that is, output 1 when the condition is satisfied, otherwise 0; In the inference stage, the policy model re-weights based on the retrieved sub-trajectories to obtain the magnitude of the reward in the input sub-trajectory and the similar sub-trajectories. There are the following relational expressions in the corresponding process: ; Among them, Represents the magnitude of the rewards in the input sub-trajectory and the similar sub-trajectory, which represents the reward for each step in the trajectory; The relevance score is re-weighted and fused with the similarity between the input trajectory and the similar task to obtain the final relevance score. There is the following relational expression in the corresponding process: ; where, represents the final relevance score, represents the weight, represents the sub-trajectory utility score adjusted by the weight ; Based on the final relevance score, the sub-trajectories in the set of retrieved similar sub-trajectories are re-ordered to obtain the set of re-weighted sub-trajectories after sorting. There is the following relational expression in the corresponding process: ; where, represents the set of re-weighted sub-trajectories after sorting, represents the top sub-trajectories sorted by similarity, represents the operation of sorting according to the magnitude of relevance.
[0021] It should be noted that in Figure 3 , represents the state at the current moment, represents the reward at the current moment; Through the design and improvement of the external memory mechanism, the present invention records the historical selections and historical reward situations of the agent in the task for subsequent retrieval and speculation in similar tasks; During the training process of the agent, by retrieving the state space similar to the current situation in the external memory mechanism, the past training information is fully utilized, and finally combined with the situation of the current task, the decision-making ability of the agent in the sparse reward environment is improved, thus solving the problems such as low sample utilization rate and difficult model convergence faced by traditional reinforcement learning methods.
[0022] Step 5: Based on the current input sub-trajectory and the set of re-weighted sub-trajectories after sorting, adjust the action distribution predicted by the policy model at the current moment to obtain the final action distribution predicted by the policy model, and use the final action distribution predicted by the policy model to optimize the policy model to obtain the optimized policy model; Use the optimized policy model to confirm the recommendation rate for the agent to execute each action.
[0023] In step 5, based on the current input sub-trajectory and the set of re-weighted sub-trajectories after sorting, adjust the action distribution predicted by the policy model at the current moment to obtain the final action distribution predicted by the policy model. There is the following relational expression in the corresponding process: ; wherein, represents the action distribution finally predicted by the policy model, represents the reinforcement learning policy.
[0024] It should be noted that a cross-attention layer is added after each self-attention layer of the policy model of the present invention to fuse the retrieved sub-trajectory information, avoiding the computational bottleneck of traditional Transformers in processing long sequences; the retrieved sub-trajectories are encoded by an independent embedding layer, and different embedding layers are used for each token type (state, dynamics, reward, and return), and the encoded sub-trajectories are passed to the cross-attention layer.
[0025] It should be understood that although the steps in the flowcharts of the various embodiments of the present invention are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0026] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0027] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0028] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.
Claims
1. A path coverage method based on a retrieval enhancement mechanism, characterized in that The method includes the following steps: Step 1: Obtain a pre-collected dataset, construct an external memory bank based on the pre-collected dataset, construct a policy model based on the Decision Transformer architecture, regard the robot as an agent, and use the policy model to predict the action distribution of the agent to obtain the action distribution predicted by the policy model at the current moment; Step 2: Based on the pre-collected dataset, sequentially process the input trajectory through the FH mechanism and the pre-trained language model to obtain a set of key-value pairs, and save the set of key-value pairs to the external memory bank; Step 3: Obtain the current input sub-trajectory, and retrieve similar sub-trajectories from the set of key-value pairs based on the current input sub-trajectory to obtain a set of retrieved similar sub-trajectories; Step 4: Based on the set of retrieved similar sub-trajectories, re-weight the similar sub-trajectories to obtain a sorted set of re-weighted sub-trajectories; Step 5: Based on the current input sub-trajectory and the sorted set of re-weighted sub-trajectories, adjust the action distribution predicted by the policy model at the current moment to obtain the finally predicted action distribution of the policy model, and use the finally predicted action distribution of the policy model to optimize the policy model to obtain an optimized policy model; Use the optimized policy model to confirm the recommendation rate of each action executed by the agent.
2. The path coverage method based on a retrieval enhancement mechanism according to claim 1, wherein In Step 1, when obtaining the pre-collected dataset, constructing the external memory bank based on the pre-collected dataset, and constructing the policy model based on the Decision Transformer architecture, the following relational expressions exist in the corresponding process: ; Among them, represents a pre-collected dataset, represents the th trajectory, represents the index of the trajectory, represents the total number of trajectories, respectively represent the state, action, reward, and return from the 0th group to the th group, represents the length of the trajectory; In the step of regarding the robot as an agent and using the policy model to predict the action distribution of the agent to obtain the action distribution predicted by the policy model at the current moment, the following relational expressions exist in the corresponding process: ; wherein, represents the action at the current moment, represents the reinforcement learning policy at the current moment, represents the current moment, represents the state sequence, represents the reward sequence, represents the action sequence, represents the reward sequence, represents the context length.
3. The path coverage method based on a retrieval enhancement mechanism according to claim 2, wherein In Step 2, based on the pre-collected dataset, sequentially process the input trajectory through the FH mechanism and the pre-trained language model to obtain a set of key-value pairs, and save the set of key-value pairs to the external memory bank. Specifically, it includes the following sub-steps: Based on a pre-collected data set, use the FH mechanism to transform all sub-trajectories from time to to obtain all the transformed sub-trajectories from time to ; Use the pre-trained language model to extract features from all sub-trajectories from the converted starting time to to obtain the embedding vectors of all sub-trajectories from the starting time to ; Based on the embedding vectors of all sub-trajectories from time to and all sub-trajectories from time to are constructed to obtain a set of key-value pairs, and the set of key-value pairs is saved to an external memory bank.
4. The path coverage method based on the retrieval enhancement mechanism according to claim 3, wherein Based on the pre-collected data set, use the FH mechanism to transform all sub-trajectories from time to to obtain all sub-trajectories after transformation from time to . The following relational expressions exist in the corresponding process: ; Among them, represents all sub-trajectories from the converted time to ; represents being processed by the FH mechanism, represents all sub-trajectories from the time to ; represents the embedding matrix, represents the transpose, represents the scaling factor, represents the random matrix, represents being processed by the softmax activation function; In the step of extracting features from all sub-trajectories from the converted starting time to to obtain the embedding vectors of all sub-trajectories from the starting time to in the corresponding process, there is the following relational expression: ; Among them, represents the embedding vectors of all sub-trajectories from time to ; represents being processed by the FH mechanism and the pre-trained language model in sequence, represents feature extraction by the pre-trained language model. In the step of constructing a set of key-value pairs based on the embedding vectors of all sub-trajectories from time to and all sub-trajectories from time to and saving the set of key-value pairs to an external memory bank, there is the following relational expression in the corresponding process: ; Among them, represents a set of key-value pairs, represents a set of keys, represents a set of values, represents all sub-trajectories from time to 5. The path coverage method based on a retrieval enhancement mechanism according to claim 4, wherein In Step 3, when obtaining the current input sub-trajectory and retrieving similar sub-trajectories from the set of key-value pairs based on the current input sub-trajectory to obtain a set of retrieved similar sub-trajectories, it specifically includes the following sub-steps: Obtain the current input sub-trajectory, and sequentially process the current input sub-trajectory through the FH mechanism and the pre-trained language model to obtain a query vector; Based on the set of key-value pairs, calculate the similarity between the key in each key-value pair and the query vector to obtain a set of retrieved similar sub-trajectories.
6. The path coverage method based on the retrieval enhancement mechanism according to claim 5, wherein When obtaining the current input sub-trajectory and sequentially processing the current input sub-trajectory through the FH mechanism and the pre-trained language model to obtain a query vector, the following relational expressions exist in the corresponding process: ; Among them, represents the query vector, represents the current input sub-trajectory; In the step of calculating the similarity between the key in each key-value pair and the query vector based on the set of key-value pairs to obtain a set of retrieved similar sub-trajectories, the following relational expressions exist in the corresponding process: ; Among them, represents the set of retrieved similar sub-trajectories, indicates that after similarity calculation, represents a key.
7. The path coverage method based on a retrieval enhancement mechanism according to claim 6, characterized in that, In Step 4, based on the set of retrieved similar sub-trajectories, re-weight the similar sub-trajectories to obtain a sorted set of re-weighted sub-trajectories. Specifically, it includes the following sub-steps: Based on the set of key-value pairs, calculate the cosine similarity between the query vector and the keys to obtain the correlation scores; During the training phase, the current input sub-trajectory is compared with the previous related sub-trajectories to calculate the similarity between the input trajectory and the similar tasks; In the inference stage, the policy model reweights based on the retrieved sub-trajectories to obtain the magnitudes of the rewards in the input sub-trajectories and the similar sub-trajectories; Reweight and fuse the correlation scores with the similarities of the input trajectories and the similar tasks to obtain the final correlation scores; Based on the final correlation scores, perform a reordering operation on the sub-trajectories in the set of retrieved similar sub-trajectories to obtain the set of reweighted sub-trajectories after sorting; 8. The path coverage method based on a retrieval enhancement mechanism according to claim 7, wherein Based on the set of key-value pairs, calculate the cosine similarity between the query vector and the keys to obtain the correlation scores, and there are the following relational expressions in the corresponding process: ; Among them, represents the correlation score, indicating that it is calculated by cosine similarity; During the training phase, when calculating the similarity between the current input sub-trajectory and the previous relevant sub-trajectories to obtain the similarity between the input trajectory and the similar tasks, the following relational expressions exist in the corresponding process: ; Among them, represents the similarity between the input trajectory and similar tasks, represents the previous number of relevant sub-trajectories, represents being processed by the task index function, represents being processed by the indication function; In the step where, in the inference stage, the policy model reweights based on the retrieved sub-trajectories to obtain the magnitudes of the rewards in the input sub-trajectories and the similar sub-trajectories, there are the following relational expressions in the corresponding process: ; Among them, represents the magnitude of the rewards in the input sub-trajectory and the similar sub-trajectory, represents the reward for each step in the trajectory; In the step of reweighting and fusing the correlation scores with the similarities of the input trajectories and the similar tasks to obtain the final correlation scores, there are the following relational expressions in the corresponding process: ; Among them, represents the final relevance score, represents the weight, represents the sub-trajectory utility score adjusted by the weight ; In the step of performing a reordering operation on the sub-trajectories in the set of retrieved similar sub-trajectories based on the final correlation scores to obtain the set of reweighted sub-trajectories after sorting, there are the following relational expressions in the corresponding process: ; Among them, represents the set after reweighted sub-trajectory sorting, represents the top sub-trajectories sorted by similarity, represents the operation of sorting by the magnitude of correlation.
9. The path coverage method based on a retrieval enhancement mechanism according to claim 8, wherein In step 5, based on the current input sub-trajectory and the set of reweighted sub-trajectories after sorting, adjust the action distribution predicted by the policy model at the current moment to obtain the final action distribution predicted by the policy model, and there are the following relational expressions in the corresponding process: ; Among them, represents the action distribution finally predicted by the policy model, represents the reinforcement learning policy.
Citation Information
Patent Citations
Reinforcement learning strategy enhancement method based on memory retrieval
CN117035050A
Building energy consumption optimization method based on Decision Transform model
CN118569313A
Mobile robot strategy simulation method based on decision Transform
CN118760163A
Base station energy consumption optimization method based on improved Decision Transform model
CN118945684A
Industrial control protocol fuzz testing method and system based on large language model
CN119415391A