Reinforcement learning method, system and device based on sub-target discovery and storage medium
By estimating the causal value of each state of the system and selecting a state with high causal value as a sub-objective, the problem of lack of guidance for complex task decomposition and agent exploration in the existing reinforcement learning methods is solved, and the task execution effect of the agent is improved.
Patent Information
- Application Number
- CN202510221500.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
When existing reinforcement learning methods deal with complex tasks, it is difficult to effectively decompose tasks, resulting in a lack of guidance in the exploration process of the agent and making training difficult.
By estimating the causal value of each state of the system, selecting a state with high causal value as the sub-target, and using the sub-target relationship diagram to guide the agent to complete complex tasks.
It improves the exploration efficiency and final goal completion rate of the agent, reduces the dependence on prior knowledge, and can find sub-objections with physical significance without introducing prior knowledge.
Smart Images

Figure CN120163202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of reinforcement learning, and in particular to a reinforcement learning method, system, device and storage medium based on sub-goal discovery. Background Art
[0002] Reinforcement learning has been widely applied to various intelligent agent decision-making problems and has achieved great success in fields such as games, autonomous driving, and robotics. In real scenarios, the rewards provided by the system are usually sparse. The intelligent agent can only know whether the final task is completed, and cannot obtain the guidance of rewards during the previous exploration process. Therefore, the exploration of the intelligent agent is usually blind, and it can only complete the final task with a very low probability, which greatly increases the training difficulty of reinforcement learning.
[0003] In order to encourage the intelligent agent to efficiently explore the environment, existing methods usually decompose complex tasks into multiple subtasks by referring to the idea of humans completing complex tasks. By training the intelligent agent to complete the subtasks one by one, the final task is then completed. However, there are still many difficulties in decomposing complex tasks while ensuring that the subtasks are effective for assisting in completing the final task.
[0004] The existing mainstream reinforcement learning training paradigm methods based on sub-goals mainly include experience replay technology and artificial prior knowledge. The experience replay technology encourages the intelligent agent to learn how to reach the non-final goal states it has reached before, enabling the intelligent agent to master the ability to return to historical states. The method using artificial prior knowledge requires humans to deconstruct the task, design subtasks or sub-goals, and provide them for the intelligent agent to train. Specifically:
[0005] (1) The experience replay technology encourages the intelligent agent to learn how to reach the non-final goal states it has reached before, enabling the intelligent agent to master the ability to return to historical states. Specifically: taking the historical states where the intelligent agent has not reached the final goal as subtasks or sub-goals, enabling the intelligent agent to learn to return to a certain state it has reached before, and improving the utilization efficiency of experience data. However, the sub-goals obtained by this method do not have clear guiding significance, and are more about the utilization of historical experience. Its sub-goals cannot guarantee a positive effect on the intelligent agent to complete the final goal.
[0006] (2) The method using artificial prior knowledge is to use prior knowledge to design a scheme for subtasks or sub-goals manually. However, its dependence on prior knowledge makes it difficult to be promoted in various different tasks or scenarios. Moreover, the present invention hopes that the intelligent agent can learn autonomously and complete tasks. Introducing too much prior knowledge does not conform to the goal of the present invention.
[0007] The above disadvantages are the main problems of the prior art. Therefore, it is necessary to design a sub-goal discovery technique that does not require prior knowledge, and at the same time ensure that the discovered sub-goals have clear physical meanings, so that they can provide clear guidance for the agent to complete complex tasks.
[0008] In view of this, the present invention is specifically proposed. Summary of the Invention
[0009] The object of the present invention is to provide a reinforcement learning method, system, device and storage medium based on sub-goal discovery, estimate the causal value of each state of the system, select the state with high value as the sub-goal, disassemble the complex task into simple tasks based on the sub-goal, so that it can provide clear guidance for the agent to complete complex tasks, thereby improving the task execution effect.
[0010] The object of the present invention is achieved by the following technical solutions:
[0011] A reinforcement learning method based on sub-goal discovery, applied to navigation tasks, includes:
[0012] Pre-training stage: Use a random policy to sample a trajectory data sequence, estimate the causal capacity of each state to which the trajectory data belongs, and select the state whose causal capacity reaches the threshold as the sub-goal of the agent. Then, according to the reachability relationship between sub-goals, where the causal capacity of a state is the entropy of the intervention-free state transition probability distribution; introduce a sub-goal predictor for predicting the sub-goal corresponding to the input state, and through the sub-goal relationship graph, determine the sub-goals that each state to which the trajectory data belongs can reach, and use this as the prediction target of the sub-goal predictor to train the sub-goal predictor;
[0013] Training stage: In each state, predict the sub-goal through the trained sub-goal predictor, and the agent makes an action decision by combining the state and the predicted sub-goal, and then trains the agent by using the reinforcement learning method in combination with the sub-goal relationship graph.
[0014] A reinforcement learning system based on sub-goal discovery, includes:
[0015] Pre-training unit, applied to the pre-training stage, includes: Use a random policy to sample a trajectory data sequence, estimate the causal capacity of each state to which the trajectory data belongs, and select the state whose causal capacity reaches the threshold as the sub-goal of the agent. Then, according to the reachability relationship between sub-goals, where the causal capacity of a state is the entropy of the intervention-free state transition probability distribution; introduce a sub-goal predictor for predicting the sub-goal corresponding to the input state, and through the sub-goal relationship graph, determine the sub-goals that each state to which the trajectory data belongs can reach, and use this as the prediction target of the sub-goal predictor to train the sub-goal predictor;
[0016] A training unit, applied in the training phase, includes: in each state, predicting sub-goals through a trained sub-goal predictor, and enabling an agent to make action decisions by combining the state and the predicted sub-goals, and then training the agent by using the method of reinforcement learning in combination with the sub-goal relationship graph.
[0017] A processing device includes: one or more processors; a memory for storing one or more programs;
[0018] Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method.
[0019] A readable storage medium stores a computer program, and when the computer program is executed by a processor, the foregoing method is implemented.
[0020] As can be seen from the technical solutions provided by the present invention above, considering the causal relationship between the behavior of the agent and its results, physical meaningful states can be found as sub-goals without introducing prior knowledge, encouraging the agent to master the causal relationship between its own behavior and the system state transition, and improving the exploration efficiency. In addition, the framework of the present invention as a pre-training method takes very little time and can be combined with most existing reinforcement learning methods to improve the overall performance of the agent. Description of the Drawings
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 It is a schematic diagram of the overall framework of a reinforcement learning method based on sub-goal discovery provided by an embodiment of the present invention;
[0023] Figure 2 It is a schematic diagram of the training process of an encoder and a decoder provided by an embodiment of the present invention;
[0024] Figure 3 It is a schematic diagram of the training process of a sub-goal predictor provided by an embodiment of the present invention;
[0025] Figure 4 It is a schematic diagram of a reinforcement learning system based on sub-goal discovery provided by an embodiment of the present invention;
[0026] Figure 5 It is a schematic diagram of a processing device provided by an embodiment of the present invention. Detailed Embodiments
[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] First, the terms that may be used in this article are described as follows:
[0029] The description of terms such as "comprising", "including", "containing", "having" or other similar semantics shall be interpreted as non-exclusive inclusion. For example, including a certain technical feature element (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, processes, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or articles, etc.) shall be interpreted as not only including the clearly listed certain technical feature element, but also including other technical feature elements well known in the art that are not clearly listed.
[0030] The term "consisting of" means excluding any technical feature element that is not clearly listed. If this term is used in a claim, this term will make the claim a closed type, making it not include technical feature elements other than the clearly listed technical feature elements, except for related conventional impurities. If this term only appears in a certain clause of the claim, then it only limits the elements clearly listed in that clause, and the elements recorded in other clauses are not excluded from the overall claim.
[0031] The following provides a detailed description of a reinforcement learning method, system, device, and storage medium based on sub-goal discovery provided by the present invention. The content not described in detail in the embodiments of the present invention belongs to the prior art well known to those of ordinary skill in the art. Conditions not specified in the embodiments of the present invention are carried out according to conventional conditions in the art or conditions recommended by manufacturers. Instruments not specified by the manufacturer in the embodiments of the present invention are all conventional products that can be obtained through commercial purchase.
[0032] Embodiment 1
[0033] The embodiment of the present invention provides a reinforcement learning method based on sub-goal discovery, which can be effectively applied to many scenarios such as autonomous driving and indoor navigation. As Figure 1 shown, the present invention mainly includes two parts:
[0034] (1) Pre-training stage: In this stage, a random policy is used to sample a sequence of trajectory data (mainly including the spatial coordinate information of the agent), estimate the causal capacity of the state to which each trajectory data belongs, and select the state whose causal capacity reaches the threshold as the sub-goal of the agent. Then, according to the reachability relationship between sub-goals, a sub-goal relationship graph is constructed. Here, the causal capacity of a state is the entropy of the non-intervention state transition probability distribution of the state; a sub-goal predictor is introduced to predict the sub-goal corresponding to the input state. Through the sub-goal relationship graph, determine the sub-goals that the state to which each trajectory data belongs can reach, and use this as the prediction target of the sub-goal predictor to train the sub-goal predictor.
[0035] The solution provided by the embodiments of the present invention can be applied to navigation tasks. Therefore, the trajectory data only needs the spatial information of the agent. The state mainly includes the spatial information of the agent and the physical information for performing actions (such as acceleration, etc.). The actions are mainly defined by the environment and are the operations that the agent can perform.
[0036] (2) Training stage: In each state, the trained sub-goal predictor is used to predict the sub-goal, and the agent combines the state and the predicted sub-goal to make an action decision, and then combines the sub-goal relationship graph to train the agent using the reinforcement learning method.
[0037] In the above solution, in the pre-training stage, the main purpose is to train the sub-goal predictor so that it can predict the best sub-goal according to the state. In the training stage, the trained sub-goal predictor is used to predict the sub-goal to guide the agent to act in the direction of the sub-goal. After that, the final goal can be achieved by combining the sub-goal relationship graph. The reinforcement learning method involved in this part can be implemented using existing solutions, and the present invention will not elaborate.
[0038] The above solution provided by the embodiments of the present invention considers the causal relationship between the agent's behavior and its results. Without introducing prior knowledge, it can find the state with physical meaning as the sub-goal, encourage the agent to master the causal relationship between its own behavior and the system state transition, and improve the exploration efficiency; moreover, it can be combined with most existing reinforcement learning methods to improve the task execution effect of the trained agent, such as the navigation performance in navigation tasks.
[0039] In order to more clearly show the technical solution provided by the present invention and the technical effects produced, the following uses specific embodiments to describe in detail the method provided by the embodiments of the present invention.
[0040] I. Overall overview of the solution.
[0041] An embodiment of the present invention provides a reinforcement learning method based on sub-goal discovery, which considers the causal relationship between the agent's behavior and its results, estimates the causal value of each state of the system, and selects states with high value as sub-goals, enabling it to provide clear guidance for the agent to complete complex tasks. Specifically: To solve the technical problems existing in the aforementioned prior art, the present invention designs an index for evaluating the causal value of each state of the system - causal capacity, that is, the entropy of the agent's state transition probability distribution without intervention in the system. By calculating the causal capacity of each state of the system, states with high causal value can be evaluated and selected as the sub-goals of the agent. The actions performed in these states with high causal value can have the greatest causal impact on the subsequent state changes of the agent. By mastering each sub-goal in the system, the agent can effectively improve its own exploration efficiency and the final goal completion rate.
[0042] II. Detailed introduction of the solution.
[0043] 1. Pre-training stage.
[0044] It can also be seen in Figure 1 , in the pre-training stage, the agent first samples pre-training data (trajectory data) in the system using a random policy to estimate the causal capacity of each state. The causal capacity is defined as the entropy of the state transition probability distribution without intervention of the system state, representing the upper limit of the causal value that all behaviors of the agent can generate in this state. The specific formula is as follows:
[0045]
[0046] Among them, p(·) is the state transition probability function, is the information entropy, S is the current state, S′ is the set of next states that state s can reach, is the causal capacity of state s.
[0047] To accurately estimate the causal capacity, the present invention designs a clustering algorithm based on the distance function d(·,·), and selects the adjacent state set as the generalized next state set:
[0048]
[0049] Among them, S nei (s) is the set of neighbor states of state s, and there is no state transition in the physical sense. S out is the set of distant states, including states that cannot be reached within one-step state transition. S adj (s) is the set of adjacent states of state s, where the states have both an obvious physical meaning change and can be reached within one-step state transition. τ nei and τ adj are distance thresholds.
[0050] After that, cluster the set of adjacent states based on the distance function d(·,·), and estimate the probability distribution of state transitions according to the frequencies of states within each category in the clustering results, and then calculate the causal capacity of the states.
[0051] Based on the calculation results of the causal capacity, the present invention selects several states with high causal capacity (i.e., exceeding a set threshold) in the system as sub-goals of the agent, and can construct a sub-goal relationship graph of the system according to the reachable relationship between sub-goals obtained from the trajectory data. During the training process, when the agent reaches the sub-goal predicted by the sub-goal predictor, it can select the next sub-goal to go to according to the relationship graph. The threshold involved here can be set by the user according to the actual situation or experience, and the present invention does not make specific numerical limitations.
[0052] In the embodiments of the present invention, the sub-goals are states selected by causal capacity. Therefore, the reachable relationship between sub-goals can be determined through the trajectory data sequence. For example, if a certain trajectory data reaches sub-goal B from sub-goal A, it is considered that sub-goals A and B are reachable.
[0053] At this stage, the present invention designs a set of state encoder-decoders and a sub-goal predictor, all of which only contain fully connected networks.
[0054] As Figure 2 shown, it is the structure of encoder p θ and decoder p φ which are trained in a self-supervised learning manner. The encoder projects the input state and sub-goal into the latent space, and the decoder reconstructs the state from the latent space vector of the input state. During training, minimize the distance between the input state and the reconstructed state (to facilitate the agent to distinguish), and at the same time, maximize the similarity of the latent space vectors of different sub-goals. The corresponding loss function is expressed as:
[0055]
[0056] where s l is the input state, s' l is the reconstructed state corresponding to s l , is the set composed of input states; are the latent space vectors of two sub-goals, λ θ and λ φ are the weights of the two parts of the loss function, sim is the similarity function. For example, the cosine similarity distance function can be adopted, and θ and φ respectively represent the parameters of the encoder and decoder.
[0057] In the embodiments of the present invention, the reconstruction results of the decoder are applicable to self-supervised training. When projecting the state into the latent space, the characteristics between the original states must be preserved, so it is required that the projected state can be reconstructed back to the original space via the decoder. Sub-goals (sub-goals are also states themselves) participate in self-supervised training just like states. However, for sub-goals, there is an additional loss function of maximizing the distance.
[0058] The encoder and the decoder are jointly trained through the above loss function. After the training is completed, the encoder is retained for the training process of the sub-goal predictor.
[0059] Design the sub-goal predictor p θ , and according to the sequence information in the pre-training data, predict the optimal sub-goal for each state. This enables the intelligent agent to clearly identify the current sub-goal to which it should go in any state of the system. For example, Figure 3 as shown, for a trajectory data sequence τ of length T, when there is a state s t+m reaching a certain sub-goal then the predicted goal before the state is the encoding result of the sub-goal ; if no state reaches any sub-goal, the predicted goal for the entire sequence is the encoding result z corresponding to the last state s t+T-1 ; by constraining the sub-goal predictor p t+T-1 to optimize the similarity between the output for each state and the sub-goal corresponding to the state, the corresponding loss function θ is expressed as:
[0060]
[0061] where t is the starting time of the sequence, m = 0,..., T - 1, ρ θ (s i ) represents the output of the sub-goal predictor p θ for the state s i , that is, the encoding result corresponding to the sub-goal of the predicted state s i , ρ θ (s t+T-1 ) represents the encoding result of the sub-goal of the encoder p θ for the sub-goal of the last state s t+T-1 , and sim is the similarity function.
[0062] After the sub-goal predictor is trained, the intelligent agent can find the current optimal sub-goal based on the trained sub-goal predictor in any state, and according to the sub-goal relationship graph, clarify the sequence of sub-goals passed through in the optimal path from the current state to the final goal.
[0063] In the embodiments of the present invention, a loss function corresponding to a codec and a sub-goal predictor is provided. Therefore, the subsequent training process can be completed in combination with the prior art, and the present invention will not elaborate on it.
[0064] 2. Training stage.
[0065] Next, the agent enters the training stage of executing policy π. Whenever the system provides a state s, the agent will predict a corresponding sub-goal θ using the sub-goal predictor p and use it together with s as the input of the execution policy, and then obtain an action a. The execution process can be expressed by the following formula:
[0066]
[0067] where, represents the sub-goal predictor p θ predicting the sub-goal π is the execution policy of the agent, which is used to decide the action a.
[0068] After that, the agent executes the action a, obtains a new state and reward, and then combines the sub-goal to decide a new action, and repeats continuously until the sub-goal is reached Then, determine the next sub-goal in combination with the sub-goal relationship graph and make an action decision until the set final sub-goal is reached; combine the rewards obtained after each execution of the action, and train the agent in the way of reinforcement learning. Considering that the reinforcement learning scheme involved in this part can be implemented by conventional techniques, it will not be elaborated here.
[0069] In the above embodiments provided by the present invention, any existing reinforcement learning method can be used to learn the execution policy, and the overall framework has a high adaptability. When the execution policy masters the ability to go from any state to the optimal sub-goal, it can complete the sub-goals one by one according to the optimal sequence of completing the final goal based on the sub-goal relationship graph, and then complete the final goal.
[0070] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0071] Embodiment 2
[0072] The present invention also provides a reinforcement learning system based on sub-goal discovery, which is mainly used to implement the method provided in the foregoing embodiments, such as Figure 4 As shown, the system mainly includes:
[0073] A pre-training unit, which is applied to the pre-training stage, including: sampling a trajectory data sequence using a random policy, estimating the causal capacity of the state to which each trajectory data belongs, and selecting the state whose causal capacity reaches the threshold as the sub-goal of the intelligent agent. Then, according to the reachable relationship between the sub-goals, where the causal capacity of a state is the entropy of the non-intervention state transition probability distribution of the state; introducing a sub-goal predictor for predicting the sub-goal corresponding to the input state, determining the sub-goals that the state to which each trajectory data belongs can reach through the sub-goal relationship graph, and using this as the prediction target of the sub-goal predictor to train the sub-goal predictor;
[0074] A training unit, which is applied to the training stage, including: in each state, predicting the sub-goal through the trained sub-goal predictor, and the intelligent agent making an action decision by combining the state and the predicted sub-goal, and then training the intelligent agent using the method of reinforcement learning in combination with the sub-goal relationship graph.
[0075] Considering that the above pre-training process and training process have been introduced in detail in the previous embodiments, they will not be elaborated here.
[0076] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.
[0077] Embodiment III
[0078] The present invention also provides a processing device, such as Figure 5 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.
[0079] Further, the processing device further includes at least one input device and at least one output device; in the processing device, the processor, the memory, the input device, and the output device are connected through a bus.
[0080] In the embodiments of the present invention, the specific types of the memory, the input device, and the output device are not limited; for example:
[0081] The input device can be a touch screen, an image acquisition device, a physical button, or a mouse, etc.;
[0082] The output device can be a display terminal;
[0083] The memory can be a Random Access Memory (RAM), or a non-volatile memory, such as a disk memory.
[0084] Embodiment 4
[0085] The present invention also provides a readable storage medium storing a computer program, which when executed by a processor implements the method provided in the foregoing embodiments.
[0086] In the embodiments of the present invention, the readable storage medium, as a computer-readable storage medium, can be disposed in the foregoing processing device, for example, as the memory in the processing device. In addition, the readable storage medium can also be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a Read-Only Memory (ROM), a magnetic disk, or an optical disc.
[0087] As described above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims. The information disclosed in the background art part of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes the prior art known to those skilled in the art.
Claims
1. A reinforcement learning method based on sub-goal discovery, applied to navigation tasks, characterized in that: include: Pre-training stage: Use random strategy to sample trajectory data sequence, estimate the causal capacity of the state to which each trajectory data belongs, and select the state whose causal capacity reaches the threshold as the sub-goal of the agent. Then, according to the reachable relationship between each sub-goal, the causal capacity of the state is the entropy of the state's non-intervention state transition probability distribution; introduce a sub-goal predictor for predicting the sub-goal corresponding to the input state, determine the sub-goal that can be reached by the state to which each trajectory data belongs through the sub-goal relationship graph, and use this as the prediction target of the sub-goal predictor to train the sub-goal predictor; Training phase: In each state, the sub-goal predictor is used to predict the sub-goal, and the agent makes action decisions based on the state and the predicted sub-goal. The agent is then trained using reinforcement learning methods combined with the sub-goal relationship graph.
2. A reinforcement learning method based on sub-goal discovery according to claim 1, characterized in that: The method for estimating the causal capacity of each state is: Among them, p(·) is the state transition probability function, is the information entropy, S is the current state, S′ is the set of next states that state s can reach, is the causal capacity of state s.
3. A reinforcement learning method based on sub-goal discovery according to claim 1 or 2, characterized in that: The unintervention state transition distribution of a state is calculated as follows: Design a clustering algorithm based on the distance function d(·,·) and select the adjacent state set as the generalized next state set: Among them, S nei (s) is the set of neighbor states of state s, S adj (s) is the set of adjacent states of state s, S out is a set of distant states, including states that cannot be reached within one-step state transition, τ nei and τ adj is the distance threshold; The adjacent state set is clustered based on the distance function d(·,·), and the frequency of states in each category in the clustering result is calculated to estimate the probability distribution of state transitions.
4. The reinforcement learning method based on sub-goal discovery according to claim 1, characterized in that: The sub-goal predictor for predicting the sub-goal corresponding to the input state is introduced, and the sub-goal that can be reached by the state of each trajectory data is determined through the sub-goal relationship graph, and this is used as the prediction target of the sub-goal predictor. Training the sub-goal predictor includes: For a trajectory data sequence τ of length T, when there is a state s t+m Reach a sub-goal Then the state s t+m The previous prediction target is a sub-target The encoding result of If no state reaches any sub-goal, the prediction target of the entire sequence is the last state s t+T-1 The corresponding encoding result z t+T-1 ; By constraining the subgoal predictor The similarity between each state output and the sub-goal corresponding to the state is optimized, and the corresponding loss function It is expressed as: Where t is the starting time of the sequence, m = 0,…,T-1, represents the subgoal predictor Prediction status i The encoding result corresponding to the sub-goal of Denotes the encoder p θ Sub-goals The last state t+T-1 The encoding result of the sub-goal, sim is the similarity function.
5. A reinforcement learning method based on sub-goal discovery according to claim 4, characterized in that: The encoder p θ Pre-training is done using self-supervised learning; Set the encoder and decoder structure. The encoder projects the input state and sub-goals into the latent space. The decoder reconstructs the latent space vector of the input state. During training, the distance between the input state and the reconstructed state is minimized, and the similarity of the latent space vectors of different sub-goals is maximized. The corresponding loss function It is expressed as: Among them, s l is the input state, s l ′ For l The corresponding reconstruction state, is the set of input states; is the latent space vector of the two sub-goals, λ θ and λ φ is the weight, sim is the similarity function; The encoder and decoder are jointly trained using the above loss function. After the training is completed, the encoder is retained for the training process of the sub-target predictor.
6. The reinforcement learning method based on sub-goal discovery according to claim 1, characterized in that: The method of predicting the sub-goal by the trained sub-goal predictor and making an action decision by the agent based on the state and the predicted sub-goal includes: For state s, the sub-goal predicted by the sub-goal predictor is recorded as in, is the sub-goal predictor, is the sub-goal; Combine state s with subgoal As the output of the agent, we get the action Among them, π is the execution strategy of the agent, which is used to decide action a.
7. The reinforcement learning method based on sub-goal discovery according to claim 1, characterized in that: The method of combining the sub-goal relationship graph with the reinforcement learning method to train the intelligent agent includes: After the agent makes an action decision based on the state and the predicted sub-goal, it executes the action and obtains the new state and reward. It then makes a new action based on the predicted sub-goal and repeats this process until the predicted sub-goal is reached. Then, it determines the next sub-goal based on the sub-goal relationship graph and makes an action decision until the final sub-goal is reached. Combined with the rewards obtained after each action is performed, the agent is trained using reinforcement learning.
8. A reinforcement learning system based on sub-goal discovery, characterized in that: include: A pre-training unit is applied in the pre-training stage, including: using a random strategy to sample a trajectory data sequence, estimating the causal capacity of the state to which each trajectory data belongs, and selecting the state whose causal capacity reaches a threshold as the sub-goal of the intelligent agent, and then according to the reachable relationship between the sub-goals, wherein the causal capacity of the state is the entropy of the probability distribution of the state's non-intervention state transition; introducing a sub-goal predictor for predicting the sub-goal corresponding to the input state, determining the sub-goal that can be reached by the state to which each trajectory data belongs through a sub-goal relationship graph, and using this as the prediction target of the sub-goal predictor, and training the sub-goal predictor; The training unit is used in the training phase, including: in each state, the sub-goal is predicted by the trained sub-goal predictor, and the intelligent agent makes action decisions based on the state and the predicted sub-goal, and then the intelligent agent is trained using the reinforcement learning method in combination with the sub-goal relationship graph.
9. A processing device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.
10. A readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.